- MantraMindAI
- Courses
- Multimodal AI
- Voice and Speech AI
Voice and Speech AI
Course Overview
Intermediate
Free Course
For engineers who want to add speech to their products. You will be able to build reliable transcription pipelines with Whisper, generate and responsibly clone voices with neural text-to-speech, and assemble a low-latency voice assistant.
Instructor: Jaidev
Sections: 3
Course Content
Section 1: Speech Recognition
How speech recognition turns audio into text, how Whisper works, and how to run transcription reliably on long, messy, real recordings. 2 lessons, about 35 minutes.
Section 2: Text-to-Speech and Audio Generation
How neural text-to-speech and voice cloning work, how to measure them, and the consent and safety controls that cloning requires. 2 lessons, about 35 minutes.
Section 3: Mini Project
Build a low-latency voice-to-voice assistant, from microphone capture to interruptible speech. 1 lesson, about 15 minutes.