Voice and Speech AI

Course Overview
Intermediate
Free Course

For engineers who want to add speech to their products. You will be able to build reliable transcription pipelines with Whisper, generate and responsibly clone voices with neural text-to-speech, and assemble a low-latency voice assistant.

Instructor: Jaidev
Sections: 3

Course Content

Section 1: Speech Recognition

How speech recognition turns audio into text, how Whisper works, and how to run transcription reliably on long, messy, real recordings. 2 lessons, about 35 minutes.

Section 2: Text-to-Speech and Audio Generation

How neural text-to-speech and voice cloning work, how to measure them, and the consent and safety controls that cloning requires. 2 lessons, about 35 minutes.

Section 3: Mini Project

Build a low-latency voice-to-voice assistant, from microphone capture to interruptible speech. 1 lesson, about 15 minutes.