Course Content
Voice and Speech AI
3 sections · 5 lessons
Mini-Project: Build a Voice-to-Voice Assistant
Four components, all of which you can get working in an afternoon. Capture microphone audio. Transcribe it. Send the text to a language model. Speak the answer. Wire them in a loop and you have a voice assistant.
Build it that way and then use it, and the experience is dreadful. You finish speaking. Nothing happens. You wonder whether it heard you. You start to repeat yourself — and then it starts talking, over you. You wait for it to finish so you can correct it, but it delivers all four sentences of an answer that stopped being relevant after the first one.
Nothing is broken. Every component is working exactly as designed. The measured round trip is 6.3 seconds, and that number is the entire problem.
Human conversation runs on a much tighter clock than people realise. Studies of turn-taking across ten unrelated languages found the median gap between one speaker finishing and the next beginning is around 200 milliseconds — less than the time it takes to plan a word, which means listeners are predicting the end of your sentence and preparing their reply before you get there. Push a response past about 800 ms and it registers as hesitation. Past two seconds and people assume the line has dropped and start talking again.
So the interesting engineering in a voice assistant is not any of the four components. It is the 6.1 seconds you have to remove from between them.
Where the 6.3 seconds actually goes
Assume a realistic turn: the user speaks for 5 seconds, and the assistant replies with 40 tokens, roughly 30 words, which at a normal 150 words per minute is 12 seconds of speech.
| Stage | Naive | Optimised | What changed |
|---|---|---|---|
| End-of-turn detection | 2,000 ms | 500 ms | VAD with 500 ms hangover, not a fixed 2 s timeout |
| Audio to server | 150 ms | 50 ms | Stream during speech instead of after it |
| ASR | 750 ms | 180 ms | faster-whisper int8; decode as audio arrives |
| LLM first token | 600 ms | 350 ms | Smaller model, cached system prompt |
| LLM remaining tokens | 1,600 ms | 0 ms | Hidden behind playback of the first sentence |
| TTS | 960 ms | 120 ms | Synthesise first sentence only, then stream |
| Transport + playback buffer | 250 ms | 80 ms | Opus, small jitter buffer |
| Time to first audio | 6,310 ms | 1,280 ms | 4.9× faster |
Check the two rows that matter most, because they are where the counter-intuitive wins live.
End-of-turn detection is the single largest naive cost, and it is pure waiting. A fixed 2-second silence timeout adds 2 seconds to every single turn before any computation begins. It is also the row people never look at, because it does not appear in any profiler — the process is idle.
"LLM remaining tokens: 0 ms" is not a trick. The first sentence of the reply is roughly 8 words, about 3.2 seconds of audio. Generating the remaining 32 tokens at 25 tokens per second takes about 1,280 ms — comfortably less than 3,200 ms of playback. The user hears continuous speech while the rest is still being produced. The work did not get faster; it got moved behind something the user was already waiting through.
Optimise time-to-first-audio, not total processing time. Everything that happens after the first syllable plays is free, provided it keeps up with playback.
The end-of-turn trade-off, honestly
Cutting the silence timeout from 2,000 ms to 500 ms is not free. Consider a user saying "my account number is... uh... four four seven one". The pause after "is" is easily 700 ms. A 500 ms timeout cuts them off mid-sentence, transcribes a fragment, and the assistant answers a question nobody asked.
| Approach | Added latency | Failure mode |
|---|---|---|
| Fixed 2,000 ms silence | 2,000 ms every turn | Feels broken; users repeat themselves |
| Fixed 500 ms silence | 500 ms every turn | Interrupts natural mid-sentence pauses |
| VAD + trailing-word heuristic | 500–900 ms | Language-specific; misses unusual phrasing |
| Semantic endpointing model | 500 ms + ~20 ms inference | Needs a model and its own failure analysis |
Semantic endpointing is the real answer: run a small classifier over the partial transcript that asks "is this a complete utterance?" and require both silence and a completeness signal. "My account number is" scores as incomplete, so you keep waiting; "my account number is four four seven one" scores complete, so you cut at 500 ms. If that is too much machinery for a first build, an adaptive timeout gets most of the benefit: 500 ms by default, extended to 1,200 ms when the transcript ends in a filler word, a conjunction, or a number that is still being read out.
Project layout
voice-assistant/ config.py settings, thresholds, model names audio.py microphone capture, VAD, playback queue asr.py speech to text llm.py streaming completion + sentence splitting speech.py text to speech, sentence-at-a-time session.py conversation state and history trimming main.py the turn loop, including barge-in Dockerfile1python -m venv .venv && source .venv/bin/activate2pip install faster-whisper sounddevice numpy webrtcvad-wheels setuptools openai coqui-tts soundfile3# webrtcvad-wheels: prebuilt webrtcvad; it still imports pkg_resources, hence setuptools4# coqui-tts: the maintained fork of Coqui TTS (same `from TTS.api import TTS`)5# ffmpeg and PortAudio are system dependencies6brew install ffmpeg portaudio # or: apt-get install ffmpeg portaudio19-devCapture and turn detection
The capture layer's job is to emit one complete user utterance and nothing else. Two rules govern it: never block inside the audio callback, and always copy the incoming buffer.
1# audio.py2import queue3import numpy as np4import sounddevice as sd5import webrtcvad67SR = 160008FRAME_MS = 209FRAME = SR * FRAME_MS // 1000 # 320 samples1011class Microphone:12 def __init__(self, aggressiveness=2):13 self.vad = webrtcvad.Vad(aggressiveness) # 0 permissive .. 3 strict14 self.q = queue.Queue()15 self.muted = False1617 def _cb(self, indata, frames, t, status):18 # Runs on the audio thread. Copy, enqueue, return. Nothing else.19 self.q.put(indata[:, 0].copy())2021 def utterances(self, silence_ms=500, min_speech_ms=300, pre_roll_ms=300):22 pre_roll = []23 pre_roll_frames = pre_roll_ms // FRAME_MS24 buf, silent_ms, speech_ms, active = [], 0, 0, False2526 with sd.InputStream(samplerate=SR, channels=1, dtype="int16",27 blocksize=FRAME, callback=self._cb):28 while True:29 frame = self.q.get()30 if self.muted: # do not listen to ourselves31 continue32 is_speech = self.vad.is_speech(frame.tobytes(), SR)3334 if not active:35 pre_roll.append(frame)36 if len(pre_roll) > pre_roll_frames:37 pre_roll.pop(0)38 if is_speech:39 active = True40 buf = pre_roll[:] + [frame] # keep the first consonant41 speech_ms, silent_ms = FRAME_MS, 042 continue4344 buf.append(frame)45 if is_speech:46 speech_ms += FRAME_MS47 silent_ms = 048 else:49 silent_ms += FRAME_MS5051 if silent_ms >= silence_ms:52 if speech_ms >= min_speech_ms: # ignore coughs and door slams53 audio = np.concatenate(buf).astype(np.float32) / 32768.054 yield audio55 buf, pre_roll, active = [], [], FalseThe pre-roll buffer is the detail that separates a working capture layer from a frustrating one. Voice activity detection needs a frame or two of speech before it fires, so by the time is_speech returns true you have already discarded the first 40–60 ms — which is exactly where a plosive consonant lives. Without pre-roll, "Play the news" is transcribed as "lay the news" often enough to be maddening, and the bug is invisible in the code because everything downstream is correct.
The min_speech_ms gate is the other one. Without it, a cough, a chair scrape or a door closing triggers a full turn: 300 ms of noise goes to the ASR model, which as a language model will confidently hallucinate a sentence, which goes to the LLM, which answers it. The assistant appears to be talking to itself.
Transcription
1# asr.py2from faster_whisper import WhisperModel34class Transcriber:5 def __init__(self, size="small", device="cuda"):6 compute = "float16" if device == "cuda" else "int8"7 self.model = WhisperModel(size, device=device, compute_type=compute)89 def transcribe(self, audio, language="en"):10 segments, _ = self.model.transcribe(11 audio,12 language=language, # skip detection: faster and more accurate13 beam_size=1, # greedy; beam search is not worth the ms here14 vad_filter=True, # second line of defence against hallucination15 condition_on_previous_text=False,16 )17 segs = list(segments)18 if not segs:19 return "", 0.020 text = " ".join(s.text.strip() for s in segs)21 confidence = sum(s.avg_logprob for s in segs) / len(segs)22 return text, confidenceThree choices here are latency decisions rather than accuracy decisions. beam_size=1 gives up a fraction of a WER point and saves roughly 40% of decode time. Passing language explicitly skips a detection pass over the first window. int8 on CPU roughly halves inference time for well under a point of WER.
Return the confidence and use it. A turn with avg_logprob below about −1.0 is usually noise or a hallucination, and the correct response is "sorry, I didn't catch that" — which users forgive instantly. Answering a hallucinated question, they do not forgive.
Streaming the language model into speech
This is the component that produces the biggest latency win, and it is mostly a sentence splitter.
1# llm.py2import os3import re4from openai import OpenAI56client = OpenAI()7MODEL = os.environ.get("LLM_MODEL", "gpt-6-luna") # a fast, cheap model; pin it in config8BOUNDARY = re.compile(r'(?<=[.!?])\s+')910SYSTEM = (11 "You are a voice assistant. Replies are spoken aloud, so keep them under "12 "three sentences. Never use lists, markdown, or symbols. Write numbers as "13 "words. Lead with the answer."14)1516def stream_sentences(history, min_chars=25):17 """Yield complete sentences as soon as the model produces them."""18 buf = ""19 stream = client.chat.completions.create(20 model=MODEL,21 messages=[{"role": "system", "content": SYSTEM}] + history,22 stream=True,23 max_completion_tokens=160,24 )25 for chunk in stream:26 if not chunk.choices: # some chunks carry only metadata27 continue28 delta = chunk.choices[0].delta.content or ""29 buf += delta30 while True:31 parts = BOUNDARY.split(buf, maxsplit=1)32 if len(parts) < 2 or len(parts[0]) < min_chars:33 break34 yield parts[0].strip()35 buf = parts[1]36 if buf.strip():37 yield buf.strip()The min_chars guard prevents a pathological case. "Sure. Let me check that for you." splits into "Sure." after five characters, which triggers a TTS call for a one-word clip. The per-call overhead then dominates, and the playback has an audible seam. Holding until 25 characters keeps the first chunk long enough to be worth synthesising.
The system prompt is doing real work too. A text chatbot answering "what are my options?" with a bulleted list is helpful; the same reply spoken aloud produces "dash upgrade your plan dash contact support". Voice output needs a prompt that bans list syntax, bans symbols, and caps length — because in text a user skims a long answer, and in speech they must sit through all of it.
Speaking, and being interruptible
1# speech.py2import queue, threading3import numpy as np4import sounddevice as sd5from TTS.api import TTS67SR = 2400089class Speaker:10 def __init__(self, speaker_wav=None, speaker="Ana Florence"):11 self.tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2").to("cuda")12 self.speaker_wav = speaker_wav # only a voice you have consent to clone13 self.speaker = None if speaker_wav else speaker # else a built-in XTTS voice14 self.q = queue.Queue()15 self.stop = threading.Event()16 threading.Thread(target=self._play_loop, daemon=True).start()1718 def _play_loop(self):19 while True:20 audio = self.q.get()21 if self.stop.is_set():22 continue # drain, do not play23 sd.play(audio, SR)24 while sd.get_stream().active:25 if self.stop.is_set():26 sd.stop() # barge-in: cut mid-word27 break28 sd.sleep(20)2930 def say(self, sentence):31 wav = self.tts.tts(text=sentence, speaker=self.speaker,32 speaker_wav=self.speaker_wav, language="en")33 self.q.put(np.asarray(wav, dtype=np.float32))3435 def interrupt(self):36 self.stop.set()37 with self.q.mutex:38 self.q.queue.clear()3940 def resume(self):41 self.stop.clear()Barge-in — letting the user cut the assistant off mid-sentence — is the feature that most separates a system that feels conversational from one that feels like a phone menu. It requires three things working together: continued microphone monitoring during playback, an interrupt path that stops the audio device rather than waiting for the current clip to finish, and cancellation of the in-flight LLM stream so tokens for an abandoned answer stop being generated and billed.
It also requires solving a problem that surprises people: the microphone hears the speaker. Without acoustic echo cancellation, the assistant's own voice triggers your VAD, gets transcribed, and gets sent to the LLM as if the user had said it. The assistant answers itself, forever. Headphones eliminate this. If you need open speakers, use a platform echo canceller (WebRTC's AEC in a browser, the OS-level one on mobile) or mute the microphone during playback — which is what the muted flag in the capture layer does, and which costs you barge-in entirely. That is the trade: half-duplex is trivial and feels robotic; full-duplex needs echo cancellation and feels human.
The turn loop
1# main.py2from audio import Microphone3from asr import Transcriber4from speech import Speaker5from llm import stream_sentences67MAX_TURNS = 12 # keep the prompt small; latency scales with it89def main():10 mic, asr, speaker = Microphone(), Transcriber(), Speaker()11 history = []1213 for audio in mic.utterances(silence_ms=500):14 text, confidence = asr.transcribe(audio)1516 if not text or confidence < -1.0:17 speaker.say("Sorry, I didn't catch that.")18 continue1920 print("user:", text)21 history.append({"role": "user", "content": text})22 history[:] = history[-MAX_TURNS:]2324 speaker.resume()25 reply = []26 for sentence in stream_sentences(history):27 print("assistant:", sentence)28 speaker.say(sentence) # queued; playback already started29 reply.append(sentence)3031 history.append({"role": "assistant", "content": " ".join(reply)})3233if __name__ == "__main__":34 main()The MAX_TURNS trim is a latency control, not a memory control. Prompt processing time grows with prompt length, so an unbounded history means turn 40 is measurably slower than turn 2, and the degradation is gradual enough that nobody attributes it to the right cause. Cap the history; if you need long-term memory, summarise older turns into a short standing note rather than carrying the transcript.
Testing each piece in isolation
End-to-end debugging of a voice pipeline is miserable, because a bad reply could originate in any of five stages. Test each stage against a fixed input and you localise faults in seconds.
1# Save a real utterance once, then use it as a fixture forever2import numpy as np3import soundfile as sf4from asr import Transcriber56def test_asr():7 audio, sr = sf.read("fixtures/what_is_the_weather.wav", dtype="float32")8 text, conf = Transcriber(size="small", device="cpu").transcribe(audio)9 assert "weather" in text.lower(), text10 assert conf > -1.0, conf1112def test_tts_is_intelligible():13 """Synthesise, transcribe, compare. Catches skipped and mangled words."""14 from speech import Speaker15 from asr import Transcriber16 sentence = "Your appointment is on Thursday the fourth at ten fifteen."17 wav = Speaker().tts.tts(text=sentence, speaker="Ana Florence", language="en")18 back, _ = Transcriber(device="cpu").transcribe(np.asarray(wav, dtype="float32"))19 assert "thursday" in back.lower() and "fifteen" in back.lower(), backThe second test is the highest-value one in the suite. Synthesise a sentence, transcribe it back, and assert the important words survived. It catches a number being read wrong, a name being mangled, or a word being skipped — and it runs in seconds on every commit, where a human listening test does not.
Symptoms and their causes
| Symptom | Cause | Fix |
|---|---|---|
| Assistant replies to things nobody said | Its own audio is reaching the microphone | Echo cancellation, or mute during playback |
| First word of every turn is clipped | No pre-roll buffer before VAD fires | Retain 300 ms of audio before speech onset |
| Random "Thank you for watching" | Whisper hallucinating over silence | VAD gate plus a minimum speech duration |
| Long pause before every reply | Fixed silence timeout too long | 500 ms VAD hangover with adaptive extension |
| Reply starts, then stutters | TTS is slower than playback consumes it | Faster vocoder, or buffer two sentences ahead |
| Gets slower over a long session | Conversation history growing unbounded | Cap turns; summarise older context |
| Users talk over the assistant | Replies too long for speech | Cap reply length in the system prompt |
| Numbers read wrong | No text normalisation before TTS | Normalise; add a pronunciation override table |
| Cuts users off mid-sentence | Endpointing on silence alone | Require silence and a completeness signal |
Almost every voice-assistant bug that looks like a model problem is a turn-boundary problem: the system decided the user had stopped speaking when they had not, or that they had started when they had not.
Deploying it
1# Dockerfile2FROM python:3.11-slim3RUN apt-get update && apt-get install -y --no-install-recommends \4 ffmpeg portaudio19-dev libsndfile1 && rm -rf /var/lib/apt/lists/*5WORKDIR /app6COPY requirements.txt .7RUN pip install --no-cache-dir -r requirements.txt8# Bake weights into the image so a cold start is not a model download9RUN python -c "from faster_whisper import WhisperModel; WhisperModel('small')"10COPY . .11CMD ["python", "main.py"]The weight-baking line matters more than it looks. Downloading a Whisper checkpoint on first request turns a cold start into a 40-second outage for whoever hit it. Bake the weights, and load models once at process start rather than per request — loading small takes several seconds and allocates around 2 GB, which is fine once and fatal per turn.
For anything beyond a local prototype, the audio should not travel over HTTP request-response. Use WebRTC or a WebSocket, so audio streams while the user is still speaking rather than being uploaded after they stop. That is the "audio to server: 150 ms → 50 ms" row in the budget table, and it is also what makes barge-in possible over a network.
There is also a different architecture to weigh: one speech-to-speech model that takes audio in and produces audio out, such as OpenAI's Realtime API (the gpt-realtime models) or Google's Gemini Live API. It removes the separate ASR and TTS hops, hears tone and hesitation that a transcript throws away, and can detect turn ends and handle barge-in on the server. The costs are the ones this project avoids: audio pricing that is usually higher than a cascade of small models, less control over each stage, and no text in the middle to log, filter or test unless you request a transcript. Build the cascade first, because it shows you where the time goes, then measure a speech-to-speech model against the same replay suite described below.
What to build next, and what it will cost
The extensions worth adding, in the order that adds the most value per unit of work:
| Extension | Effort | Why it matters |
|---|---|---|
| Tool calling (calendar, search, orders) | Medium | Turns a chatbot into an assistant; the main reason users return |
| Semantic endpointing | Medium | Removes the last big latency cost without cutting people off |
| Wake word | Low | Lets it run continuously without transcribing the room |
| Per-user voice selection | Low | Zero-shot embeddings; 2 KB per user, deletable on request |
| Multilingual turn handling | Medium | Detect language per turn; switch ASR and TTS together |
| Interruption-aware history | Low | Record what was actually heard, not what was generated |
That last one is subtle and worth doing early. When a user barges in after three words of a four-sentence reply, the model's history says it delivered all four sentences. It will then refer back to information the user never heard. Truncate the stored assistant turn at the point playback was interrupted.
On cost, run the arithmetic before you scale, because voice is priced per minute and per character rather than per request. Take a ten-turn conversation with 5 seconds of user speech and about 180 characters of reply per turn. Assume example hosted rates of 0.006 dollars per minute of transcription and 15 dollars per million characters of synthesis — real prices change often, so check your provider's pricing page before you plan on these. Transcription is 50 seconds, or 0.83 minutes, giving 0.005 dollars. Synthesis is 1,800 characters, giving 0.027 dollars. With a small language model contributing under a cent, the conversation costs roughly 0.04 dollars. At ten thousand conversations a day that is about 400 dollars a day, or 146,000 a year — at which point self-hosting an ASR model on a single GPU, which costs a small fraction of that, stops being a preference and becomes the obvious answer.
Finally, add the two things a prototype never has and a product cannot ship without. One is a transcript log per session, with timestamps for each stage, so you can answer "why was that turn slow" without reproducing it. The other is a fixed set of recorded utterances — twenty is enough — that you replay through the whole pipeline on every change, asserting both the transcript and the time to first audio. A model swap that improves accuracy by a point and adds 400 ms to every turn is a regression, and without that check you will ship it believing it was an upgrade.