Voice and Speech AI

Course Content

Mini-Project: Build a Voice-to-Voice Assistant


Four components, all of which you can get working in an afternoon. Capture microphone audio. Transcribe it. Send the text to a language model. Speak the answer. Wire them in a loop and you have a voice assistant.

Build it that way and then use it, and the experience is dreadful. You finish speaking. Nothing happens. You wonder whether it heard you. You start to repeat yourself — and then it starts talking, over you. You wait for it to finish so you can correct it, but it delivers all four sentences of an answer that stopped being relevant after the first one.

Nothing is broken. Every component is working exactly as designed. The measured round trip is 6.3 seconds, and that number is the entire problem.

Human conversation runs on a much tighter clock than people realise. Studies of turn-taking across ten unrelated languages found the median gap between one speaker finishing and the next beginning is around 200 milliseconds — less than the time it takes to plan a word, which means listeners are predicting the end of your sentence and preparing their reply before you get there. Push a response past about 800 ms and it registers as hesitation. Past two seconds and people assume the line has dropped and start talking again.

So the interesting engineering in a voice assistant is not any of the four components. It is the 6.1 seconds you have to remove from between them.

Where the 6.3 seconds actually goesTransport andplayback: 0.4 sFixed silencetimeout: 2.0 sWhispertranscription: 0.75 sLanguage model,full reply: 2.2 sTTS of thewhole reply: 0.96 stopbottomStreaming the reply into TTS sentence by sentence hides the last 1.6 seconds of generation behind playback.
The largest cost is waiting, not computing — shorten the end-of-turn wait and start speaking before the reply is finished.

Where the 6.3 seconds actually goes

Assume a realistic turn: the user speaks for 5 seconds, and the assistant replies with 40 tokens, roughly 30 words, which at a normal 150 words per minute is 12 seconds of speech.

StageNaiveOptimisedWhat changed
End-of-turn detection2,000 ms500 msVAD with 500 ms hangover, not a fixed 2 s timeout
Audio to server150 ms50 msStream during speech instead of after it
ASR750 ms180 msfaster-whisper int8; decode as audio arrives
LLM first token600 ms350 msSmaller model, cached system prompt
LLM remaining tokens1,600 ms0 msHidden behind playback of the first sentence
TTS960 ms120 msSynthesise first sentence only, then stream
Transport + playback buffer250 ms80 msOpus, small jitter buffer
Time to first audio6,310 ms1,280 ms4.9× faster

Check the two rows that matter most, because they are where the counter-intuitive wins live.

End-of-turn detection is the single largest naive cost, and it is pure waiting. A fixed 2-second silence timeout adds 2 seconds to every single turn before any computation begins. It is also the row people never look at, because it does not appear in any profiler — the process is idle.

"LLM remaining tokens: 0 ms" is not a trick. The first sentence of the reply is roughly 8 words, about 3.2 seconds of audio. Generating the remaining 32 tokens at 25 tokens per second takes about 1,280 ms — comfortably less than 3,200 ms of playback. The user hears continuous speech while the rest is still being produced. The work did not get faster; it got moved behind something the user was already waiting through.

Optimise time-to-first-audio, not total processing time. Everything that happens after the first syllable plays is free, provided it keeps up with playback.

The end-of-turn trade-off, honestly

Cutting the silence timeout from 2,000 ms to 500 ms is not free. Consider a user saying "my account number is... uh... four four seven one". The pause after "is" is easily 700 ms. A 500 ms timeout cuts them off mid-sentence, transcribes a fragment, and the assistant answers a question nobody asked.

ApproachAdded latencyFailure mode
Fixed 2,000 ms silence2,000 ms every turnFeels broken; users repeat themselves
Fixed 500 ms silence500 ms every turnInterrupts natural mid-sentence pauses
VAD + trailing-word heuristic500–900 msLanguage-specific; misses unusual phrasing
Semantic endpointing model500 ms + ~20 ms inferenceNeeds a model and its own failure analysis

Semantic endpointing is the real answer: run a small classifier over the partial transcript that asks "is this a complete utterance?" and require both silence and a completeness signal. "My account number is" scores as incomplete, so you keep waiting; "my account number is four four seven one" scores complete, so you cut at 500 ms. If that is too much machinery for a first build, an adaptive timeout gets most of the benefit: 500 ms by default, extended to 1,200 ms when the transcript ends in a filler word, a conjunction, or a number that is still being read out.

Project layout

Text
voice-assistant/  config.py       settings, thresholds, model names  audio.py        microphone capture, VAD, playback queue  asr.py          speech to text  llm.py          streaming completion + sentence splitting  speech.py       text to speech, sentence-at-a-time  session.py      conversation state and history trimming  main.py         the turn loop, including barge-in  Dockerfile
Bash
python -m venv .venv && source .venv/bin/activatepip install faster-whisper sounddevice numpy webrtcvad-wheels setuptools openai coqui-tts soundfile# webrtcvad-wheels: prebuilt webrtcvad; it still imports pkg_resources, hence setuptools# coqui-tts: the maintained fork of Coqui TTS (same `from TTS.api import TTS`)# ffmpeg and PortAudio are system dependenciesbrew install ffmpeg portaudio        # or: apt-get install ffmpeg portaudio19-dev

Capture and turn detection

The capture layer's job is to emit one complete user utterance and nothing else. Two rules govern it: never block inside the audio callback, and always copy the incoming buffer.

Python
# audio.pyimport queueimport numpy as npimport sounddevice as sdimport webrtcvadSR = 16000FRAME_MS = 20FRAME = SR * FRAME_MS // 1000        # 320 samplesclass Microphone:    def __init__(self, aggressiveness=2):        self.vad = webrtcvad.Vad(aggressiveness)   # 0 permissive .. 3 strict        self.q = queue.Queue()        self.muted = False    def _cb(self, indata, frames, t, status):        # Runs on the audio thread. Copy, enqueue, return. Nothing else.        self.q.put(indata[:, 0].copy())    def utterances(self, silence_ms=500, min_speech_ms=300, pre_roll_ms=300):        pre_roll = []        pre_roll_frames = pre_roll_ms // FRAME_MS        buf, silent_ms, speech_ms, active = [], 0, 0, False        with sd.InputStream(samplerate=SR, channels=1, dtype="int16",                            blocksize=FRAME, callback=self._cb):            while True:                frame = self.q.get()                if self.muted:                       # do not listen to ourselves                    continue                is_speech = self.vad.is_speech(frame.tobytes(), SR)                if not active:                    pre_roll.append(frame)                    if len(pre_roll) > pre_roll_frames:                        pre_roll.pop(0)                    if is_speech:                        active = True                        buf = pre_roll[:] + [frame]  # keep the first consonant                        speech_ms, silent_ms = FRAME_MS, 0                    continue                buf.append(frame)                if is_speech:                    speech_ms += FRAME_MS                    silent_ms = 0                else:                    silent_ms += FRAME_MS                if silent_ms >= silence_ms:                    if speech_ms >= min_speech_ms:   # ignore coughs and door slams                        audio = np.concatenate(buf).astype(np.float32) / 32768.0                        yield audio                    buf, pre_roll, active = [], [], False

The pre-roll buffer is the detail that separates a working capture layer from a frustrating one. Voice activity detection needs a frame or two of speech before it fires, so by the time is_speech returns true you have already discarded the first 40–60 ms — which is exactly where a plosive consonant lives. Without pre-roll, "Play the news" is transcribed as "lay the news" often enough to be maddening, and the bug is invisible in the code because everything downstream is correct.

The min_speech_ms gate is the other one. Without it, a cough, a chair scrape or a door closing triggers a full turn: 300 ms of noise goes to the ASR model, which as a language model will confidently hallucinate a sentence, which goes to the LLM, which answers it. The assistant appears to be talking to itself.

Transcription

Python
# asr.pyfrom faster_whisper import WhisperModelclass Transcriber:    def __init__(self, size="small", device="cuda"):        compute = "float16" if device == "cuda" else "int8"        self.model = WhisperModel(size, device=device, compute_type=compute)    def transcribe(self, audio, language="en"):        segments, _ = self.model.transcribe(            audio,            language=language,          # skip detection: faster and more accurate            beam_size=1,                # greedy; beam search is not worth the ms here            vad_filter=True,            # second line of defence against hallucination            condition_on_previous_text=False,        )        segs = list(segments)        if not segs:            return "", 0.0        text = " ".join(s.text.strip() for s in segs)        confidence = sum(s.avg_logprob for s in segs) / len(segs)        return text, confidence

Three choices here are latency decisions rather than accuracy decisions. beam_size=1 gives up a fraction of a WER point and saves roughly 40% of decode time. Passing language explicitly skips a detection pass over the first window. int8 on CPU roughly halves inference time for well under a point of WER.

Return the confidence and use it. A turn with avg_logprob below about −1.0 is usually noise or a hallucination, and the correct response is "sorry, I didn't catch that" — which users forgive instantly. Answering a hallucinated question, they do not forgive.

Streaming the language model into speech

This is the component that produces the biggest latency win, and it is mostly a sentence splitter.

Python
# llm.pyimport osimport refrom openai import OpenAIclient = OpenAI()MODEL = os.environ.get("LLM_MODEL", "gpt-6-luna")   # a fast, cheap model; pin it in configBOUNDARY = re.compile(r'(?<=[.!?])\s+')SYSTEM = (    "You are a voice assistant. Replies are spoken aloud, so keep them under "    "three sentences. Never use lists, markdown, or symbols. Write numbers as "    "words. Lead with the answer.")def stream_sentences(history, min_chars=25):    """Yield complete sentences as soon as the model produces them."""    buf = ""    stream = client.chat.completions.create(        model=MODEL,        messages=[{"role": "system", "content": SYSTEM}] + history,        stream=True,        max_completion_tokens=160,    )    for chunk in stream:        if not chunk.choices:                 # some chunks carry only metadata            continue        delta = chunk.choices[0].delta.content or ""        buf += delta        while True:            parts = BOUNDARY.split(buf, maxsplit=1)            if len(parts) < 2 or len(parts[0]) < min_chars:                break            yield parts[0].strip()            buf = parts[1]    if buf.strip():        yield buf.strip()

The min_chars guard prevents a pathological case. "Sure. Let me check that for you." splits into "Sure." after five characters, which triggers a TTS call for a one-word clip. The per-call overhead then dominates, and the playback has an audible seam. Holding until 25 characters keeps the first chunk long enough to be worth synthesising.

The system prompt is doing real work too. A text chatbot answering "what are my options?" with a bulleted list is helpful; the same reply spoken aloud produces "dash upgrade your plan dash contact support". Voice output needs a prompt that bans list syntax, bans symbols, and caps length — because in text a user skims a long answer, and in speech they must sit through all of it.

Speaking, and being interruptible

Python
# speech.pyimport queue, threadingimport numpy as npimport sounddevice as sdfrom TTS.api import TTSSR = 24000class Speaker:    def __init__(self, speaker_wav=None, speaker="Ana Florence"):        self.tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2").to("cuda")        self.speaker_wav = speaker_wav      # only a voice you have consent to clone        self.speaker = None if speaker_wav else speaker   # else a built-in XTTS voice        self.q = queue.Queue()        self.stop = threading.Event()        threading.Thread(target=self._play_loop, daemon=True).start()    def _play_loop(self):        while True:            audio = self.q.get()            if self.stop.is_set():                continue                       # drain, do not play            sd.play(audio, SR)            while sd.get_stream().active:                if self.stop.is_set():                    sd.stop()                  # barge-in: cut mid-word                    break                sd.sleep(20)    def say(self, sentence):        wav = self.tts.tts(text=sentence, speaker=self.speaker,                           speaker_wav=self.speaker_wav, language="en")        self.q.put(np.asarray(wav, dtype=np.float32))    def interrupt(self):        self.stop.set()        with self.q.mutex:            self.q.queue.clear()    def resume(self):        self.stop.clear()

Barge-in — letting the user cut the assistant off mid-sentence — is the feature that most separates a system that feels conversational from one that feels like a phone menu. It requires three things working together: continued microphone monitoring during playback, an interrupt path that stops the audio device rather than waiting for the current clip to finish, and cancellation of the in-flight LLM stream so tokens for an abandoned answer stop being generated and billed.

It also requires solving a problem that surprises people: the microphone hears the speaker. Without acoustic echo cancellation, the assistant's own voice triggers your VAD, gets transcribed, and gets sent to the LLM as if the user had said it. The assistant answers itself, forever. Headphones eliminate this. If you need open speakers, use a platform echo canceller (WebRTC's AEC in a browser, the OS-level one on mobile) or mute the microphone during playback — which is what the muted flag in the capture layer does, and which costs you barge-in entirely. That is the trade: half-duplex is trivial and feels robotic; full-duplex needs echo cancellation and feels human.

The turn loop

Python
# main.pyfrom audio import Microphonefrom asr import Transcriberfrom speech import Speakerfrom llm import stream_sentencesMAX_TURNS = 12          # keep the prompt small; latency scales with itdef main():    mic, asr, speaker = Microphone(), Transcriber(), Speaker()    history = []    for audio in mic.utterances(silence_ms=500):        text, confidence = asr.transcribe(audio)        if not text or confidence < -1.0:            speaker.say("Sorry, I didn't catch that.")            continue        print("user:", text)        history.append({"role": "user", "content": text})        history[:] = history[-MAX_TURNS:]        speaker.resume()        reply = []        for sentence in stream_sentences(history):            print("assistant:", sentence)            speaker.say(sentence)              # queued; playback already started            reply.append(sentence)        history.append({"role": "assistant", "content": " ".join(reply)})if __name__ == "__main__":    main()

The MAX_TURNS trim is a latency control, not a memory control. Prompt processing time grows with prompt length, so an unbounded history means turn 40 is measurably slower than turn 2, and the degradation is gradual enough that nobody attributes it to the right cause. Cap the history; if you need long-term memory, summarise older turns into a short standing note rather than carrying the transcript.

Testing each piece in isolation

End-to-end debugging of a voice pipeline is miserable, because a bad reply could originate in any of five stages. Test each stage against a fixed input and you localise faults in seconds.

Python
# Save a real utterance once, then use it as a fixture foreverimport numpy as npimport soundfile as sffrom asr import Transcriberdef test_asr():    audio, sr = sf.read("fixtures/what_is_the_weather.wav", dtype="float32")    text, conf = Transcriber(size="small", device="cpu").transcribe(audio)    assert "weather" in text.lower(), text    assert conf > -1.0, confdef test_tts_is_intelligible():    """Synthesise, transcribe, compare. Catches skipped and mangled words."""    from speech import Speaker    from asr import Transcriber    sentence = "Your appointment is on Thursday the fourth at ten fifteen."    wav = Speaker().tts.tts(text=sentence, speaker="Ana Florence", language="en")    back, _ = Transcriber(device="cpu").transcribe(np.asarray(wav, dtype="float32"))    assert "thursday" in back.lower() and "fifteen" in back.lower(), back

The second test is the highest-value one in the suite. Synthesise a sentence, transcribe it back, and assert the important words survived. It catches a number being read wrong, a name being mangled, or a word being skipped — and it runs in seconds on every commit, where a human listening test does not.

Symptoms and their causes

SymptomCauseFix
Assistant replies to things nobody saidIts own audio is reaching the microphoneEcho cancellation, or mute during playback
First word of every turn is clippedNo pre-roll buffer before VAD firesRetain 300 ms of audio before speech onset
Random "Thank you for watching"Whisper hallucinating over silenceVAD gate plus a minimum speech duration
Long pause before every replyFixed silence timeout too long500 ms VAD hangover with adaptive extension
Reply starts, then stuttersTTS is slower than playback consumes itFaster vocoder, or buffer two sentences ahead
Gets slower over a long sessionConversation history growing unboundedCap turns; summarise older context
Users talk over the assistantReplies too long for speechCap reply length in the system prompt
Numbers read wrongNo text normalisation before TTSNormalise; add a pronunciation override table
Cuts users off mid-sentenceEndpointing on silence aloneRequire silence and a completeness signal

Almost every voice-assistant bug that looks like a model problem is a turn-boundary problem: the system decided the user had stopped speaking when they had not, or that they had started when they had not.

Deploying it

Bash
# DockerfileFROM python:3.11-slimRUN apt-get update && apt-get install -y --no-install-recommends \      ffmpeg portaudio19-dev libsndfile1 && rm -rf /var/lib/apt/lists/*WORKDIR /appCOPY requirements.txt .RUN pip install --no-cache-dir -r requirements.txt# Bake weights into the image so a cold start is not a model downloadRUN python -c "from faster_whisper import WhisperModel; WhisperModel('small')"COPY . .CMD ["python", "main.py"]

The weight-baking line matters more than it looks. Downloading a Whisper checkpoint on first request turns a cold start into a 40-second outage for whoever hit it. Bake the weights, and load models once at process start rather than per request — loading small takes several seconds and allocates around 2 GB, which is fine once and fatal per turn.

For anything beyond a local prototype, the audio should not travel over HTTP request-response. Use WebRTC or a WebSocket, so audio streams while the user is still speaking rather than being uploaded after they stop. That is the "audio to server: 150 ms → 50 ms" row in the budget table, and it is also what makes barge-in possible over a network.

There is also a different architecture to weigh: one speech-to-speech model that takes audio in and produces audio out, such as OpenAI's Realtime API (the gpt-realtime models) or Google's Gemini Live API. It removes the separate ASR and TTS hops, hears tone and hesitation that a transcript throws away, and can detect turn ends and handle barge-in on the server. The costs are the ones this project avoids: audio pricing that is usually higher than a cascade of small models, less control over each stage, and no text in the middle to log, filter or test unless you request a transcript. Build the cascade first, because it shows you where the time goes, then measure a speech-to-speech model against the same replay suite described below.

What to build next, and what it will cost

The extensions worth adding, in the order that adds the most value per unit of work:

ExtensionEffortWhy it matters
Tool calling (calendar, search, orders)MediumTurns a chatbot into an assistant; the main reason users return
Semantic endpointingMediumRemoves the last big latency cost without cutting people off
Wake wordLowLets it run continuously without transcribing the room
Per-user voice selectionLowZero-shot embeddings; 2 KB per user, deletable on request
Multilingual turn handlingMediumDetect language per turn; switch ASR and TTS together
Interruption-aware historyLowRecord what was actually heard, not what was generated

That last one is subtle and worth doing early. When a user barges in after three words of a four-sentence reply, the model's history says it delivered all four sentences. It will then refer back to information the user never heard. Truncate the stored assistant turn at the point playback was interrupted.

On cost, run the arithmetic before you scale, because voice is priced per minute and per character rather than per request. Take a ten-turn conversation with 5 seconds of user speech and about 180 characters of reply per turn. Assume example hosted rates of 0.006 dollars per minute of transcription and 15 dollars per million characters of synthesis — real prices change often, so check your provider's pricing page before you plan on these. Transcription is 50 seconds, or 0.83 minutes, giving 0.005 dollars. Synthesis is 1,800 characters, giving 0.027 dollars. With a small language model contributing under a cent, the conversation costs roughly 0.04 dollars. At ten thousand conversations a day that is about 400 dollars a day, or 146,000 a year — at which point self-hosting an ASR model on a single GPU, which costs a small fraction of that, stops being a preference and becomes the obvious answer.

Finally, add the two things a prototype never has and a product cannot ship without. One is a transcript log per session, with timestamps for each stage, so you can answer "why was that turn slow" without reproducing it. The other is a fixed set of recorded utterances — twenty is enough — that you replay through the whole pipeline on every change, asserting both the transcript and the time to first audio. A model swap that improves accuracy by a point and adds 400 ms to every turn is a regression, and without that check you will ship it believing it was an upgrade.