Course Content
Capstone Project: Multimodal Assistant
1 sections · 6 lessons
Voice Integration
You wire up the microphone, speak clearly for six seconds — "what's the weather like in Boston today" — and print the transcript:
>>> result = model.transcribe("recording.wav")>>> repr(result["text"])' Thank you.'Not an error. Not an empty string. A polite, confident, entirely fabricated sentence. This is Whisper's most notorious behaviour: fed near-silence or noise, it does not return nothing, it returns whatever text most often accompanies silence in its training data — "Thank you.", "Thanks for watching!", a subtitle credit line.
There are two bugs behind that one symptom, and both are in your code, not the model's. Your capture loop opened the stream at 44,100 Hz but wrote a WAV header claiming 16,000 Hz, so the file plays back at roughly a third of the correct speed and Whisper hears a slow rumble. And your microphone gain was low enough that the loudest sample in the file was about 3,200 out of a possible 32,768 — around 10% of full scale, which is quiet enough to look like room tone.
Text and images sit still. Audio does not. A microphone does not hand you "the sentence the user said"; it hands you tens of thousands of integers per second, with no word boundaries, no punctuation, and no marker for when speech began or ended. This stage is about the two translations at the edges of your assistant: audio in, text out, and text in, audio out. Everything between them is the text pipeline you already have.
A voice assistant is a text assistant with a translator bolted onto each end. If your text loop already works, adding voice means writing two adapters — not rebuilding the reasoning.
Install the system libraries first
This is the first part of the project with dependencies pip cannot satisfy, because they are C libraries owned by the operating system.
1# macOS2brew install portaudio ffmpeg34# Debian / Ubuntu5sudo apt-get install -y portaudio19-dev ffmpeg67# verify BEFORE writing any code8ffmpeg -version9python -c "import pyaudio; print('pyaudio ok')"PortAudio is what pyaudio binds to for microphone capture; ffmpeg is what Whisper shells out to when decoding anything that is not a plain WAV. Skip either and you will get a build failure during pip install or, worse, a FileNotFoundError from deep inside a library at runtime. On macOS you also need to grant your terminal microphone permission the first time — deny it and recording succeeds while producing pure silence, which lands you straight back at "Thank you."
What a microphone actually gives you
Digital audio is a stream of amplitude measurements. Three numbers define it: the sample rate (measurements per second), the sample width (bytes per measurement), and the channel count. This project uses 16,000 Hz, 16-bit signed integers, mono.
The data rate follows directly: 16,000×2×1=32,000 bytes per second. A five-second utterance is 160 kB. Compare that with CD-quality stereo at 44,100 Hz: 44,100×2×2=176,400 bytes per second, 5.5 times more data — for zero benefit, because Whisper resamples everything to 16 kHz internally before it does anything else. You would be paying to move bytes that get thrown away.
Why 16,000 specifically
The Nyquist–Shannon sampling theorem says a sample rate of f can faithfully represent frequencies up to f/2. At 16,000 Hz that ceiling is 8,000 Hz. Human speech fits underneath it: the vocal fold fundamental sits between roughly 85 Hz (a low male voice) and 255 Hz (a high female voice), the formants that distinguish one vowel from another live between about 300 Hz and 3,500 Hz, and the highest-frequency content that matters — the hiss of s and f — tops out around 8,000 Hz. Sample at 8,000 Hz instead and your ceiling drops to 4,000 Hz, taking the fricatives with it: "sip" and "ship" become genuinely ambiguous. Sample at 44,100 Hz and you capture cymbals and birdsong that no speech model uses.
Whisper, and what "base" costs you
Whisper ships in six sizes. The trade is accuracy against speed and memory, and there is no universally right choice — there is a right choice for your latency budget.
| Model | Parameters | Approx. VRAM | Relative speed | Use it when |
|---|---|---|---|---|
tiny | 39 M | ~1 GB | ~10× | Latency matters more than accuracy; clear English, quiet room |
base | 74 M | ~1 GB | ~7× | Default for this project — usable on a CPU laptop |
small | 244 M | ~2 GB | ~4× | Accents or mild background noise, and you have a GPU |
medium | 769 M | ~5 GB | ~2× | Multilingual work, batch transcription |
large | 1550 M | ~10 GB | 1× | Offline transcription where quality is the only concern |
turbo | 809 M | ~6 GB | ~8× | Near-large accuracy at much higher speed on a GPU; transcription only, no translation |
Relative speeds are the Whisper project's own figures, measured on an A100 GPU against large; on a laptop CPU every size is much slower, but the ordering holds.
One implementation detail changes how you design around Whisper: it processes audio in fixed 30-second windows, padding shorter clips with silence. A 2-second clip therefore costs almost exactly the same compute as a 28-second one. Chunking a long recording into many 2-second pieces is the worst possible strategy — fifteen 2-second chunks cost fifteen full windows, while one 30-second recording costs one.
1# src/voice_handler.py2import wave3import tempfile4from pathlib import Path56import numpy as np7import whisper89from config.settings import settings10from src.exceptions import VoiceError11from src.utils import logger1213RATE = 16_000 # Whisper's native rate; anything else gets resampled14CHANNELS = 115SAMPLE_WIDTH = 2 # bytes, i.e. 16-bit signed PCM16CHUNK = 1024 # frames per read from the device171819class WhisperTranscriber:20 """Loads the model ONCE. Reuse the instance; never build one per request."""2122 def __init__(self, model_size: str | None = None):23 size = model_size or settings.whisper_model24 logger.info("Loading Whisper %s ...", size)25 self.model = whisper.load_model(size)26 logger.info("Whisper %s ready", size)2728 def transcribe_file(self, path: str | Path) -> str:29 try:30 result = self.model.transcribe(31 str(path),32 language="en", # skip language detection: saves ~0.3 s33 fp16=False, # CPU cannot do fp16; silences a warning34 condition_on_previous_text=False, # stops loop-y repetitions35 )36 except Exception as e:37 raise VoiceError(f"transcription failed for {path}: {e}") from e3839 text = result["text"].strip()40 logger.info("Transcribed %d chars from %s", len(text), path)41 return text4243 def transcribe_bytes(self, pcm: bytes) -> str:44 """Whisper wants a path or an array, not raw bytes -- so write a WAV."""45 tmp_path = None46 try:47 with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tmp:48 tmp_path = tmp.name49 with wave.open(tmp_path, "wb") as wf:50 wf.setnchannels(CHANNELS)51 wf.setsampwidth(SAMPLE_WIDTH)52 wf.setframerate(RATE) # MUST match how it was captured53 wf.writeframes(pcm)54 return self.transcribe_file(tmp_path)55 finally:56 if tmp_path:57 Path(tmp_path).unlink(missing_ok=True)Three arguments there are doing real work. language="en" skips the detection pass, which is a separate forward pass costing roughly 0.3 seconds. fp16=False avoids a half-precision path CPUs cannot execute. And condition_on_previous_text=False stops Whisper feeding each window's output back as context for the next — the mechanism behind the failure where a model repeats one phrase twenty times after a single mis-transcription.
Capturing from the microphone
The bug from the opening was a mismatch between capture settings and file header. Define the constants once and use those same constants in both places, so drift is impossible.
1import pyaudio234class MicrophoneRecorder:5 def __init__(self):6 self.audio = pyaudio.PyAudio()78 def record(self, seconds: float = 5.0) -> bytes:9 stream = self.audio.open(10 format=pyaudio.paInt16, # matches SAMPLE_WIDTH = 211 channels=CHANNELS,12 rate=RATE, # matches the WAV header above13 input=True,14 frames_per_buffer=CHUNK,15 )16 logger.info("Recording %.1f s ...", seconds)17 frames = []18 try:19 for _ in range(int(RATE / CHUNK * seconds)):20 frames.append(stream.read(CHUNK, exception_on_overflow=False))21 finally:22 stream.stop_stream()23 stream.close()2425 pcm = b"".join(frames)26 peak = int(np.abs(np.frombuffer(pcm, dtype=np.int16)).max())27 logger.debug("Captured %d bytes, peak %d/32768 (%.1f%% FS)",28 len(pcm), peak, 100 * peak / 32768)29 if peak < 500:30 raise VoiceError(31 f"recording is essentially silent (peak {peak}/32768) -- "32 "check microphone permission, device selection, and gain"33 )34 return pcm3536 def close(self):37 self.audio.terminate()That peak check is the single highest-value line in the file. It converts the worst class of audio bug — a silent recording that produces a plausible hallucinated transcript — into an immediate, named error that tells you exactly what to look at. Without it you debug the wrong layer for an hour, convinced the model is broken.
Preprocessing: gain and silence
Normalisation, with the arithmetic
Take the opening failure: a recording whose loudest sample is 3,200 on a 16-bit scale that runs to 32,768. In decibels relative to full scale that is
You do not want to scale it all the way to full scale, because any louder moment later would clip. Target −3 dBFS instead, which is an amplitude ratio of 10−3/20=0.708, so the peak should land at 0.708×32768=23,200. The gain you need is 23,200/3,200=7.25.
1def normalise(pcm: bytes, target_dbfs: float = -3.0) -> bytes:2 samples = np.frombuffer(pcm, dtype=np.int16).astype(np.float32)3 peak = np.abs(samples).max()4 if peak < 1:5 return pcm # pure digital silence, nothing to scale6 target_peak = (10 ** (target_dbfs / 20)) * 327677 gain = target_peak / peak8 logger.debug("Normalising: peak %.0f, gain %.2fx", peak, gain)9 scaled = np.clip(samples * gain, -32768, 32767).astype(np.int16)10 return scaled.tobytes()Trimming silence
A user who presses record, thinks for three seconds, speaks for three, and hesitates for one has produced an 8-second file that is 4 seconds of nothing — 50% waste. Trimming does not speed Whisper up, because of the 30-second window, but it removes the silence that triggers hallucinated filler text, and it stops a long pause being mistaken for the end of the utterance.
1def trim_silence(pcm: bytes, threshold: int = 500,2 pad_ms: int = 100) -> bytes:3 samples = np.frombuffer(pcm, dtype=np.int16)4 loud = np.where(np.abs(samples) > threshold)[0]5 if loud.size == 0:6 return b"" # nothing but silence: return nothing7 pad = int(RATE * pad_ms / 1000) # keep 100 ms either side8 start = max(0, loud[0] - pad)9 end = min(samples.size, loud[-1] + pad)10 logger.debug("Trimmed %.2f s -> %.2f s",11 samples.size / RATE, (end - start) / RATE)12 return samples[start:end].tobytes()Returning b"" rather than the original when everything is below threshold is deliberate: an empty result is something the caller can check, whereas passing silence to Whisper produces confident nonsense.
Speaking back
Text-to-speech splits into two philosophies, and picking the wrong one for your situation is a common early mistake.
pyttsx3 (offline) | gTTS (free API) | Commercial API | |
|---|---|---|---|
| Needs network | No | Yes | Yes |
| Typical synthesis latency | 0.1–0.3 s | 0.5–1.0 s | 0.4–1.5 s |
| Voice quality | Clearly synthetic | Acceptable | Near-human |
| Cost | Free | Free, unofficial, may break | Per character |
| Works in CI / a container | Needs a system speech engine | Yes | Yes |
| Good for | Development, offline demos, tests | Prototypes | Anything a user will actually listen to |
Run the numbers before committing to a paid voice. An average assistant reply is around 400 characters. At 100 replies a day that is 40,000 characters daily and roughly 1.2 million a month — comfortably past the free tier of most commercial providers, and enough that per-character pricing becomes a line item you should have modelled before you built on it. The design answer is the same one that made the vision boundary useful: put TTS behind an interface with two implementations, develop against the free one, and switch by changing a setting.
1from abc import ABC, abstractmethod234class TextToSpeech(ABC):5 @abstractmethod6 def speak(self, text: str) -> None: ...789class OfflineTTS(TextToSpeech):10 def __init__(self, rate: int = 175):11 import pyttsx312 self.engine = pyttsx3.init()13 self.engine.setProperty("rate", rate) # words per minute1415 def speak(self, text: str) -> None:16 if not text.strip():17 return18 try:19 self.engine.say(text)20 self.engine.runAndWait()21 except Exception as e:22 raise VoiceError(f"speech synthesis failed: {e}") from e232425class NullTTS(TextToSpeech):26 """Used in tests and headless containers. Records what would be said."""27 def __init__(self): self.spoken: list[str] = []28 def speak(self, text: str) -> None: self.spoken.append(text)NullTTS is not filler. Without it, every test that touches the voice path needs a working audio device, which means your suite fails in CI and inside Docker for reasons that have nothing to do with your code.
The loop, and where the seconds go
1class VoiceHandler:2 def __init__(self, tts: TextToSpeech | None = None):3 self.transcriber = WhisperTranscriber()4 self.recorder = MicrophoneRecorder()5 self.tts = tts or OfflineTTS()67 def listen(self, seconds: float = 5.0) -> str:8 pcm = self.recorder.record(seconds)9 pcm = trim_silence(normalise(pcm))10 if not pcm:11 raise VoiceError("no speech detected in the recording")12 return self.transcriber.transcribe_bytes(pcm)1314 def respond(self, text: str) -> None:15 self.tts.speak(text)Now measure the full turn honestly. Times below are a mid-range laptop CPU with Whisper base:
| Stage | Time | Counts toward perceived delay? |
|---|---|---|
| User speaks | 4.0 s | No — the user is busy |
| Fixed-duration recording runs out | 1.0 s | Yes — pure dead air |
| Normalise and trim | 0.02 s | Negligible |
Whisper base on 30 s window | 1.3 s | Yes |
| LLM generates a reply | 2.1 s | Yes |
| TTS synthesis | 0.3 s | Yes |
| Playback begins | 0.1 s | Yes |
Adding the stages the user actually waits through: 1.0+0.02+1.3+2.1+0.3+0.1=4.82 seconds of silence after they stop talking. Human conversational turn-taking runs on gaps of about 200 milliseconds, so nearly five seconds feels broken, and users start repeating themselves — which produces overlapping audio and makes it worse.
The table tells you where to spend effort. The LLM call is the largest block at 2.1 s, and streaming it so synthesis starts on the first complete sentence recovers around 1.5 s. Replacing the fixed 5-second recording with silence detection that stops 0.4 s after the user does recovers most of the 1.0 s. Dropping to tiny saves about 0.8 s at a real accuracy cost. Doing all three lands just under 2 s, which is tolerable. Guessing instead of measuring, and people optimise the 0.02 s preprocessing step.
Latency in a voice pipeline is additive and dominated by one or two stages. Measure each stage separately before optimising anything, or you will spend a day making the fast part faster.
Tests without a microphone
Almost everything here is testable offline if you generate audio instead of recording it.
1# tests/test_voice.py2import numpy as np3import pytest4from src.voice_handler import normalise, trim_silence, NullTTS, RATE567def sine(seconds: float, freq: int = 440, amplitude: int = 3200) -> bytes:8 t = np.linspace(0, seconds, int(RATE * seconds), endpoint=False)9 return (amplitude * np.sin(2 * np.pi * freq * t)).astype(np.int16).tobytes()101112def silence(seconds: float) -> bytes:13 return np.zeros(int(RATE * seconds), dtype=np.int16).tobytes()141516def test_normalise_reaches_target_peak():17 out = np.frombuffer(normalise(sine(1.0, amplitude=3200)), dtype=np.int16)18 peak = np.abs(out).max()19 assert 22_000 < peak < 24_000 # -3 dBFS is 23,200 of 32,767202122def test_normalise_does_not_clip():23 out = np.frombuffer(normalise(sine(1.0, amplitude=32000)), dtype=np.int16)24 assert np.abs(out).max() <= 32_767252627def test_trim_removes_leading_and_trailing_silence():28 pcm = silence(2.0) + sine(1.0) + silence(1.5)29 trimmed = trim_silence(pcm)30 seconds = len(trimmed) / 2 / RATE31 assert 1.1 < seconds < 1.4 # 1 s of tone plus 100 ms padding3233def test_pure_silence_trims_to_nothing():34 assert trim_silence(silence(3.0)) == b""353637def test_null_tts_records_without_a_device():38 tts = NullTTS()39 tts.speak("hello")40 assert tts.spoken == ["hello"]Do not assert on exact transcript text. Whisper's output varies with version and hardware, including leading whitespace and punctuation. If you must test transcription end to end, commit one short WAV of clear speech and assert that the lower-cased result contains an expected keyword.
When things go wrong here
| Symptom | Cause | Fix |
|---|---|---|
| Transcript is "Thank you." or "Thanks for watching" | Near-silent or noise-only input; Whisper hallucinates on silence | The peak check in record(); trim_silence returning b""; verify mic permission |
| Transcript is garbled or in the wrong language | Capture rate and WAV header rate disagree | One RATE constant used by both the stream and setframerate |
| One phrase repeated many times | condition_on_previous_text feeding a bad window forward | Set it to False |
OSError: [Errno -9981] Input overflowed | Read loop cannot keep up with the device | exception_on_overflow=False, and raise frames_per_buffer to 2048 |
FileNotFoundError: [Errno 2] ... 'ffmpeg' | ffmpeg missing from PATH | Install it; restart the shell so PATH is refreshed |
pip install pyaudio fails compiling | PortAudio headers absent | brew install portaudio or apt-get install portaudio19-dev, then reinstall |
| Recording is silent on macOS only | Terminal or IDE lacks microphone permission | System Settings, Privacy, Microphone; restart the app after granting |
| First transcription takes 40 s, later ones 1.3 s | Model weights downloading and loading | Construct WhisperTranscriber once at startup; pre-warm on a short clip |
| TTS silent in Docker or CI | No audio device in the container | Inject NullTTS; return audio bytes to the client instead of playing them |
| Speech is cut off mid-sentence | Fixed recording duration expired while the user was still talking | Silence-based endpointing, or raise the duration and rely on trimming |
Acceptance criteria for this stage
ffmpeg -versionandimport pyaudioboth succeed before any voice code runs.- A 5-second recording of clear speech produces a transcript containing the expected keywords, with no "Thank you." artefact.
- Recording with the microphone muted raises
VoiceErrornaming the peak level — it never returns a fabricated transcript. normalise()on a signal peaking at 3,200 produces a peak between 22,000 and 24,000, and on one peaking at 32,000 produces no sample beyond ±32,767.trim_silence()on 2 s silence + 1 s tone + 1.5 s silence returns between 1.1 s and 1.4 s of audio; on pure silence it returnsb"".- The full voice test suite passes on a machine with no microphone and no speakers.
- Timing instrumentation logs each stage separately, and the sum of stages after the user stops speaking is under 5 seconds on your hardware.
- Whisper is loaded exactly once per process — grep the logs for "Whisper ... ready" and confirm a single occurrence per run.
What this buys you for the rest of the build
The whole stage collapses into two method calls: listen() returns a string, respond() consumes a string. Nothing beyond voice_handler.py knows that PortAudio, sample rates, or 30-second windows exist. That is what lets the same orchestrator serve a terminal user typing, a voice user speaking, and an HTTP client posting a WAV file, with the routing logic written once.
The interface choice around TTS pays off twice more. The abstract base class plus NullTTS is what makes the assistant runnable inside a container that has no sound card — you swap in an implementation that returns audio bytes over the wire instead of playing them locally, and every other line of code is unchanged. And it is what keeps the test suite honest: tests that need real hardware get skipped, then rot, then stop protecting you.
The habit worth carrying forward is the peak check. Audio failures are uniquely dangerous because the failure mode is not an exception — it is a plausible wrong answer. Every stage that can silently produce garbage deserves a cheap assertion at its boundary that turns "wrong" into "loudly broken". In this pipeline that assertion costs one line and saves an afternoon every time it fires.