Capstone Project: Multimodal Assistant

Voice Integration


You wire up the microphone, speak clearly for six seconds — "what's the weather like in Boston today" — and print the transcript:

Text
>>> result = model.transcribe("recording.wav")>>> repr(result["text"])' Thank you.'

Not an error. Not an empty string. A polite, confident, entirely fabricated sentence. This is Whisper's most notorious behaviour: fed near-silence or noise, it does not return nothing, it returns whatever text most often accompanies silence in its training data — "Thank you.", "Thanks for watching!", a subtitle credit line.

There are two bugs behind that one symptom, and both are in your code, not the model's. Your capture loop opened the stream at 44,100 Hz but wrote a WAV header claiming 16,000 Hz, so the file plays back at roughly a third of the correct speed and Whisper hears a slow rumble. And your microphone gain was low enough that the loudest sample in the file was about 3,200 out of a possible 32,768 — around 10% of full scale, which is quiet enough to look like room tone.

Text and images sit still. Audio does not. A microphone does not hand you "the sentence the user said"; it hands you tens of thousands of integers per second, with no word boundaries, no punctuation, and no marker for when speech began or ended. This stage is about the two translations at the edges of your assistant: audio in, text out, and text in, audio out. Everything between them is the text pipeline you already have.

A voice assistant is a text assistant with a translator bolted onto each end. If your text loop already works, adding voice means writing two adapters — not rebuilding the reasoning.

A six-second capture, before Whisper sees it00.010.42-0.380.55-0.510.02001234567trimstarts herepeak 0.55trailingsilence16,000 samples per second, mono, 16-bit; gain normalises the peak, then leading and trailing silence is cut.
Whisper transcribes a fixed 30-second window, so silence you fail to trim is padding the model still has to read.

Install the system libraries first

This is the first part of the project with dependencies pip cannot satisfy, because they are C libraries owned by the operating system.

Bash
# macOSbrew install portaudio ffmpeg# Debian / Ubuntusudo apt-get install -y portaudio19-dev ffmpeg# verify BEFORE writing any codeffmpeg -versionpython -c "import pyaudio; print('pyaudio ok')"

PortAudio is what pyaudio binds to for microphone capture; ffmpeg is what Whisper shells out to when decoding anything that is not a plain WAV. Skip either and you will get a build failure during pip install or, worse, a FileNotFoundError from deep inside a library at runtime. On macOS you also need to grant your terminal microphone permission the first time — deny it and recording succeeds while producing pure silence, which lands you straight back at "Thank you."

What a microphone actually gives you

Digital audio is a stream of amplitude measurements. Three numbers define it: the sample rate (measurements per second), the sample width (bytes per measurement), and the channel count. This project uses 16,000 Hz, 16-bit signed integers, mono.

The data rate follows directly: 16,000×2×1=32,00016{,}000 \times 2 \times 1 = 32{,}000 bytes per second. A five-second utterance is 160 kB. Compare that with CD-quality stereo at 44,100 Hz: 44,100×2×2=176,40044{,}100 \times 2 \times 2 = 176{,}400 bytes per second, 5.5 times more data — for zero benefit, because Whisper resamples everything to 16 kHz internally before it does anything else. You would be paying to move bytes that get thrown away.

Why 16,000 specifically

The Nyquist–Shannon sampling theorem says a sample rate of ff can faithfully represent frequencies up to f/2f/2. At 16,000 Hz that ceiling is 8,000 Hz. Human speech fits underneath it: the vocal fold fundamental sits between roughly 85 Hz (a low male voice) and 255 Hz (a high female voice), the formants that distinguish one vowel from another live between about 300 Hz and 3,500 Hz, and the highest-frequency content that matters — the hiss of s and f — tops out around 8,000 Hz. Sample at 8,000 Hz instead and your ceiling drops to 4,000 Hz, taking the fricatives with it: "sip" and "ship" become genuinely ambiguous. Sample at 44,100 Hz and you capture cymbals and birdsong that no speech model uses.

Whisper, and what "base" costs you

Whisper ships in six sizes. The trade is accuracy against speed and memory, and there is no universally right choice — there is a right choice for your latency budget.

ModelParametersApprox. VRAMRelative speedUse it when
tiny39 M~1 GB~10×Latency matters more than accuracy; clear English, quiet room
base74 M~1 GB~7×Default for this project — usable on a CPU laptop
small244 M~2 GB~4×Accents or mild background noise, and you have a GPU
medium769 M~5 GB~2×Multilingual work, batch transcription
large1550 M~10 GB1×Offline transcription where quality is the only concern
turbo809 M~6 GB~8×Near-large accuracy at much higher speed on a GPU; transcription only, no translation

Relative speeds are the Whisper project's own figures, measured on an A100 GPU against large; on a laptop CPU every size is much slower, but the ordering holds.

One implementation detail changes how you design around Whisper: it processes audio in fixed 30-second windows, padding shorter clips with silence. A 2-second clip therefore costs almost exactly the same compute as a 28-second one. Chunking a long recording into many 2-second pieces is the worst possible strategy — fifteen 2-second chunks cost fifteen full windows, while one 30-second recording costs one.

Python
# src/voice_handler.pyimport waveimport tempfilefrom pathlib import Pathimport numpy as npimport whisperfrom config.settings import settingsfrom src.exceptions import VoiceErrorfrom src.utils import loggerRATE = 16_000        # Whisper's native rate; anything else gets resampledCHANNELS = 1SAMPLE_WIDTH = 2     # bytes, i.e. 16-bit signed PCMCHUNK = 1024         # frames per read from the deviceclass WhisperTranscriber:    """Loads the model ONCE. Reuse the instance; never build one per request."""    def __init__(self, model_size: str | None = None):        size = model_size or settings.whisper_model        logger.info("Loading Whisper %s ...", size)        self.model = whisper.load_model(size)        logger.info("Whisper %s ready", size)    def transcribe_file(self, path: str | Path) -> str:        try:            result = self.model.transcribe(                str(path),                language="en",          # skip language detection: saves ~0.3 s                fp16=False,             # CPU cannot do fp16; silences a warning                condition_on_previous_text=False,  # stops loop-y repetitions            )        except Exception as e:            raise VoiceError(f"transcription failed for {path}: {e}") from e        text = result["text"].strip()        logger.info("Transcribed %d chars from %s", len(text), path)        return text    def transcribe_bytes(self, pcm: bytes) -> str:        """Whisper wants a path or an array, not raw bytes -- so write a WAV."""        tmp_path = None        try:            with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tmp:                tmp_path = tmp.name            with wave.open(tmp_path, "wb") as wf:                wf.setnchannels(CHANNELS)                wf.setsampwidth(SAMPLE_WIDTH)                wf.setframerate(RATE)      # MUST match how it was captured                wf.writeframes(pcm)            return self.transcribe_file(tmp_path)        finally:            if tmp_path:                Path(tmp_path).unlink(missing_ok=True)

Three arguments there are doing real work. language="en" skips the detection pass, which is a separate forward pass costing roughly 0.3 seconds. fp16=False avoids a half-precision path CPUs cannot execute. And condition_on_previous_text=False stops Whisper feeding each window's output back as context for the next — the mechanism behind the failure where a model repeats one phrase twenty times after a single mis-transcription.

Capturing from the microphone

The bug from the opening was a mismatch between capture settings and file header. Define the constants once and use those same constants in both places, so drift is impossible.

Python
import pyaudioclass MicrophoneRecorder:    def __init__(self):        self.audio = pyaudio.PyAudio()    def record(self, seconds: float = 5.0) -> bytes:        stream = self.audio.open(            format=pyaudio.paInt16,   # matches SAMPLE_WIDTH = 2            channels=CHANNELS,            rate=RATE,                # matches the WAV header above            input=True,            frames_per_buffer=CHUNK,        )        logger.info("Recording %.1f s ...", seconds)        frames = []        try:            for _ in range(int(RATE / CHUNK * seconds)):                frames.append(stream.read(CHUNK, exception_on_overflow=False))        finally:            stream.stop_stream()            stream.close()        pcm = b"".join(frames)        peak = int(np.abs(np.frombuffer(pcm, dtype=np.int16)).max())        logger.debug("Captured %d bytes, peak %d/32768 (%.1f%% FS)",                     len(pcm), peak, 100 * peak / 32768)        if peak < 500:            raise VoiceError(                f"recording is essentially silent (peak {peak}/32768) -- "                "check microphone permission, device selection, and gain"            )        return pcm    def close(self):        self.audio.terminate()

That peak check is the single highest-value line in the file. It converts the worst class of audio bug — a silent recording that produces a plausible hallucinated transcript — into an immediate, named error that tells you exactly what to look at. Without it you debug the wrong layer for an hour, convinced the model is broken.

Preprocessing: gain and silence

Normalisation, with the arithmetic

Take the opening failure: a recording whose loudest sample is 3,200 on a 16-bit scale that runs to 32,768. In decibels relative to full scale that is

20log⁡10 ⁣(320032768)=20log⁡10(0.0977)=−20.2 dBFS20 \log_{10}\!\left(\frac{3200}{32768}\right) = 20 \log_{10}(0.0977) = -20.2 \text{ dBFS}

You do not want to scale it all the way to full scale, because any louder moment later would clip. Target −3 dBFS instead, which is an amplitude ratio of 10−3/20=0.70810^{-3/20} = 0.708, so the peak should land at 0.708×32768=23,2000.708 \times 32768 = 23{,}200. The gain you need is 23,200/3,200=7.2523{,}200 / 3{,}200 = 7.25.

Python
def normalise(pcm: bytes, target_dbfs: float = -3.0) -> bytes:    samples = np.frombuffer(pcm, dtype=np.int16).astype(np.float32)    peak = np.abs(samples).max()    if peak < 1:        return pcm                       # pure digital silence, nothing to scale    target_peak = (10 ** (target_dbfs / 20)) * 32767    gain = target_peak / peak    logger.debug("Normalising: peak %.0f, gain %.2fx", peak, gain)    scaled = np.clip(samples * gain, -32768, 32767).astype(np.int16)    return scaled.tobytes()

Trimming silence

A user who presses record, thinks for three seconds, speaks for three, and hesitates for one has produced an 8-second file that is 4 seconds of nothing — 50% waste. Trimming does not speed Whisper up, because of the 30-second window, but it removes the silence that triggers hallucinated filler text, and it stops a long pause being mistaken for the end of the utterance.

Python
def trim_silence(pcm: bytes, threshold: int = 500,                 pad_ms: int = 100) -> bytes:    samples = np.frombuffer(pcm, dtype=np.int16)    loud = np.where(np.abs(samples) > threshold)[0]    if loud.size == 0:        return b""                       # nothing but silence: return nothing    pad = int(RATE * pad_ms / 1000)      # keep 100 ms either side    start = max(0, loud[0] - pad)    end = min(samples.size, loud[-1] + pad)    logger.debug("Trimmed %.2f s -> %.2f s",                 samples.size / RATE, (end - start) / RATE)    return samples[start:end].tobytes()

Returning b"" rather than the original when everything is below threshold is deliberate: an empty result is something the caller can check, whereas passing silence to Whisper produces confident nonsense.

Speaking back

Text-to-speech splits into two philosophies, and picking the wrong one for your situation is a common early mistake.

pyttsx3 (offline)gTTS (free API)Commercial API
Needs networkNoYesYes
Typical synthesis latency0.1–0.3 s0.5–1.0 s0.4–1.5 s
Voice qualityClearly syntheticAcceptableNear-human
CostFreeFree, unofficial, may breakPer character
Works in CI / a containerNeeds a system speech engineYesYes
Good forDevelopment, offline demos, testsPrototypesAnything a user will actually listen to

Run the numbers before committing to a paid voice. An average assistant reply is around 400 characters. At 100 replies a day that is 40,000 characters daily and roughly 1.2 million a month — comfortably past the free tier of most commercial providers, and enough that per-character pricing becomes a line item you should have modelled before you built on it. The design answer is the same one that made the vision boundary useful: put TTS behind an interface with two implementations, develop against the free one, and switch by changing a setting.

Python
from abc import ABC, abstractmethodclass TextToSpeech(ABC):    @abstractmethod    def speak(self, text: str) -> None: ...class OfflineTTS(TextToSpeech):    def __init__(self, rate: int = 175):        import pyttsx3        self.engine = pyttsx3.init()        self.engine.setProperty("rate", rate)   # words per minute    def speak(self, text: str) -> None:        if not text.strip():            return        try:            self.engine.say(text)            self.engine.runAndWait()        except Exception as e:            raise VoiceError(f"speech synthesis failed: {e}") from eclass NullTTS(TextToSpeech):    """Used in tests and headless containers. Records what would be said."""    def __init__(self): self.spoken: list[str] = []    def speak(self, text: str) -> None: self.spoken.append(text)

NullTTS is not filler. Without it, every test that touches the voice path needs a working audio device, which means your suite fails in CI and inside Docker for reasons that have nothing to do with your code.

The loop, and where the seconds go

Python
class VoiceHandler:    def __init__(self, tts: TextToSpeech | None = None):        self.transcriber = WhisperTranscriber()        self.recorder = MicrophoneRecorder()        self.tts = tts or OfflineTTS()    def listen(self, seconds: float = 5.0) -> str:        pcm = self.recorder.record(seconds)        pcm = trim_silence(normalise(pcm))        if not pcm:            raise VoiceError("no speech detected in the recording")        return self.transcriber.transcribe_bytes(pcm)    def respond(self, text: str) -> None:        self.tts.speak(text)

Now measure the full turn honestly. Times below are a mid-range laptop CPU with Whisper base:

StageTimeCounts toward perceived delay?
User speaks4.0 sNo — the user is busy
Fixed-duration recording runs out1.0 sYes — pure dead air
Normalise and trim0.02 sNegligible
Whisper base on 30 s window1.3 sYes
LLM generates a reply2.1 sYes
TTS synthesis0.3 sYes
Playback begins0.1 sYes

Adding the stages the user actually waits through: 1.0+0.02+1.3+2.1+0.3+0.1=4.821.0 + 0.02 + 1.3 + 2.1 + 0.3 + 0.1 = 4.82 seconds of silence after they stop talking. Human conversational turn-taking runs on gaps of about 200 milliseconds, so nearly five seconds feels broken, and users start repeating themselves — which produces overlapping audio and makes it worse.

The table tells you where to spend effort. The LLM call is the largest block at 2.1 s, and streaming it so synthesis starts on the first complete sentence recovers around 1.5 s. Replacing the fixed 5-second recording with silence detection that stops 0.4 s after the user does recovers most of the 1.0 s. Dropping to tiny saves about 0.8 s at a real accuracy cost. Doing all three lands just under 2 s, which is tolerable. Guessing instead of measuring, and people optimise the 0.02 s preprocessing step.

Latency in a voice pipeline is additive and dominated by one or two stages. Measure each stage separately before optimising anything, or you will spend a day making the fast part faster.

Tests without a microphone

Almost everything here is testable offline if you generate audio instead of recording it.

Python
# tests/test_voice.pyimport numpy as npimport pytestfrom src.voice_handler import normalise, trim_silence, NullTTS, RATEdef sine(seconds: float, freq: int = 440, amplitude: int = 3200) -> bytes:    t = np.linspace(0, seconds, int(RATE * seconds), endpoint=False)    return (amplitude * np.sin(2 * np.pi * freq * t)).astype(np.int16).tobytes()def silence(seconds: float) -> bytes:    return np.zeros(int(RATE * seconds), dtype=np.int16).tobytes()def test_normalise_reaches_target_peak():    out = np.frombuffer(normalise(sine(1.0, amplitude=3200)), dtype=np.int16)    peak = np.abs(out).max()    assert 22_000 < peak < 24_000        # -3 dBFS is 23,200 of 32,767def test_normalise_does_not_clip():    out = np.frombuffer(normalise(sine(1.0, amplitude=32000)), dtype=np.int16)    assert np.abs(out).max() <= 32_767def test_trim_removes_leading_and_trailing_silence():    pcm = silence(2.0) + sine(1.0) + silence(1.5)    trimmed = trim_silence(pcm)    seconds = len(trimmed) / 2 / RATE    assert 1.1 < seconds < 1.4           # 1 s of tone plus 100 ms paddingdef test_pure_silence_trims_to_nothing():    assert trim_silence(silence(3.0)) == b""def test_null_tts_records_without_a_device():    tts = NullTTS()    tts.speak("hello")    assert tts.spoken == ["hello"]

Do not assert on exact transcript text. Whisper's output varies with version and hardware, including leading whitespace and punctuation. If you must test transcription end to end, commit one short WAV of clear speech and assert that the lower-cased result contains an expected keyword.

When things go wrong here

SymptomCauseFix
Transcript is "Thank you." or "Thanks for watching"Near-silent or noise-only input; Whisper hallucinates on silenceThe peak check in record(); trim_silence returning b""; verify mic permission
Transcript is garbled or in the wrong languageCapture rate and WAV header rate disagreeOne RATE constant used by both the stream and setframerate
One phrase repeated many timescondition_on_previous_text feeding a bad window forwardSet it to False
OSError: [Errno -9981] Input overflowedRead loop cannot keep up with the deviceexception_on_overflow=False, and raise frames_per_buffer to 2048
FileNotFoundError: [Errno 2] ... 'ffmpeg'ffmpeg missing from PATHInstall it; restart the shell so PATH is refreshed
pip install pyaudio fails compilingPortAudio headers absentbrew install portaudio or apt-get install portaudio19-dev, then reinstall
Recording is silent on macOS onlyTerminal or IDE lacks microphone permissionSystem Settings, Privacy, Microphone; restart the app after granting
First transcription takes 40 s, later ones 1.3 sModel weights downloading and loadingConstruct WhisperTranscriber once at startup; pre-warm on a short clip
TTS silent in Docker or CINo audio device in the containerInject NullTTS; return audio bytes to the client instead of playing them
Speech is cut off mid-sentenceFixed recording duration expired while the user was still talkingSilence-based endpointing, or raise the duration and rely on trimming

Acceptance criteria for this stage

  1. ffmpeg -version and import pyaudio both succeed before any voice code runs.
  2. A 5-second recording of clear speech produces a transcript containing the expected keywords, with no "Thank you." artefact.
  3. Recording with the microphone muted raises VoiceError naming the peak level — it never returns a fabricated transcript.
  4. normalise() on a signal peaking at 3,200 produces a peak between 22,000 and 24,000, and on one peaking at 32,000 produces no sample beyond ±32,767.
  5. trim_silence() on 2 s silence + 1 s tone + 1.5 s silence returns between 1.1 s and 1.4 s of audio; on pure silence it returns b"".
  6. The full voice test suite passes on a machine with no microphone and no speakers.
  7. Timing instrumentation logs each stage separately, and the sum of stages after the user stops speaking is under 5 seconds on your hardware.
  8. Whisper is loaded exactly once per process — grep the logs for "Whisper ... ready" and confirm a single occurrence per run.

What this buys you for the rest of the build

The whole stage collapses into two method calls: listen() returns a string, respond() consumes a string. Nothing beyond voice_handler.py knows that PortAudio, sample rates, or 30-second windows exist. That is what lets the same orchestrator serve a terminal user typing, a voice user speaking, and an HTTP client posting a WAV file, with the routing logic written once.

The interface choice around TTS pays off twice more. The abstract base class plus NullTTS is what makes the assistant runnable inside a container that has no sound card — you swap in an implementation that returns audio bytes over the wire instead of playing them locally, and every other line of code is unchanged. And it is what keeps the test suite honest: tests that need real hardware get skipped, then rot, then stop protecting you.

The habit worth carrying forward is the peak check. Audio failures are uniquely dangerous because the failure mode is not an exception — it is a plausible wrong answer. Every stage that can silently produce garbage deserves a cheap assertion at its boundary that turns "wrong" into "loudly broken". In this pipeline that assertion costs one line and saves an afternoon every time it fires.