Voice and Speech AI

Course Content

ASR Concepts & Whisper Architecture


Here is a task that sounds easy. Write a program that turns the sentence "I scream for ice cream" into text. You have the microphone recording. Go.

Your first instinct is probably to look for word boundaries — find the silences, cut there, then match each chunk against a dictionary of word sounds. Open the waveform and the plan dies immediately. There is no silence between "I" and "scream". There is no silence between "ice" and "cream" either. The acoustic signal for "I scream" and "ice cream" is, for many speakers, essentially identical. Two different sentences, one waveform.

So you try harder. You decide to match sounds, not words. You record the /s/ sound and template-match it across the audio. It fails again, because the /s/ in "scream" is physically different from the /s/ in "ice" — the mouth is already moving toward the next sound while it produces the current one. Linguists call this coarticulation: sounds bleed into their neighbours. There is no clean, reusable template for any phoneme.

Now add a speaker with a Glaswegian accent, a fan running in the background, and a mobile phone that clips everything above 4 kHz. Automatic Speech Recognition (ASR) is the problem of getting text out of that. It is not a lookup problem. It is a probabilistic inference problem, and understanding why is the fastest route to understanding why modern systems are built the way they are.

What Whisper's encoder actually readsframe0frame1frame2frame3frame4frame5frame6frame7012345670 to 25 mshop 10 ms80 melbins each30 seconds at 16 kHz becomes 3,000 frames of 80 mel bins — a picture, not a sound.
Frames overlap because a 25 ms window is long enough to be worth a spectrum but short enough that speech has not moved.

What makes speech genuinely hard

It helps to name the difficulties precisely, because each one shaped a design decision in systems like Whisper.

DifficultyWhat it meansConcrete example
No word boundariesSpeech is continuous; silence does not mark words"recognise speech" vs "wreck a nice beach"
CoarticulationEach sound is reshaped by its neighboursThe /k/ in "key" and "caw" are acoustically different
Speaker variationPitch, rate, accent, vocal tract length all differAdult male fundamental frequency ~110 Hz, adult female ~200 Hz, child ~300 Hz
HomophonesIdentical sound, different text"their / there / they're" — only context disambiguates
Noise and channelRooms add reverb; codecs remove frequenciesTelephone audio is band-limited to roughly 300–3400 Hz
DisfluencyReal speech has "um", restarts, half-words"I want to— can you book, uh, a flight?"

Speech recognition is not transcription of sounds into letters. It is inference of the most likely sentence given a noisy, ambiguous, speaker-specific signal.

Measuring failure: Word Error Rate

You cannot improve what you cannot score. The standard metric is Word Error Rate (WER), which counts the minimum edits needed to turn the system's output into the correct reference:

WER=S+D+IN\text{WER} = \frac{S + D + I}{N}

where SS is substitutions, DD deletions, II insertions, and NN the number of words in the reference.

Work one through. Reference: "i want to book a flight to boston tomorrow" — that is 9 words, so N=9N = 9. The system outputs "i want book a flight to austin tomorrow". Aligning them, "to" (third word) is missing, so D=1D = 1; "boston" became "austin", so S=1S = 1; nothing was added, so I=0I = 0.

WER=1+1+09=29=0.222=22.2%\text{WER} = \frac{1 + 1 + 0}{9} = \frac{2}{9} = 0.222 = 22.2\%

Two things about WER surprise people. First, it is unbounded above. If the reference is the 3-word "call john smith" and a hallucinating model emits "call john smith please hold on now", that is I=4I = 4 and WER =4/3=133%= 4/3 = 133\%. Second, it weights all errors equally: getting "boston" wrong instead of "austin" costs exactly as much as dropping a "to", even though one sends the customer to the wrong airport and the other is invisible. For languages without whitespace word boundaries — Mandarin, Japanese, Thai — the field uses Character Error Rate (CER) with the same formula applied per character.

How the field got here

The architectural history matters because it explains what Whisper stopped doing.

EraApproachComponents you had to build
1970s–80sTemplate matching, DTWOne template per word; speaker-dependent; tiny vocabulary
1990s–2000sGMM-HMMAcoustic model, pronunciation lexicon, n-gram language model, decoder
2012–2015DNN-HMM (hybrid)Same four pieces, neural acoustic model — roughly 30% relative WER drop
2015–2020End-to-end CTC / RNN-TOne network; lexicon becomes optional
2020–nowEncoder-decoder TransformersOne network; language modelling absorbed into the decoder

The GMM-HMM era required a pronunciation lexicon — a hand-written file saying that "tomato" is /təˈmɑːtoʊ/. Building one for a new language took linguists months. Every end-to-end system since 2015 has been an attempt to delete that file.

Turning sound into numbers

Sound is a pressure wave. A microphone converts pressure to voltage, and an analogue-to-digital converter measures that voltage at fixed intervals. Two parameters define the result.

Sample rate is how many measurements per second. Bit depth is how finely each measurement is quantised. Standard speech ASR uses 16,000 samples per second at 16 bits per sample, mono.

Run the numbers on what that costs. 16,000 samples/second × 2 bytes/sample = 32,000 bytes per second. A 30-second clip is 960,000 bytes, roughly 0.96 MB. An hour of speech is about 115 MB of raw PCM. That is why every ASR pipeline compresses before it does anything else.

Why 16 kHz and not more

The Nyquist–Shannon sampling theorem says a sample rate of fsf_s can faithfully represent frequencies up to fs/2f_s/2, called the Nyquist frequency. Sample at 16 kHz and you capture everything up to 8 kHz.

Is 8 kHz enough for speech? Vowel energy and the formants that distinguish one vowel from another sit below 4 kHz. The fricatives — /s/, /f/, /ʃ/ — carry energy up to about 8 kHz, and they are exactly the sounds telephone audio (band-limited near 3.4 kHz) mangles, which is why "s" and "f" are so easily confused on a bad phone line. Above 8 kHz there is essentially nothing that changes which word was said. Music production uses 44.1 kHz because instruments and cymbals have real content up to 20 kHz; speech does not need it, and doubling the sample rate doubles the compute for no accuracy gain.

Get this wrong in the other direction and you do not just lose information — you actively corrupt it. Feed a 10 kHz component into a 16 kHz sampler without filtering first and it aliases: it reappears as a phantom tone at ∣16,000−10,000∣=6,000|16{,}000 - 10{,}000| = 6{,}000 Hz, right in the middle of the speech band. This is why resampling must always low-pass filter before it downsamples, and why librosa.load(path, sr=16000) is safe while naively taking every third sample is not.

Bit depth follows a similarly clean rule. Dynamic range in decibels is approximately 6.02×bits6.02 \times \text{bits}. At 16 bits that is about 96 dB — the gap between a whisper and a shout, comfortably. At 8 bits it is 48 dB, and quiet consonants disappear into quantisation noise.

Python
import librosaimport numpy as np# Loads, converts to mono, and resamples with proper anti-alias filteringaudio, sr = librosa.load("meeting.wav", sr=16000, mono=True)print(f"sample rate : {sr}")print(f"samples     : {len(audio)}")print(f"duration    : {len(audio) / sr:.2f} s")print(f"dtype       : {audio.dtype}")   # float32, range roughly [-1.0, 1.0]# Peak-normalise so a quiet recording is not fed in at near-zero amplitudepeak = np.max(np.abs(audio))if peak > 0:    audio = audio / peak

From waveform to spectrogram

You now have 480,000 floating-point numbers for 30 seconds of audio. Why not feed them straight to a neural network?

Two reasons, and the second is fatal. First, the raw waveform is a terrible representation of what the ear cares about: the same spoken word recorded twice will have wildly different sample values because the waveform's phase shifts, even though it sounds identical. Second, sequence length. Transformer self-attention costs O(n2)O(n^2). Attending over 480,000 positions means roughly 2.3×10112.3 \times 10^{11} pairwise comparisons per layer. It is not a tuning problem; it is arithmetically impossible.

The fix is the Short-Time Fourier Transform (STFT). Slice the audio into short overlapping windows, and for each window ask "how much energy is at each frequency?" Stack the answers side by side and you get a spectrogram: time on the x-axis, frequency on the y-axis, energy as brightness. Audio becomes an image.

Whisper's exact settings are worth memorising because they explain everything downstream:

ParameterValueWhy
Window length25 ms = 400 samplesLong enough to resolve pitch, short enough that speech is roughly stationary
Hop length10 ms = 160 samplesGives 100 frames per second; enough to catch a fast plosive
FFT size400 (the window, no padding)Yields 400/2+1=201400/2 + 1 = 201 frequency bins
Frequency resolution16000 / 400 = 40 Hz per binFine enough to separate adjacent formants
Mel bins80 (128 in large-v3 and turbo)201 linear bins compressed to a perceptual scale

The mel scale, and why it is not a detail

Human hearing is not linear in frequency. You can easily distinguish 200 Hz from 300 Hz; you cannot distinguish 7000 Hz from 7100 Hz. The mel scale warps frequency to match perception:

m(f)=2595log⁡10 ⁣(1+f700)m(f) = 2595 \log_{10}\!\left(1 + \frac{f}{700}\right)

The scale is anchored so that 1000 Hz maps to 1000 mel. Check the compression with real arithmetic. Take two bands that are both exactly 900 Hz wide:

  • 100–1000 Hz. m(100)=2595log⁡10(1.1429)=150.5m(100) = 2595 \log_{10}(1.1429) = 150.5 mel. m(1000)=1000.0m(1000) = 1000.0 mel. Width: 849.5 mel.
  • 7000–7900 Hz. m(7000)=2595log⁡10(11.0)=2702.4m(7000) = 2595 \log_{10}(11.0) = 2702.4 mel. m(7900)=2595log⁡10(12.286)=2827.0m(7900) = 2595 \log_{10}(12.286) = 2827.0 mel. Width: 124.6 mel.

The same 900 Hz of physical bandwidth gets 849.5 mel of representational budget down low and 124.6 up high — a ratio of about 6.8 to 1. Spending 80 evenly-spaced mel filters therefore means spending most of your resolution exactly where vowels and formants live, and almost none on the region where nothing linguistically distinctive happens.

The last step is a logarithm on the magnitudes. Loudness is perceived logarithmically, and raw spectrogram energies span many orders of magnitude, which makes them miserable input for a network. Taking log⁡\log compresses the range and turns multiplicative channel effects (a microphone that boosts all high frequencies by a factor of three) into additive offsets a network can learn to ignore. The result is the log-mel spectrogram, and it is literally what Whisper eats.

MFCCs, and why they are now mostly historical

Classical ASR went one step further: apply a Discrete Cosine Transform to the log-mel values and keep the first 13 coefficients. These are Mel-Frequency Cepstral Coefficients. The DCT decorrelates the mel bands, which mattered enormously when the acoustic model was a Gaussian Mixture Model that assumed a diagonal covariance matrix — correlated inputs broke that assumption badly.

Neural networks have no such assumption, and the DCT throws away real information for nothing. Modern systems skip it. If you see MFCCs in a modern codebase it is either a speaker-identification system (where the compactness helps) or inherited code nobody revisited.

Python
import librosaimport numpy as npaudio, sr = librosa.load("speech.wav", sr=16000)# Log-mel spectrogram: the modern default, and Whisper's input formatmel = librosa.feature.melspectrogram(    y=audio, sr=sr, n_fft=400, hop_length=160, n_mels=80, fmax=8000)log_mel = librosa.power_to_db(mel, ref=np.max)print("log-mel shape:", log_mel.shape)   # (80, frames) -> 100 frames per second# MFCCs: the classical alternative, shown for contrastmfcc = librosa.feature.mfcc(y=audio, sr=sr, n_mfcc=13, n_fft=400, hop_length=160)print("mfcc shape   :", mfcc.shape)      # (13, frames)

Whisper's architecture

Whisper is an encoder-decoder Transformer — the same shape as a machine translation model — trained to translate audio into text. Nothing about the architecture is novel. What is novel is the training data and the task formulation.

The encoder

Whisper always processes exactly 30 seconds of audio. Shorter clips are zero-padded; longer files are chunked. That fixed window is what makes the shapes constant:

Text
30 s of audio @ 16 kHz          480,000 samples  |  v  STFT: 25 ms window, 10 ms hoplog-mel spectrogram             80 x 3000     (100 frames/s)  |  v  two 1-D convolutions, second with stride 2encoder input sequence          1500 x d_model  (50 positions/s)  |  v  N Transformer encoder blocks with sinusoidal position embeddingsaudio representation            1500 x d_model

Trace the compression. Raw audio arrives at 16,000 numbers per second. The encoder emits 50 vectors per second. That is a 320× reduction in sequence length, and because attention is quadratic, it cuts the attention cost by 3202=102,400×320^2 = 102{,}400\times. This single design choice is what makes the model trainable at all.

Each encoder position covers 20 ms of audio — about the duration of a single phoneme. That is not a coincidence.

The decoder, and the trick that makes it multi-task

The decoder generates text tokens one at a time, attending both to what it has written so far (self-attention) and to the 1500 encoder positions (cross-attention). Cross-attention is the mechanism that lets the decoder "look at" the right moment in the audio while emitting a word, and it is also, usefully, what word-level timestamp tools exploit later.

The clever part is that Whisper does not have separate models for transcription, translation, and language detection. It has one model, and the task is specified by special tokens in the decoder prompt:

Text
<|startoftranscript|> <|en|> <|transcribe|> <|0.00|> Hello there. <|2.48|> <|endoftext|>                        ^        ^              ^                   language    task        timestamp token
TokenEffect
<|en|>, <|de|>, …Declares the source language. Omit it and the model predicts it from the audio.
<|transcribe|>Output text in the source language
<|translate|>Output English regardless of source language
<|notimestamps|>Emit plain text only
<|0.00|> … <|30.00|>1501 timestamp tokens in 0.02 s steps

Notice that the timestamp granularity is 0.02 s = 20 ms, exactly one encoder position. The vocabulary was designed backwards from the encoder's frame rate.

Whisper turns "transcribe", "translate" and "detect language" into three different prompts for one network, rather than three different networks.

Model sizes

ModelParametersLayersdmodeld_{model}Approx. VRAMRelative speed
tiny39 M4384~1 GB~10×
base74 M6512~1 GB~7×
small244 M12768~2 GB~4×
medium769 M241024~5 GB~2×
large-v31550 M321280~10 GB1×
large-v3-turbo (turbo)809 M32 encoder, 4 decoder1280~6 GB~8×

The .en variants (tiny.en, base.en, small.en, medium.en) are English-only and measurably better on English than their multilingual twins at the same size, because none of their capacity is spent on 98 other languages. There is no large.en. If you are certain your audio is English and you are using a small model, the .en variant is free accuracy.

turbo (large-v3-turbo, released in 2024) keeps the large-v3 encoder but cuts the decoder from 32 layers to 4. Since decoding is the sequential part, it runs about eight times faster than large-v3 for a small accuracy loss, and it is the usual choice when you want large-model quality at a moderate cost. One catch: it was not trained for translation, so use medium or large-v3 for task="translate".

Why Whisper is unusually robust

The training recipe explains the behaviour you will observe in practice. Whisper large-v2 was trained on 680,000 hours of audio — about 78 years of continuous speech — scraped from the web with existing transcripts, not hand-annotated in a studio. Roughly 117,000 of those hours were non-English, spanning 96 languages, and 125,000 hours were "X to English" translation pairs. large-v3 extended this to roughly 1 million weakly-labelled hours plus 4 million pseudo-labelled hours.

This is weak supervision: the labels are imperfect. Some are auto-generated by other ASR systems, some are subtitles that paraphrase rather than transcribe. The team filtered out machine-generated transcripts (they teach a model to imitate another model's mistakes) but otherwise embraced the mess.

The messiness is the point. Studio-clean training data produces a model that collapses on a phone call. Whisper's training set already contains podcasts with music beds, lecture recordings with room echo, videos with two people talking over each other, and speakers with every accent that appears online. It never sees a clean distribution, so it never learns to depend on one. The measurable consequence: Whisper is often worse than a specialised model on a clean benchmark like LibriSpeech test-clean, and dramatically better on real-world audio the specialist has never seen. It generalises rather than specialises.

Where Whisper goes wrong

Four failure modes cause almost all production incidents, and each has a specific cause.

Hallucination on silence. Whisper's decoder is a language model. Given 30 seconds of near-silence or background noise, it will confidently emit fluent text that was never spoken — commonly "Thank you for watching!" or "Subtitles by the Amara.org community", phrases that ended thousands of the YouTube videos in its training set. The model is not lying; it is doing exactly what a language model does when the audio evidence is uninformative. The fix is to gate transcription behind a Voice Activity Detector so silent regions are never sent to the model at all.

Repetition loops. The decoder can enter a cycle, repeating a phrase until it hits the token limit. This is standard autoregressive degeneration. The mitigation built into Whisper is temperature fallback: decode greedily first, and if the output has suspiciously low average log-probability or high compression ratio (a strong signal of repetition), re-decode with increasing temperature.

Chunk-boundary damage. Files longer than 30 seconds are split. A naive split at exactly 30.0 s will cut a word in half, and both halves transcribe badly. Overlapping windows or silence-aware splitting fixes this.

Language misdetection. Auto-detection reads only the first 30 seconds. A recording that opens with 20 seconds of English pleasantries before switching to Portuguese will be transcribed as English throughout — badly. If you know the language, pass it explicitly; it is both faster and more accurate.

A confident-sounding transcript of silence is the single most dangerous ASR failure, because nothing downstream can tell it apart from a real one.

Running it

Bash
pip install -U openai-whisper      # reference implementationpip install -U faster-whisper     # CTranslate2 port: ~4x faster, less memory# ffmpeg is required for audio decodingbrew install ffmpeg               # or: apt-get install ffmpeg
Python
import whispermodel = whisper.load_model("small")# Transcribe. Passing the language skips detection and avoids the 30 s trap.result = model.transcribe(    "interview.mp3",    language="en",    temperature=(0.0, 0.2, 0.4, 0.6, 0.8, 1.0),  # greedy first, then fall back    condition_on_previous_text=False,             # reduces repetition drift)print(result["text"])for seg in result["segments"]:    start, end = seg["start"], seg["end"]    conf = seg["avg_logprob"]        # closer to 0 is more confident    print(f"[{start:6.2f} -> {end:6.2f}] (lp={conf:5.2f}) {seg['text'].strip()}")

Switching the task to translation is a one-word change, because it is the same model with a different prompt token:

Python
# Spanish audio in, English text outresult = model.transcribe("entrevista.mp3", task="translate")print(result["text"])

The faster-whisper port has the same semantics with a streaming generator interface and built-in voice activity detection, which directly addresses the hallucination problem:

Python
from faster_whisper import WhisperModelmodel = WhisperModel("small", device="cuda", compute_type="float16")segments, info = model.transcribe(    "interview.mp3",    vad_filter=True,                              # drop silence before decoding    vad_parameters={"min_silence_duration_ms": 500},)print(f"detected language: {info.language} (p={info.language_probability:.2f})")for seg in segments:                              # lazily evaluated    print(f"[{seg.start:.2f} -> {seg.end:.2f}] {seg.text}")

Choosing, when it is your system

The decision that actually matters is not "which model is best" but "which error rate can my product tolerate at what cost".

Start by asking whether a human reads the output. A transcript displayed to a user tolerates a 10% WER, because readers repair errors from context without noticing. A transcript feeding a keyword alert or an automated action tolerates almost none, because a substituted proper noun silently produces a wrong outcome. Those two products should not use the same model.

Then measure on your audio, not a benchmark. Take fifty clips from your real source — the actual microphones, the actual room, the actual accents — transcribe them by hand once, and compute WER for tiny, base, small, and medium. The result is frequently that small matches medium on your domain while running twice as fast, or that base is catastrophically worse on your accent while being fine on LibriSpeech. Published benchmark numbers cannot tell you this.

Put more than Whisper in that test. At the time of writing (2026), open models such as NVIDIA's Parakeet and Canary families beat Whisper on English leaderboards and run much faster, though they cover fewer languages. Hosted APIs — OpenAI's gpt-transcribe and gpt-4o-transcribe (with whisper-1 kept as a legacy option), Google Cloud Speech-to-Text, Deepgram, AssemblyAI — trade a per-minute fee for zero operations work. Model names in this list change every few months; the ranking on your own fifty clips is the one that counts.

Finally, budget for the failure modes rather than hoping they disappear. Every production ASR service should have a voice activity detector in front of the model, a length-and-repetition sanity check on the output, an explicit language setting where the language is known, and a confidence threshold below which a segment is flagged rather than trusted. None of these are optimisations. They are the difference between a system that degrades gracefully and one that invents a sentence nobody said and passes it downstream as fact.