Course Content
Voice and Speech AI
3 sections · 5 lessons
Transcription Pipelines & Real-World Deployment
A team ships a podcast transcription feature. It works beautifully in the demo: a 90-second clip, transcribed in four seconds, punctuation intact. They launch. The first real customer uploads a three-hour interview.
The worker process allocates 691 MB decoding the audio, another gigabyte for the intermediate arrays, and gets killed by the container's memory limit. They raise the limit to 8 GB. Now it survives, but takes 41 minutes, and the HTTP request timed out at 30 seconds so the customer already refreshed and uploaded it twice more. They add a job queue. Now it finishes, and the transcript contains the phrase "the authent ication flow" — a word sliced in half at a chunk boundary — 104 separate times.
Every one of those failures is a pipeline problem, not a model problem. The model was correct throughout. What was missing was everything around it: how audio gets in, how it is divided, what happens when a request fails, how the pieces are reassembled, and how you find out any of this went wrong.
A production transcription system is roughly 15% model and 85% plumbing. This is the plumbing.
The shape of the pipeline
Every serious transcription system, regardless of scale, has the same six stages. Naming them separately matters because each one fails differently and each one needs its own metric.
ingest normalise segment transcribe stitch format | | | | | | file / URL -> 16 kHz mono -> chunks -> ASR model -> merge + fix -> SRT / JSON mic stream loudness with (batched or overlaps VTT / text normalised overlap parallel) dedupe | | | | | | fails: 404 fails: codec fails: cuts fails: OOM, fails: dupe fails: bad auth, huge unsupported, mid-word, rate limit, words at timestamps file wrong sr silence gap hallucination seams driftThe bottom row is the part people skip. Design each stage around its failure, not around its happy path.
Getting audio in
Three ingest modes cover almost everything, and they have genuinely different constraints.
| Mode | Latency profile | Main risk | Right approach |
|---|---|---|---|
| File upload | Batch; seconds to hours | Memory blowup on long files | Stream from disk, never load whole |
| Live microphone | Continuous; must keep up | Buffer overrun, dropped frames | Ring buffer + VAD-triggered segments |
| Remote URL | Batch, plus network | Untrusted size, slow origin, redirects | Stream with a hard byte cap and timeout |
Files: the memory trap
Do the arithmetic that killed the team above. Three hours is 10,800 seconds. At 16,000 samples per second that is 172,800,000 samples. Decoded to float32 at 4 bytes each, that is 691 MB for the array alone — before the decoder's own buffers, before the mel spectrogram, before the model weights. A single naive librosa.load() on a long file is the most common cause of an OOM kill in transcription services.
The fix is to never hold the whole file. soundfile can read in blocks, and the pipeline only ever needs one chunk resident at a time:
1import soundfile as sf2import numpy as np3import librosa45def stream_chunks(path, chunk_seconds=30, target_sr=16000):6 """Yield (start_time, audio) without ever loading the whole file."""7 info = sf.info(path)8 block = int(chunk_seconds * info.samplerate)9 position = 01011 with sf.SoundFile(path) as f:12 while True:13 data = f.read(block, dtype="float32")14 if len(data) == 0:15 break16 n_read = len(data) # count in the file's own rate17 if data.ndim > 1: # downmix stereo to mono18 data = data.mean(axis=1)19 if info.samplerate != target_sr:20 data = librosa.resample(21 data, orig_sr=info.samplerate, target_sr=target_sr22 )23 yield position / info.samplerate, data24 position += n_readPeak memory is now one chunk — 30 seconds at 16 kHz float32 is 1.92 MB — regardless of whether the file is three minutes or thirteen hours.
Remote URLs: never trust the other end
Downloading from a user-supplied URL is a security surface, not a convenience. Two rules: stream rather than buffer, and enforce a byte cap yourself rather than trusting the Content-Length header, which a hostile origin will happily lie about.
1import requests, tempfile23MAX_BYTES = 500 * 1024 * 1024 # 500 MB hard cap45def fetch_audio(url, timeout=30):6 with requests.get(url, stream=True, timeout=timeout, allow_redirects=True) as r:7 r.raise_for_status()8 total = 09 tmp = tempfile.NamedTemporaryFile(suffix=".audio", delete=False)10 for block in r.iter_content(chunk_size=1 << 20):11 total += len(block)12 if total > MAX_BYTES:13 tmp.close()14 raise ValueError(f"audio exceeds {MAX_BYTES} byte cap")15 tmp.write(block)16 tmp.close()17 return tmp.nameLive audio: the ring buffer
Microphone input arrives whether or not you are ready for it. If your transcription step takes 400 ms and the callback fires every 100 ms, you must buffer or you drop audio. The standard shape is a callback that only ever appends to a thread-safe queue, and a separate consumer that assembles segments:
1import queue, threading2import numpy as np3import sounddevice as sd45SR = 160006audio_q = queue.Queue()78def callback(indata, frames, time_info, status):9 if status:10 print("stream status:", status) # xruns show up here first11 audio_q.put(indata[:, 0].copy()) # copy: the buffer is reused1213def consume(min_seconds=2.0, silence_rms=0.01, silence_seconds=0.7):14 buf, silent_for = [], 0.015 while True:16 block = audio_q.get()17 buf.append(block)18 rms = float(np.sqrt(np.mean(block ** 2)))19 block_seconds = len(block) / SR20 silent_for = silent_for + block_seconds if rms < silence_rms else 0.02122 held = sum(len(b) for b in buf) / SR23 if held >= min_seconds and silent_for >= silence_seconds:24 yield np.concatenate(buf)25 buf, silent_for = [], 0.02627stream = sd.InputStream(samplerate=SR, channels=1, blocksize=1600, callback=callback)The critical line is indata[:, 0].copy(). The audio library reuses that buffer for the next callback; keeping a reference instead of a copy gives you a queue full of identical, corrupted blocks. This bug is subtle, silent, and extremely common.
Chunking: where transcripts get quietly destroyed
Whisper processes 30 seconds at a time, so long audio must be divided. How you divide determines your error rate.
Fixed-size chunking, and its exact cost
The naive approach cuts every 30.0 seconds. Consider what happens when a speaker says "authentication" spanning 29.8 s to 30.4 s. Chunk one ends mid-word and the model transcribes the fragment as "authent". Chunk two starts mid-word and produces "ication". One cut, two errors.
Quantify it on a three-hour file. 10,800 seconds divided into 30-second chunks gives 360 chunks and therefore 359 internal boundaries. At normal speaking rates of about 150 words per minute, roughly a third of arbitrary cut points land inside a word rather than in a gap — call it 120 damaged boundaries producing 240 word errors. The file contains around 27,000 words, so this contributes
to your Word Error Rate, on top of whatever the model's own error rate is. If your model achieves 5% WER, you have just made it 5.9% for free, and the errors are concentrated in a distinctive, embarrassing pattern that users notice immediately.
Overlap: the cheap fix
Extend each chunk backwards by a couple of seconds so every boundary is covered twice, then discard the duplicate region from the second chunk. With 30-second chunks and 2-second overlap, the stride becomes 28 seconds. For 10,800 seconds:
That is 386 chunks instead of 360 — 7.2% more compute to eliminate a 0.89% WER penalty. It is close to the best trade available anywhere in the pipeline.
Silence-aware chunking: the better fix
Better still, cut where nobody is speaking. Voice activity detection finds the gaps, and you split at the largest gap near your target boundary rather than at the boundary itself. This eliminates mid-word cuts entirely instead of merely patching them, and has a second benefit: silent stretches never reach the model, which is the primary defence against hallucinated text.
1import numpy as np2import librosa34def silence_aware_chunks(audio, sr=16000, target=28.0, max_len=30.0, search=4.0):5 """Split near `target` seconds, but snap to the quietest nearby point."""6 # Frame-level energy, 10 ms hop7 hop = int(0.01 * sr)8 rms = librosa.feature.rms(y=audio, frame_length=int(0.025 * sr), hop_length=hop)[0]910 chunks, start = [], 011 total = len(audio)12 while start < total:13 ideal = start + int(target * sr)14 if ideal >= total:15 chunks.append((start / sr, audio[start:]))16 break1718 lo = max(start + int(sr), ideal - int(search * sr))19 hi = min(total, ideal + int(search * sr), start + int(max_len * sr))2021 window = rms[lo // hop: hi // hop]22 if len(window) == 0:23 cut = ideal24 else:25 cut = lo + int(np.argmin(window)) * hop # quietest frame in range2627 chunks.append((start / sr, audio[start:cut]))28 start = cut29 return chunksChunk boundaries are not an implementation detail — a naive split adds roughly a percentage point of word error rate before the model has done anything wrong.
Doing the chunks in parallel — and when that helps
Here is where a lot of code goes wrong for a reason worth understanding. Whether parallelism helps depends entirely on where the time is going.
| Setup | Bottleneck | Does a thread pool help? | What to do instead |
|---|---|---|---|
| Hosted API (OpenAI, etc.) | Network round-trip | Yes, enormously — threads wait, they do not compute | 8–16 workers, respect rate limits |
| Single local GPU | GPU compute | No — the device serialises anyway | Increase batch size within one call |
| Multiple local GPUs | GPU compute | Yes — one worker pinned per device | Thread pool sized to device count |
| CPU inference | CPU compute | Partly — the GIL is released inside native ops | Process pool, or let the runtime use all cores |
The API case is the dramatic one. Suppose 360 chunks, each taking about 3 seconds of round-trip time. Serially that is 1,080 seconds — 18 minutes. With 8 concurrent workers it is roughly 135 seconds, because the threads spend virtually all of their time blocked on the network rather than holding the interpreter lock.
1from concurrent.futures import ThreadPoolExecutor, as_completed23def transcribe_all(chunks, transcribe_fn, workers=8):4 results = [None] * len(chunks)5 with ThreadPoolExecutor(max_workers=workers) as pool:6 futures = {7 pool.submit(transcribe_fn, audio): (i, offset)8 for i, (offset, audio) in enumerate(chunks)9 }10 for fut in as_completed(futures):11 i, offset = futures[fut]12 try:13 text, segments = fut.result()14 except Exception as e: # one bad chunk must not kill the job15 results[i] = {"offset": offset, "text": "", "error": str(e)}16 continue17 # Shift local timestamps into whole-file time18 for s in segments:19 s["start"] += offset20 s["end"] += offset21 results[i] = {"offset": offset, "text": text, "segments": segments}22 return resultsTwo details make this correct rather than merely concurrent. Results are written into a pre-sized list indexed by chunk number, so completion order does not scramble the transcript. And every chunk's local timestamps are shifted by its offset — forget this and your subtitles all claim to start at zero.
Failing well
Transient failures are certain at scale: rate limits, connection resets, a GPU that throws once. Retrying immediately makes things worse, because whatever caused the failure is usually still happening.
Exponential backoff waits longer after each failure. With a base delay of 1 second and a factor of 2, the waits are 1, 2, 4, 8, 16 seconds — a total of 31 seconds across five retries, giving an overloaded service real time to recover.
But exponential backoff alone has a nasty failure mode. If 200 clients all hit a rate limit at the same instant, they all sleep exactly 1 second and all retry at exactly the same instant — a thundering herd that re-triggers the limit forever. Jitter is the fix: randomise the delay so the retries spread out.
1import time, random, logging23RETRIABLE = (ConnectionError, TimeoutError)45def with_backoff(fn, *args, attempts=6, base=1.0, cap=30.0, **kwargs):6 for n in range(attempts):7 try:8 return fn(*args, **kwargs)9 except RETRIABLE as e:10 if n == attempts - 1:11 logging.error("giving up after %d attempts: %s", attempts, e)12 raise13 # Full jitter: uniform in [0, min(cap, base * 2^n)]14 delay = random.uniform(0, min(cap, base * (2 ** n)))15 logging.warning("attempt %d failed (%s); retrying in %.2fs", n + 1, e, delay)16 time.sleep(delay)Note the exception tuple. Retrying a ValueError from a corrupt file, or a 401 from a bad API key, wastes up to 31 seconds to fail identically. Only retry things that might succeed next time.
Retry only transient errors, back off exponentially, and always add jitter — synchronised retries are how a brief blip becomes a sustained outage.
Post-processing: the raw transcript is not the product
Model output needs work before anyone should see it. These steps are boring and they are where perceived quality actually comes from.
| Step | Problem it solves | Example |
|---|---|---|
| Seam deduplication | Overlapped chunks transcribe the same words twice | "...and then we and then we deployed" |
| Domain vocabulary | Proper nouns and jargon are systematically wrong | "Kubernetes" → "cuber netties" |
| Number normalisation | Spoken form is unreadable in text | "twenty twenty four" → "2024" |
| Filler removal | Verbatim transcripts read badly | Strip "um", "uh", "you know" — optionally |
| Hallucination screening | Silence produces confident invented text | Drop segments with low avg log-probability |
| Speaker attribution | Multi-speaker audio is unreadable as a wall | Merge diarisation output onto segments |
Seam deduplication deserves an actual algorithm. When chunks overlap by 2 seconds, the tail of chunk n and the head of chunk n+1 describe the same audio. Compare the last few words of one against the first few of the next and drop the longest match:
1def dedupe_seam(prev_text, next_text, max_overlap=12):2 """Drop the longest word-sequence that ends prev_text and starts next_text."""3 a, b = prev_text.split(), next_text.split()4 limit = min(max_overlap, len(a), len(b))5 for k in range(limit, 0, -1):6 if [w.lower().strip(".,!?") for w in a[-k:]] == \7 [w.lower().strip(".,!?") for w in b[:k]]:8 return " ".join(b[k:])9 return next_textHallucination screening is the highest-value filter. Whisper reports avg_logprob per segment (closer to zero means more confident) and no_speech_prob. A segment with no_speech_prob above about 0.6 and avg_logprob below about −1.0 is almost always invented. Dropping those costs you a handful of real words and saves you from shipping sentences nobody said.
Output formats
Two subtitle formats cover nearly all use. They are nearly identical and the differences will bite you.
| SRT | WebVTT | |
|---|---|---|
| Header | None | WEBVTT plus blank line, required |
| Cue numbers | Required, 1-indexed | Optional |
| Decimal separator | Comma: 00:01:02,500 | Full stop: 00:01:02.500 |
| Hours field | Always present | Optional below one hour |
| Where used | Video players, legacy tooling | HTML5 <track> element |
The comma-versus-full-stop difference is the single most common subtitle bug. A WebVTT file with commas parses as zero cues in a browser, silently — the video plays with no subtitles and no error.
Timestamp formatting is pure arithmetic. Take 3,725.5 seconds: dividing by 3,600 gives 1 hour with 125.5 seconds left; dividing that by 60 gives 2 minutes with 5.5 seconds left. So the value is 01:02:05,500 in SRT and 01:02:05.500 in WebVTT.
1def fmt(seconds, vtt=False):2 total_ms = int(round(seconds * 1000)) # round once, so 1.9996 s -> 00:00:02,0003 h, rest = divmod(total_ms, 3_600_000)4 m, rest = divmod(rest, 60_000)5 s, ms = divmod(rest, 1000)6 sep = "." if vtt else ","7 return f"{h:02d}:{m:02d}:{s:02d}{sep}{ms:03d}"89def to_srt(segments):10 out = []11 for i, seg in enumerate(segments, start=1):12 out.append(str(i))13 out.append(f"{fmt(seg['start'])} --> {fmt(seg['end'])}")14 out.append(seg["text"].strip())15 out.append("")16 return "\n".join(out)1718def to_vtt(segments):19 out = ["WEBVTT", ""]20 for seg in segments:21 out.append(f"{fmt(seg['start'], vtt=True)} --> {fmt(seg['end'], vtt=True)}")22 out.append(seg["text"].strip())23 out.append("")24 return "\n".join(out)One more constraint that has nothing to do with code: readability. A subtitle cue should hold at most two lines of about 42 characters and stay on screen for at least one second. Whisper's segments routinely exceed this, so a production formatter re-splits long segments at clause boundaries and redistributes the timing proportionally.
Making it fast enough
The metric to track is real-time factor (RTF): processing time divided by audio duration. RTF of 0.1 means a one-hour recording takes six minutes. RTF above 1.0 means you can never keep up with live audio.
| Change | Typical effect | Cost |
|---|---|---|
| medium → small model | ~2× faster | 1–3 points of WER on hard audio |
| float32 → float16 on GPU | ~2× faster, half the VRAM | Negligible accuracy change |
| faster-whisper (CTranslate2) | ~4× faster than the reference implementation | None; same weights |
| int8 quantisation on CPU | ~2–3× faster, ~4× less memory | Small WER increase, usually under 1 point |
| VAD filtering | Proportional to silence removed | None — also reduces hallucination |
| Batching chunks on GPU | Large, if VRAM allows | Higher peak memory |
VAD filtering is the one people forget and it is often the biggest single win. A recorded meeting is frequently 40% silence. Removing it removes 40% of the compute and produces a better transcript, because the silence was the part the model was inventing text over.
On GPU memory: keep the model loaded once per worker process rather than reloading per request. Loading medium takes several seconds and allocates about 5 GB; doing that per request means most of your latency is initialisation. Load at process start, transcribe many times.
Knowing it is working
A transcription service can fail without erroring. It returns 200, it returns text, and the text is wrong. Standard application monitoring catches none of this, so you need domain-specific signals.
| Signal | What it detects | Alert when |
|---|---|---|
| RTF per job | Model or hardware degradation | p95 rises above your SLA budget |
Mean avg_logprob | Audio quality drop, wrong language | Falls below your historical baseline |
| Empty-transcript rate | Broken ingest, silent uploads | Any sustained increase |
| Compression ratio of output | Repetition loops | Ratio above ~2.4 on any segment |
| Golden-set WER | Silent regression after a change | Rises more than 1 point |
| Retry rate | Upstream instability | Exceeds a few percent of chunks |
The golden set is the one that matters most and the one teams skip. Hand-transcribe thirty clips from your real audio source, store them in the repository, and re-run WER on every deploy. Without it, a change to chunking or a model upgrade can degrade quality by five points and nobody notices for a month.
1import logging, time, json23def transcribe_job(path, transcribe_fn):4 t0 = time.perf_counter()5 result = transcribe_fn(path)6 elapsed = time.perf_counter() - t07 duration = result["duration"]8 segs = result["segments"]910 logging.info(json.dumps({11 "event": "transcription_complete",12 "audio_seconds": round(duration, 2),13 "wall_seconds": round(elapsed, 2),14 "rtf": round(elapsed / max(duration, 1e-6), 3),15 "segments": len(segs),16 "chars": sum(len(s["text"]) for s in segs),17 "mean_logprob": round(18 sum(s["avg_logprob"] for s in segs) / max(len(segs), 1), 3),19 "language": result.get("language"),20 }))21 return resultStructured logs, one JSON object per job. Every field above is queryable, and the ratio between chars and audio_seconds alone will catch a hallucinating model within minutes — real speech produces roughly 12 to 18 characters per second, and a repetition loop produces hundreds.
What this looks like when you build it
The system you should actually build is smaller than the list above suggests, because the stages compose. Ingest streams from disk or a capped download. Segment with VAD and a small overlap. Transcribe through a retry wrapper, in parallel only if the bottleneck is network. Stitch with offset correction and seam deduplication. Screen segments by confidence. Format to VTT or SRT with the right decimal separator. Log one structured record per job.
What separates a working service from a demo is that each of those stages assumes the next one might receive garbage. The chunker assumes a file might be thirteen hours. The transcriber assumes a chunk might be pure silence. The stitcher assumes a chunk might have failed entirely and returned nothing. The formatter assumes a segment might be 40 seconds long and need re-splitting.
Build in that order and add the golden-set WER check before you add anything clever. Most teams reach for a bigger model when their transcripts are bad; more often the model is fine and the chunker is cutting words in half.