Voice and Speech AI

Course Content

Transcription Pipelines & Real-World Deployment


A team ships a podcast transcription feature. It works beautifully in the demo: a 90-second clip, transcribed in four seconds, punctuation intact. They launch. The first real customer uploads a three-hour interview.

The worker process allocates 691 MB decoding the audio, another gigabyte for the intermediate arrays, and gets killed by the container's memory limit. They raise the limit to 8 GB. Now it survives, but takes 41 minutes, and the HTTP request timed out at 30 seconds so the customer already refreshed and uploaded it twice more. They add a job queue. Now it finishes, and the transcript contains the phrase "the authent ication flow" — a word sliced in half at a chunk boundary — 104 separate times.

Every one of those failures is a pipeline problem, not a model problem. The model was correct throughout. What was missing was everything around it: how audio gets in, how it is divided, what happens when a request fails, how the pieces are reassembled, and how you find out any of this went wrong.

A production transcription system is roughly 15% model and 85% plumbing. This is the plumbing.

A 90-second demo against a three-hour interviewBatch a whole file• Decode to 16 kHz mono first• Never load three hours into RAM• Chunk on silence, not on the clock• Chunks can run in parallelStream live audio• Fixed ring buffer, oldest dropped• Latency budget caps chunk length• Overlap so words are not cut in half• Partial hypotheses get revised
Fixed-size chunking destroys transcripts by cutting mid-word; overlapping and cutting at silence is the whole fix.

The shape of the pipeline

Every serious transcription system, regardless of scale, has the same six stages. Naming them separately matters because each one fails differently and each one needs its own metric.

Text
  ingest        normalise      segment        transcribe      stitch        format    |               |             |               |              |             | file / URL  ->  16 kHz mono ->  chunks  ->   ASR model  ->  merge + fix  ->  SRT / JSON mic stream      loudness        with          (batched or     overlaps       VTT / text                 normalised      overlap        parallel)      dedupe    |               |             |               |              |             | fails: 404      fails: codec   fails: cuts    fails: OOM,    fails: dupe   fails: bad auth, huge      unsupported,   mid-word,      rate limit,    words at      timestamps file            wrong sr       silence gap    hallucination  seams         drift

The bottom row is the part people skip. Design each stage around its failure, not around its happy path.

Getting audio in

Three ingest modes cover almost everything, and they have genuinely different constraints.

ModeLatency profileMain riskRight approach
File uploadBatch; seconds to hoursMemory blowup on long filesStream from disk, never load whole
Live microphoneContinuous; must keep upBuffer overrun, dropped framesRing buffer + VAD-triggered segments
Remote URLBatch, plus networkUntrusted size, slow origin, redirectsStream with a hard byte cap and timeout

Files: the memory trap

Do the arithmetic that killed the team above. Three hours is 10,800 seconds. At 16,000 samples per second that is 172,800,000 samples. Decoded to float32 at 4 bytes each, that is 691 MB for the array alone — before the decoder's own buffers, before the mel spectrogram, before the model weights. A single naive librosa.load() on a long file is the most common cause of an OOM kill in transcription services.

The fix is to never hold the whole file. soundfile can read in blocks, and the pipeline only ever needs one chunk resident at a time:

Python
import soundfile as sfimport numpy as npimport librosadef stream_chunks(path, chunk_seconds=30, target_sr=16000):    """Yield (start_time, audio) without ever loading the whole file."""    info = sf.info(path)    block = int(chunk_seconds * info.samplerate)    position = 0    with sf.SoundFile(path) as f:        while True:            data = f.read(block, dtype="float32")            if len(data) == 0:                break            n_read = len(data)                      # count in the file's own rate            if data.ndim > 1:                       # downmix stereo to mono                data = data.mean(axis=1)            if info.samplerate != target_sr:                data = librosa.resample(                    data, orig_sr=info.samplerate, target_sr=target_sr                )            yield position / info.samplerate, data            position += n_read

Peak memory is now one chunk — 30 seconds at 16 kHz float32 is 1.92 MB — regardless of whether the file is three minutes or thirteen hours.

Remote URLs: never trust the other end

Downloading from a user-supplied URL is a security surface, not a convenience. Two rules: stream rather than buffer, and enforce a byte cap yourself rather than trusting the Content-Length header, which a hostile origin will happily lie about.

Python
import requests, tempfileMAX_BYTES = 500 * 1024 * 1024   # 500 MB hard capdef fetch_audio(url, timeout=30):    with requests.get(url, stream=True, timeout=timeout, allow_redirects=True) as r:        r.raise_for_status()        total = 0        tmp = tempfile.NamedTemporaryFile(suffix=".audio", delete=False)        for block in r.iter_content(chunk_size=1 << 20):            total += len(block)            if total > MAX_BYTES:                tmp.close()                raise ValueError(f"audio exceeds {MAX_BYTES} byte cap")            tmp.write(block)        tmp.close()        return tmp.name

Live audio: the ring buffer

Microphone input arrives whether or not you are ready for it. If your transcription step takes 400 ms and the callback fires every 100 ms, you must buffer or you drop audio. The standard shape is a callback that only ever appends to a thread-safe queue, and a separate consumer that assembles segments:

Python
import queue, threadingimport numpy as npimport sounddevice as sdSR = 16000audio_q = queue.Queue()def callback(indata, frames, time_info, status):    if status:        print("stream status:", status)      # xruns show up here first    audio_q.put(indata[:, 0].copy())         # copy: the buffer is reuseddef consume(min_seconds=2.0, silence_rms=0.01, silence_seconds=0.7):    buf, silent_for = [], 0.0    while True:        block = audio_q.get()        buf.append(block)        rms = float(np.sqrt(np.mean(block ** 2)))        block_seconds = len(block) / SR        silent_for = silent_for + block_seconds if rms < silence_rms else 0.0        held = sum(len(b) for b in buf) / SR        if held >= min_seconds and silent_for >= silence_seconds:            yield np.concatenate(buf)            buf, silent_for = [], 0.0stream = sd.InputStream(samplerate=SR, channels=1, blocksize=1600, callback=callback)

The critical line is indata[:, 0].copy(). The audio library reuses that buffer for the next callback; keeping a reference instead of a copy gives you a queue full of identical, corrupted blocks. This bug is subtle, silent, and extremely common.

Chunking: where transcripts get quietly destroyed

Whisper processes 30 seconds at a time, so long audio must be divided. How you divide determines your error rate.

Fixed-size chunking, and its exact cost

The naive approach cuts every 30.0 seconds. Consider what happens when a speaker says "authentication" spanning 29.8 s to 30.4 s. Chunk one ends mid-word and the model transcribes the fragment as "authent". Chunk two starts mid-word and produces "ication". One cut, two errors.

Quantify it on a three-hour file. 10,800 seconds divided into 30-second chunks gives 360 chunks and therefore 359 internal boundaries. At normal speaking rates of about 150 words per minute, roughly a third of arbitrary cut points land inside a word rather than in a gap — call it 120 damaged boundaries producing 240 word errors. The file contains around 27,000 words, so this contributes

24027,000=0.89%\frac{240}{27{,}000} = 0.89\%

to your Word Error Rate, on top of whatever the model's own error rate is. If your model achieves 5% WER, you have just made it 5.9% for free, and the errors are concentrated in a distinctive, embarrassing pattern that users notice immediately.

Overlap: the cheap fix

Extend each chunk backwards by a couple of seconds so every boundary is covered twice, then discard the duplicate region from the second chunk. With 30-second chunks and 2-second overlap, the stride becomes 28 seconds. For 10,800 seconds:

chunks=⌈10,800−228⌉=⌈385.6⌉=386\text{chunks} = \left\lceil \frac{10{,}800 - 2}{28} \right\rceil = \lceil 385.6 \rceil = 386

That is 386 chunks instead of 360 — 7.2% more compute to eliminate a 0.89% WER penalty. It is close to the best trade available anywhere in the pipeline.

Silence-aware chunking: the better fix

Better still, cut where nobody is speaking. Voice activity detection finds the gaps, and you split at the largest gap near your target boundary rather than at the boundary itself. This eliminates mid-word cuts entirely instead of merely patching them, and has a second benefit: silent stretches never reach the model, which is the primary defence against hallucinated text.

Python
import numpy as npimport librosadef silence_aware_chunks(audio, sr=16000, target=28.0, max_len=30.0, search=4.0):    """Split near `target` seconds, but snap to the quietest nearby point."""    # Frame-level energy, 10 ms hop    hop = int(0.01 * sr)    rms = librosa.feature.rms(y=audio, frame_length=int(0.025 * sr), hop_length=hop)[0]    chunks, start = [], 0    total = len(audio)    while start < total:        ideal = start + int(target * sr)        if ideal >= total:            chunks.append((start / sr, audio[start:]))            break        lo = max(start + int(sr), ideal - int(search * sr))        hi = min(total, ideal + int(search * sr), start + int(max_len * sr))        window = rms[lo // hop: hi // hop]        if len(window) == 0:            cut = ideal        else:            cut = lo + int(np.argmin(window)) * hop   # quietest frame in range        chunks.append((start / sr, audio[start:cut]))        start = cut    return chunks

Chunk boundaries are not an implementation detail — a naive split adds roughly a percentage point of word error rate before the model has done anything wrong.

Doing the chunks in parallel — and when that helps

Here is where a lot of code goes wrong for a reason worth understanding. Whether parallelism helps depends entirely on where the time is going.

SetupBottleneckDoes a thread pool help?What to do instead
Hosted API (OpenAI, etc.)Network round-tripYes, enormously — threads wait, they do not compute8–16 workers, respect rate limits
Single local GPUGPU computeNo — the device serialises anywayIncrease batch size within one call
Multiple local GPUsGPU computeYes — one worker pinned per deviceThread pool sized to device count
CPU inferenceCPU computePartly — the GIL is released inside native opsProcess pool, or let the runtime use all cores

The API case is the dramatic one. Suppose 360 chunks, each taking about 3 seconds of round-trip time. Serially that is 1,080 seconds — 18 minutes. With 8 concurrent workers it is roughly 135 seconds, because the threads spend virtually all of their time blocked on the network rather than holding the interpreter lock.

Python
from concurrent.futures import ThreadPoolExecutor, as_completeddef transcribe_all(chunks, transcribe_fn, workers=8):    results = [None] * len(chunks)    with ThreadPoolExecutor(max_workers=workers) as pool:        futures = {            pool.submit(transcribe_fn, audio): (i, offset)            for i, (offset, audio) in enumerate(chunks)        }        for fut in as_completed(futures):            i, offset = futures[fut]            try:                text, segments = fut.result()            except Exception as e:                 # one bad chunk must not kill the job                results[i] = {"offset": offset, "text": "", "error": str(e)}                continue            # Shift local timestamps into whole-file time            for s in segments:                s["start"] += offset                s["end"] += offset            results[i] = {"offset": offset, "text": text, "segments": segments}    return results

Two details make this correct rather than merely concurrent. Results are written into a pre-sized list indexed by chunk number, so completion order does not scramble the transcript. And every chunk's local timestamps are shifted by its offset — forget this and your subtitles all claim to start at zero.

Failing well

Transient failures are certain at scale: rate limits, connection resets, a GPU that throws once. Retrying immediately makes things worse, because whatever caused the failure is usually still happening.

Exponential backoff waits longer after each failure. With a base delay of 1 second and a factor of 2, the waits are 1, 2, 4, 8, 16 seconds — a total of 31 seconds across five retries, giving an overloaded service real time to recover.

But exponential backoff alone has a nasty failure mode. If 200 clients all hit a rate limit at the same instant, they all sleep exactly 1 second and all retry at exactly the same instant — a thundering herd that re-triggers the limit forever. Jitter is the fix: randomise the delay so the retries spread out.

Python
import time, random, loggingRETRIABLE = (ConnectionError, TimeoutError)def with_backoff(fn, *args, attempts=6, base=1.0, cap=30.0, **kwargs):    for n in range(attempts):        try:            return fn(*args, **kwargs)        except RETRIABLE as e:            if n == attempts - 1:                logging.error("giving up after %d attempts: %s", attempts, e)                raise            # Full jitter: uniform in [0, min(cap, base * 2^n)]            delay = random.uniform(0, min(cap, base * (2 ** n)))            logging.warning("attempt %d failed (%s); retrying in %.2fs", n + 1, e, delay)            time.sleep(delay)

Note the exception tuple. Retrying a ValueError from a corrupt file, or a 401 from a bad API key, wastes up to 31 seconds to fail identically. Only retry things that might succeed next time.

Retry only transient errors, back off exponentially, and always add jitter — synchronised retries are how a brief blip becomes a sustained outage.

Post-processing: the raw transcript is not the product

Model output needs work before anyone should see it. These steps are boring and they are where perceived quality actually comes from.

StepProblem it solvesExample
Seam deduplicationOverlapped chunks transcribe the same words twice"...and then we and then we deployed"
Domain vocabularyProper nouns and jargon are systematically wrong"Kubernetes" → "cuber netties"
Number normalisationSpoken form is unreadable in text"twenty twenty four" → "2024"
Filler removalVerbatim transcripts read badlyStrip "um", "uh", "you know" — optionally
Hallucination screeningSilence produces confident invented textDrop segments with low avg log-probability
Speaker attributionMulti-speaker audio is unreadable as a wallMerge diarisation output onto segments

Seam deduplication deserves an actual algorithm. When chunks overlap by 2 seconds, the tail of chunk nn and the head of chunk n+1n{+}1 describe the same audio. Compare the last few words of one against the first few of the next and drop the longest match:

Python
def dedupe_seam(prev_text, next_text, max_overlap=12):    """Drop the longest word-sequence that ends prev_text and starts next_text."""    a, b = prev_text.split(), next_text.split()    limit = min(max_overlap, len(a), len(b))    for k in range(limit, 0, -1):        if [w.lower().strip(".,!?") for w in a[-k:]] == \           [w.lower().strip(".,!?") for w in b[:k]]:            return " ".join(b[k:])    return next_text

Hallucination screening is the highest-value filter. Whisper reports avg_logprob per segment (closer to zero means more confident) and no_speech_prob. A segment with no_speech_prob above about 0.6 and avg_logprob below about −1.0 is almost always invented. Dropping those costs you a handful of real words and saves you from shipping sentences nobody said.

Output formats

Two subtitle formats cover nearly all use. They are nearly identical and the differences will bite you.

SRTWebVTT
HeaderNoneWEBVTT plus blank line, required
Cue numbersRequired, 1-indexedOptional
Decimal separatorComma: 00:01:02,500Full stop: 00:01:02.500
Hours fieldAlways presentOptional below one hour
Where usedVideo players, legacy toolingHTML5 <track> element

The comma-versus-full-stop difference is the single most common subtitle bug. A WebVTT file with commas parses as zero cues in a browser, silently — the video plays with no subtitles and no error.

Timestamp formatting is pure arithmetic. Take 3,725.5 seconds: dividing by 3,600 gives 1 hour with 125.5 seconds left; dividing that by 60 gives 2 minutes with 5.5 seconds left. So the value is 01:02:05,500 in SRT and 01:02:05.500 in WebVTT.

Python
def fmt(seconds, vtt=False):    total_ms = int(round(seconds * 1000))   # round once, so 1.9996 s -> 00:00:02,000    h, rest = divmod(total_ms, 3_600_000)    m, rest = divmod(rest, 60_000)    s, ms = divmod(rest, 1000)    sep = "." if vtt else ","    return f"{h:02d}:{m:02d}:{s:02d}{sep}{ms:03d}"def to_srt(segments):    out = []    for i, seg in enumerate(segments, start=1):        out.append(str(i))        out.append(f"{fmt(seg['start'])} --> {fmt(seg['end'])}")        out.append(seg["text"].strip())        out.append("")    return "\n".join(out)def to_vtt(segments):    out = ["WEBVTT", ""]    for seg in segments:        out.append(f"{fmt(seg['start'], vtt=True)} --> {fmt(seg['end'], vtt=True)}")        out.append(seg["text"].strip())        out.append("")    return "\n".join(out)

One more constraint that has nothing to do with code: readability. A subtitle cue should hold at most two lines of about 42 characters and stay on screen for at least one second. Whisper's segments routinely exceed this, so a production formatter re-splits long segments at clause boundaries and redistributes the timing proportionally.

Making it fast enough

The metric to track is real-time factor (RTF): processing time divided by audio duration. RTF of 0.1 means a one-hour recording takes six minutes. RTF above 1.0 means you can never keep up with live audio.

ChangeTypical effectCost
medium → small model~2× faster1–3 points of WER on hard audio
float32 → float16 on GPU~2× faster, half the VRAMNegligible accuracy change
faster-whisper (CTranslate2)~4× faster than the reference implementationNone; same weights
int8 quantisation on CPU~2–3× faster, ~4× less memorySmall WER increase, usually under 1 point
VAD filteringProportional to silence removedNone — also reduces hallucination
Batching chunks on GPULarge, if VRAM allowsHigher peak memory

VAD filtering is the one people forget and it is often the biggest single win. A recorded meeting is frequently 40% silence. Removing it removes 40% of the compute and produces a better transcript, because the silence was the part the model was inventing text over.

On GPU memory: keep the model loaded once per worker process rather than reloading per request. Loading medium takes several seconds and allocates about 5 GB; doing that per request means most of your latency is initialisation. Load at process start, transcribe many times.

Knowing it is working

A transcription service can fail without erroring. It returns 200, it returns text, and the text is wrong. Standard application monitoring catches none of this, so you need domain-specific signals.

SignalWhat it detectsAlert when
RTF per jobModel or hardware degradationp95 rises above your SLA budget
Mean avg_logprobAudio quality drop, wrong languageFalls below your historical baseline
Empty-transcript rateBroken ingest, silent uploadsAny sustained increase
Compression ratio of outputRepetition loopsRatio above ~2.4 on any segment
Golden-set WERSilent regression after a changeRises more than 1 point
Retry rateUpstream instabilityExceeds a few percent of chunks

The golden set is the one that matters most and the one teams skip. Hand-transcribe thirty clips from your real audio source, store them in the repository, and re-run WER on every deploy. Without it, a change to chunking or a model upgrade can degrade quality by five points and nobody notices for a month.

Python
import logging, time, jsondef transcribe_job(path, transcribe_fn):    t0 = time.perf_counter()    result = transcribe_fn(path)    elapsed = time.perf_counter() - t0    duration = result["duration"]    segs = result["segments"]    logging.info(json.dumps({        "event": "transcription_complete",        "audio_seconds": round(duration, 2),        "wall_seconds": round(elapsed, 2),        "rtf": round(elapsed / max(duration, 1e-6), 3),        "segments": len(segs),        "chars": sum(len(s["text"]) for s in segs),        "mean_logprob": round(            sum(s["avg_logprob"] for s in segs) / max(len(segs), 1), 3),        "language": result.get("language"),    }))    return result

Structured logs, one JSON object per job. Every field above is queryable, and the ratio between chars and audio_seconds alone will catch a hallucinating model within minutes — real speech produces roughly 12 to 18 characters per second, and a repetition loop produces hundreds.

What this looks like when you build it

The system you should actually build is smaller than the list above suggests, because the stages compose. Ingest streams from disk or a capped download. Segment with VAD and a small overlap. Transcribe through a retry wrapper, in parallel only if the bottleneck is network. Stitch with offset correction and seam deduplication. Screen segments by confidence. Format to VTT or SRT with the right decimal separator. Log one structured record per job.

What separates a working service from a demo is that each of those stages assumes the next one might receive garbage. The chunker assumes a file might be thirteen hours. The transcriber assumes a chunk might be pure silence. The stitcher assumes a chunk might have failed entirely and returned nothing. The formatter assumes a segment might be 40 seconds long and need re-splitting.

Build in that order and add the golden-set WER check before you add anything clever. Most teams reach for a bigger model when their transcripts are bad; more often the model is fine and the chunker is cutting words in half.