Voice and Speech AI

Course Content

Neural Text-to-Speech (TTS) Models


Try to build a text-to-speech system the obvious way. Record a voice actor saying every word in English — a big job, but finite. Store the clips. When text arrives, look up each word and play the clips in order.

Feed it the sentence: "Dr. Smith lives on Martin Luther King Dr."

Your system says "Doctor Smith lives on Martin Luther King Doctor". The abbreviation is identical in both positions and means two different things. Fine, you add a rule: "Dr." at the end of an address is "Drive". Now feed it "I read the book yesterday" and it says "reed" instead of "red" — the same five letters, two pronunciations, distinguished only by tense, which is not in your dictionary.

Push past both. Play the clips for "the quick brown fox" back to back and it sounds like a ransom note. Each word was recorded in isolation with a falling pitch, so your sentence has four separate endings — no rising intonation on a question, no stress on the important word, no shortening of unstressed syllables. Every word is correct and the result is unlistenable.

Those three failures map onto the three hard problems in speech synthesis: text normalisation (what characters actually mean), grapheme-to-phoneme conversion (how letters map to sounds), and prosody (the melody, rhythm and stress that make speech sound like a person rather than a lookup table). Every TTS architecture is an attempt at those three, and neural TTS won because it solved the third almost by accident.

From a string to a waveform, in two modelsRaw text: "Dr. Smith paid 1,024 pounds in 2019"Front-end: normalisenumbers, dates and abbreviationsGrapheme-to-phoneme: the sounds, not the spellingAcoustic model: phonemes to a mel-spectrogramVocoder: mel-spectrogram to 22 kHz samples
The split is the point: the acoustic model decides what is said and how, and the vocoder only has to make it sound real.

The shape of a TTS system

Modern neural TTS has three stages, and knowing where each one lives makes every later decision easier.

Text
  "I'll pay $5 on 3/4."        |        v   TEXT FRONT-END      normalisation + phonemisation  "AY1 L P EY1 F AY1 V D AA1 L ER0 Z AA1 N M AA1 R CH F AO1 R TH"        |        v   ACOUSTIC MODEL      Tacotron 2 / FastSpeech 2 / Glow-TTS  mel-spectrogram, 80 x T       ~86 frames per second        |        v   VOCODER             HiFi-GAN / Vocos / WaveNet  waveform, 22,050 samples/s

The middle stage decides what the speech sounds like — rhythm, pitch, timbre — compressed into a mel-spectrogram, a picture of how energy spreads across frequency over time. The vocoder's job is narrower and harder than it looks: turn that picture back into a pressure wave. The mel-spectrogram discarded phase, so the vocoder must invent it, and inventing it badly is what produces the characteristic robotic buzz.

Count the ratio. At 22,050 Hz with a 256-sample hop, the acoustic model produces 22050/256=86.1322050 / 256 = 86.13 mel frames per second. The vocoder must emit 22,050 samples per second. Every single mel frame becomes 256 audio samples. The vocoder is doing 256× upsampling, and it is where most of the compute and most of the perceptual quality lives.

Why the old approaches failed

EraApproachHow it worksWhy it lost
1980s–2000sFormant synthesisModel the vocal tract as filters; generate sound from physics rulesUnmistakably robotic. Tiny footprint — ran on 1980s hardware
1990s–2010sConcatenative (unit selection)Record hours of one speaker, cut into diphones, search for the best sequenceExcellent in-domain, audible glitches at joins; one voice per 20+ studio hours
2000sStatistical parametric (HMM)Learn statistics of acoustic parameters, generate from the modelNever glitches, always muffled — averaging over training data smooths away detail
2016–nowNeural end-to-endOne network learns text → acoustics from paired dataWon on quality; costs GPUs and data

The concatenative-versus-parametric contrast is the clearest illustration of a trade-off that keeps reappearing. Concatenative systems play real recorded audio, so at their best they are indistinguishable from a person — and at their worst, when the search finds no good unit, they leave an audible seam mid-word. Parametric systems generate from a statistical model, so they never glitch, but the model predicts the average of all the ways a sound could be produced, and the average of many sharp spectra is a blurry one. That blur is what "muffled" means acoustically.

Concatenative synthesis is occasionally perfect and occasionally broken; parametric synthesis is reliably mediocre. Neural synthesis is the first approach that is reliably good.

The front-end, which nobody wants to build

Before any model runs, raw text must become something pronounceable. This is unglamorous rule-and-data work and it causes a large fraction of real TTS bugs.

Text normalisation

Written text is full of tokens that are not words. Each needs expansion, and the correct expansion is context-dependent:

InputCorrect spoken formWhy it is hard
3/4"three quarters" or "March fourth" or "three out of four"Requires surrounding context
1984"nineteen eighty-four" (year) or "one thousand nine hundred eighty-four" (quantity)Same digits, different reading
Dr."Doctor" or "Drive"Position in the sentence decides
$5.2M"five point two million dollars"Currency symbol moves to the end; suffix expands
iOS 17.4"i O S seventeen point four"Letters read individually, digits not
2-3"two to three" or "two minus three" or "two dash three"Semantics, not syntax

The currency case shows the shape of the problem. The symbol appears before the number in writing and the word appears after it in speech, and the plural depends on the value: 1 dollar, 5 dollars, 0.5 dollars. No character-level rule handles this; it needs a real number-to-words module.

Grapheme-to-phoneme conversion

English orthography is famously unfaithful to pronunciation. The letter sequence "ough" is pronounced four different ways in though, through, tough and thought. A model trained on characters must learn all of this from data; a model trained on phonemes gets it handed over.

Phonemisation converts text to a phonetic alphabet — ARPAbet or IPA — usually via a pronunciation dictionary (CMUdict holds around 134,000 entries) with a learned fallback for out-of-vocabulary words. The gain is largest where you would expect: names, technical terms, and words the model saw rarely in training.

Python
from g2p_en import G2pg2p = G2p()print(g2p("I read the book"))# ['AY1', ' ', 'R', 'IY1', 'D', ' ', 'DH', 'AH0', ' ', 'B', 'UH1', 'K']   (' ' marks a word gap)print(g2p("Kubernetes orchestrates containers"))# out-of-dictionary words fall back to a learned model

Note that g2p_en read "read" as R IY1 D, the present tense. It already runs a part-of-speech tagger to choose between homographs, but "I read the book" is genuinely ambiguous without the surrounding sentences, so the tagger guesses present tense — tagging helps, and still leaves residual errors. This is why character-input models did not simply disappear: they can learn contextual pronunciation a dictionary cannot express.

Acoustic models: text to mel-spectrogram

Three architectures define the design space, and they differ on one axis that determines almost everything: how the model decides how long each sound should last.

Tacotron 2 — attention decides duration implicitly

Tacotron 2 is an encoder-decoder with attention. The encoder reads the character or phoneme sequence. The decoder produces one mel frame at a time, and at each step an attention mechanism decides which input symbol it is currently "reading". Duration is never predicted; it emerges from the attention lingering on a symbol for several frames.

The result was a landmark — the paper reported a Mean Opinion Score of 4.53 against 4.58 for real recorded human speech, statistically indistinguishable to listeners. But two properties limit it in production.

First, it is autoregressive: frame tt needs frame t−1t-1. Five seconds of speech is 5×86.13≈4315 \times 86.13 \approx 431 mel frames, and therefore 431 strictly sequential decoder passes. You cannot parallelise them and a GPU sits mostly idle.

Second, and worse, attention can fail. Nothing forces it to move monotonically forward through the text, so it sometimes skips a word, repeats one, or gets stuck babbling until the stop token fires. Rare on ordinary sentences; common enough on long inputs, unusual punctuation or repeated words to be a genuine production hazard.

FastSpeech 2 — duration predicted explicitly

FastSpeech 2 removes attention from the generation path. It predicts, for each input phoneme, an explicit duration in mel frames, then simply repeats that phoneme's encoding that many times before decoding everything in one parallel pass.

Text
phonemes:   HH   AH0   L   OW1durations:   4     3    5    9      (frames, predicted)                |                v  length regulator: repeat each encodingexpanded:   HH HH HH HH AH0 AH0 AH0 L L L L L OW1 x9   -> 21 frames                |                v  decoder, all 21 frames at oncemel-spectrogram 80 x 21

At 22,050 Hz with a 256-sample hop each frame is 256/22050=11.6256/22050 = 11.6 ms, so those 21 frames are 244 ms of audio — a plausible "hello".

Three consequences follow. The whole utterance decodes in one forward pass rather than 431, roughly an order of magnitude faster. Words can no longer be skipped or repeated, because every phoneme is guaranteed to appear for its predicted duration. And speed becomes a dial: multiply every duration by 1/α1/\alpha. With α=1.25\alpha = 1.25, a 431-frame utterance becomes 431×0.8=345431 \times 0.8 = 345 frames, so 5.0 seconds becomes 4.0 — and pitch is unaffected, unlike naive playback-rate change, which makes everyone sound like a chipmunk.

FastSpeech 2 also predicts pitch and energy per frame as separate outputs, which is what lets you control intonation and loudness independently of the words. This is called a variance adaptor, and it is the mechanism behind most controllable-prosody features you will encounter.

The cost: duration labels have to come from somewhere. FastSpeech 2 extracts them by forced alignment against the training audio, which is an extra pipeline stage and an extra place to go wrong.

Glow-TTS — alignment learned without a teacher

Glow-TTS is a normalising flow: an invertible network mapping a simple distribution to the mel distribution. Being invertible, it can run backwards to compute exact likelihoods, which lets it find the optimal monotonic alignment between phonemes and frames during training via monotonic alignment search — no external aligner, and no attention failures, because monotonicity is enforced by construction.

It is fast, parallel and stable, and it is the conceptual ancestor of VITS, which fuses flow-based modelling with a vocoder into one end-to-end network running text straight to waveform, no mel-spectrogram in between. VITS and its descendants remain the workhorse of lightweight open-source TTS for that reason: fewer stages, fewer places for quality to leak. Many of the newest high-quality open models take a different route — a language model that generates discrete audio-codec tokens (XTTS, used in the next lesson, works this way) or flow matching — at a higher compute cost.

Tacotron 2FastSpeech 2Glow-TTS / VITS
GenerationAutoregressiveParallelParallel
AlignmentLearned attentionExternal alignerMonotonic alignment search
Skips / repeats wordsOccasionallyNeverNever
Speed controlNot directlyDuration multiplierDuration multiplier
Inference speedSlowestFastFast
Prosody varietyNatural, variedFlatter without extra conditioningVaried (stochastic duration in VITS)
Use whenResearch, maximum naturalness offlineProduction, latency mattersProduction default today

The "flatter" note on FastSpeech 2 is a real trade. Deterministic duration prediction gives the average rhythm for a sentence, and average rhythm sounds mechanical. VITS uses a stochastic duration predictor that samples a different plausible rhythm each time, so the same sentence generated twice sounds subtly different — as it would from a person.

Vocoders: mel-spectrogram to waveform

This is where the "256 samples per frame" number bites.

WaveNet, and the wall it hit

WaveNet models audio one sample at a time, conditioning each sample on all the samples before it. The quality was a step change in 2016. The speed was catastrophic: generating one second of 24 kHz audio requires 24,000 sequential forward passes. Even at an optimistic 0.1 ms per pass, that is 2.4 seconds of compute for 1 second of audio — a real-time factor of 2.4, meaning it can never keep up with speech, let alone stream. The original implementation was far slower than that.

HiFi-GAN, and why a GAN was the right answer

HiFi-GAN generates the entire waveform in a single parallel pass using transposed convolutions that upsample the mel-spectrogram by 8×8×2×2=2568 \times 8 \times 2 \times 2 = 256× — exactly the hop length. No sequential dependency at all.

The training is the interesting part. A plain regression loss on the waveform fails badly: two waveforms can sound identical while differing sample-by-sample (a small phase shift), so the loss punishes correct outputs and the model hedges by producing something smooth and buzzy. HiFi-GAN instead trains against discriminators that judge whether audio sounds real, including a multi-period discriminator that reshapes the waveform into 2D by periods 2, 3, 5, 7 and 11. Speech is quasi-periodic, and viewing it through several prime-numbered periods catches artefacts a single view misses.

The result is real-time factors around 0.006 on a GPU: roughly 167× faster than real time, at quality close to WaveNet's.

Vocos, and the shortcut

Vocos observes that the transposed convolutions doing the upsampling are the expensive part, and that they reinvent what the inverse Fourier transform already does perfectly. So it predicts Fourier coefficients — magnitude and phase — at the mel frame rate and hands them to an inverse STFT. All neural computation happens at 86 frames per second instead of 22,050, a 256× reduction in the network's work.

VocoderGenerationTypical RTF (GPU)QualityWhen to use
Griffin-LimIterative phase estimation, no neural net~0.05, CPU onlyPoor — audibly metallicDebugging your acoustic model
WaveNetSample-by-sample, autoregressive> 1.0ExcellentOffline reference quality only
WaveRNNSample-by-sample, small RNN~0.5–1.0Very goodOn-device, historically
HiFi-GANParallel convolutional~0.006ExcellentProduction default
VocosParallel, inverse STFT head~0.001ExcellentLowest latency, CPU-friendly

Griffin-Lim earns its row for one practical reason: it needs no training. If your speech sounds wrong, run the mel-spectrogram through Griffin-Lim. Intelligible-but-metallic means the acoustic model is fine and the vocoder is at fault; gibberish means the acoustic model is. That five-minute test saves hours.

Autoregressive vocoders lost not because they sounded worse, but because generating audio one sample at a time is arithmetically incompatible with real-time speech.

Three tools, three trade-offs

Python
# 1. gTTS - wraps Google Translate's TTS. Simplest possible option.from gtts import gTTSgTTS("Your order has shipped.", lang="en", tld="co.uk").save("out.mp3")# No control over voice, speed or prosody. Requires network. Not for production.# 2. Hugging Face - full control, runs locally, weights are yoursimport torch, soundfile as sffrom transformers import SpeechT5Processor, SpeechT5ForTextToSpeech, SpeechT5HifiGanfrom datasets import load_datasetproc = SpeechT5Processor.from_pretrained("microsoft/speecht5_tts")model = SpeechT5ForTextToSpeech.from_pretrained("microsoft/speecht5_tts")vocoder = SpeechT5HifiGan.from_pretrained("microsoft/speecht5_hifigan")# Current `datasets` no longer runs loading scripts, so read the Parquet exportemb_ds = load_dataset("Matthijs/cmu-arctic-xvectors",                      revision="refs/convert/parquet", split="validation")speaker = torch.tensor(emb_ds[7306]["xvector"]).unsqueeze(0)   # 512-dim identityinputs = proc(text="Your order has shipped.", return_tensors="pt")speech = model.generate_speech(inputs["input_ids"], speaker, vocoder=vocoder)sf.write("out.wav", speech.numpy(), samplerate=16000)
Python
# 3. Hosted API - highest quality per unit of effort, streams, pay per useimport osfrom openai import OpenAIclient = OpenAI()with client.audio.speech.with_streaming_response.create(    model=os.environ.get("TTS_MODEL", "gpt-4o-mini-tts"),  # older: tts-1, tts-1-hd    voice="marin",    input="Your order has shipped.",    instructions="Warm and brisk, like a helpful shop assistant.",  # newer models only    response_format="opus",   # opus for streaming, mp3 for storage, wav for processing) as response:    response.stream_to_file("out.opus")
gTTSLocal (SpeechT5 / VITS)Hosted API
Setup effortOne lineModel download, GPU configAPI key
QualityAcceptableGood to excellentExcellent
LatencyNetwork round-tripLowest, if on GPUNetwork + generation
Cost modelFree, unofficial, may breakFixed hardwarePer character or per token
Data leaves your networkYesNoYes
Voice controlNoneTotal — you can fine-tuneFixed voice list; tone via text instructions on newer models
OfflineNoYesNo

For regulated work the deciding row is "data leaves your network". A medical or legal application synthesising patient text cannot send it to a third-party endpoint, whatever the quality advantage.

Controlling how it sounds

Three levers, in increasing order of power.

SSML (Speech Synthesis Markup Language) is XML annotating text with prosodic instructions. Support varies wildly — hosted cloud voices implement it well; most open-source models ignore it entirely.

Text
<speak>  Your order has <emphasis level="strong">shipped</emphasis>.  <break time="400ms"/>  It arrives on <say-as interpret-as="date" format="md">3/4</say-as>.  <prosody rate="90%" pitch="-2st">Track it in your account.</prosody></speak>

Duration and pitch scaling acts directly on the acoustic model's outputs, which is why FastSpeech-family models expose it. Scaling durations changes speed without changing pitch; shifting the predicted pitch contour changes voice height without changing speed. Naive audio-rate manipulation cannot separate the two.

Reference-audio conditioning is the most powerful: give the model a clip of someone speaking in the style you want, extract a style embedding, and condition generation on it. This is how "read this in an excited tone" features work, and it is the same machinery that underlies voice cloning.

Streaming, and where the latency actually goes

For a written article, total generation time is what matters. For a conversation, only time to first audio matters — the gap between the user finishing their sentence and hearing the first syllable of a reply. A system that generates a 60-second answer in 3 seconds feels broken if it plays nothing for those 3 seconds.

The fix is to pipeline at sentence granularity: generate the first sentence, start playing it, generate the rest while it plays. An eight-word first sentence is roughly 2.5 seconds of audio; at RTF 0.05 it synthesises in about 125 ms, and its playback covers generating everything after it.

StageNon-streaming (60 s of speech)Sentence-streamed
Front-end normalisation~15 ms (whole text)~3 ms (first sentence)
Acoustic model~1,600 ms~70 ms
Vocoder~360 ms~15 ms
Encode + buffer~40 ms~15 ms
Time to first audio~2,015 ms~103 ms

Same model, same hardware, same total compute — a twentyfold difference in perceived responsiveness, purely from ordering the work differently. The one requirement is that text arrives incrementally too: wait for a language model to finish before you start synthesising and the advantage is gone.

Evaluating a voice

TTS quality is fundamentally subjective, so the primary metric is human: Mean Opinion Score. Listeners rate samples from 1 (bad) to 5 (excellent) and you average. Natural human speech typically scores around 4.5, not 5.0, because recording conditions and rater strictness put a ceiling on it.

The statistics matter more than teams expect. With NN raters and rating standard deviation near 1.0, the 95% confidence interval on the mean is roughly

±1.96×σN=±1.96×1.030=±0.36\pm 1.96 \times \frac{\sigma}{\sqrt{N}} = \pm 1.96 \times \frac{1.0}{\sqrt{30}} = \pm 0.36

So with 30 raters, a system scoring 4.1 and one scoring 4.3 are statistically indistinguishable. Resolving a 0.1 MOS gap needs roughly (1.96/0.05)2≈1,537(1.96 / 0.05)^2 \approx 1{,}537 raters. Almost every "beats theirs by 0.15 MOS" claim is under-powered noise. For small differences use paired A/B preference tests instead — far more sensitive, because each rater controls for their own scale.

MetricMeasuresCostWatch out for
MOSOverall naturalnessHigh — human ratersNeeds large N; scale drifts between studies
A/B preferenceWhich of two is betterMediumCannot compare across experiments
UTMOS / MOSNetPredicted MOS from a modelFreeTrained on specific data; can be gamed
Mel-cepstral distortionSpectral distance from a referenceFreeNeeds paired reference; punishes valid variation
ASR word error rateIntelligibilityFreeCatches skipped words; says nothing about naturalness
Speaker similarityDoes it sound like the target voiceFreeOnly meaningful for cloning

Running ASR over your synthesised audio and computing WER against the input text is the cheapest high-value check you can automate. It says nothing about whether the voice is pleasant, but it catches a skipped word, a mangled number, or a mispronounced name — the failures that actually reach users.

Three beliefs worth correcting

"Higher MOS means better for my product." MOS measures naturalness on read sentences. A navigation system needs a voice that stays intelligible at 60 mph with the window open, which tracks clarity and dynamic range far more than naturalness. Pick a metric matching your real listening condition.

"The vocoder is a solved detail." It is where phase is invented and perceptual quality is won or lost. Swapping Griffin-Lim for HiFi-GAN, same acoustic model, is the single largest quality jump available in most pipelines.

"Neural TTS removed the need for a text front-end." End-to-end models learn the pronunciations that appear in their training data. Your product names, ticker symbols and customers' surnames do not appear there. Every production system ends up with a normalisation layer and a pronunciation override dictionary.

Choosing, when you have to ship

Start with your latency budget. If a user will hear the audio within a second of speaking, you need a parallel acoustic model, a parallel vocoder and sentence-level streaming — which rules out anything autoregressive and rules out waiting for complete text. If you are generating an audiobook overnight, none of that applies and you optimise purely for quality.

Then decide where the audio can be generated. An on-premises requirement collapses the option space to open-source models, and the practical default there is a single end-to-end model — a VITS-family voice such as Piper, or a compact newer model such as Kokoro: one end-to-end network is far less to operate than an acoustic model plus a separately-versioned vocoder plus a forced aligner.

Whatever you pick, build two things before tuning anything. One is a pronunciation override table mapping your domain's difficult tokens to explicit phonemes — you will need it within a week of launch, and retrofitting it means re-testing everything. The other is an automated intelligibility check: synthesise a fixed set of sentences containing your product's names, numbers and edge cases, transcribe them with ASR, and fail the build if word error rate rises. Human listening tests are too slow for every change; that check runs in seconds and catches the errors users complain about.