Course Content
Voice and Speech AI
3 sections · 5 lessons
Voice Cloning & Advanced Audio Synthesis
You have thirty seconds of a colleague's voice from a voicemail. You want a system that can say any new sentence in that voice. Take the obvious route: fine-tune a text-to-speech model on the recording.
It fails before it starts. Thirty seconds at 22,050 Hz with a 256-sample hop is about 2,580 mel frames of training signal. A modest TTS decoder has tens of millions of parameters. The model does not learn the voice; it memorises the clip, and asked to say anything new it produces noise, or it produces the original sentence regardless of what you typed. The classical answer was to record twenty hours in a studio and fine-tune for a day on a GPU — which is why, until about 2018, having a synthetic version of your voice was a thing that happened to celebrities and nobody else.
The trick that broke this open is a change in what you are learning. Do not learn the voice. Learn, once, across thousands of speakers, what makes voices different from each other — and then represent any particular voice as a short vector in that space. The TTS model is trained once, on many speakers, to be conditioned on such a vector. Cloning a new voice then requires no training at all: extract the vector from a few seconds of audio and pass it in.
That works. Six seconds of reference audio is now enough to produce a recognisable clone in roughly 300 milliseconds. It works so well that in 2019 a UK energy company's chief executive transferred €220,000 to a fraudulent account because the voice on the phone was, unmistakably, his German parent-company boss. Both facts are the same fact. This lesson covers the mechanism and it covers what the mechanism obliges you to do, because the two are not separable.
Speaker embeddings: compressing "who" out of audio
A speaker embedding is a fixed-length vector — commonly 192 or 512 dimensions — that captures speaker identity and discards content. Feed it audio of Alice saying "hello" and audio of Alice reading a legal contract, and you should get nearly the same vector. Feed it Bob saying "hello" and you should get a distant one.
The compression is severe and that severity is the point. A 30-second reference clip at 16 kHz in float32 is 1,920,000 bytes. A 512-dimensional float32 embedding is 2,048 bytes. That is a 937× reduction, and everything that survives it is, by construction, speaker identity: vocal tract length, glottal characteristics, habitual pitch range, accent, speaking-rate tendencies. The words are gone.
How the encoder learns to do this
You cannot supervise this directly, because there is no ground-truth "identity vector" for a person. Instead you train with a metric-learning objective: sample audio clips, and push the network to make two clips from the same speaker land close together while two clips from different speakers land far apart.
The GE2E (Generalised End-to-End) loss does this in batches structured as N speakers × M utterances each. For every utterance, compute its similarity to the centroid of its own speaker and to the centroids of all other speakers, then apply a softmax that rewards the correct centroid. Training on thousands of speakers — VoxCeleb2 has 6,112 — forces the network to discover the axes along which human voices actually vary, rather than memorising any particular person.
| Method | Era | How it works | Dimensions | Limitation |
|---|---|---|---|---|
| i-vector | 2011 | Factor analysis over GMM supervectors; entirely statistical | 400–600 | Needs long utterances; degrades badly under noise |
| d-vector | 2014 | Average the hidden layer of a speaker-classification DNN | 256 | Trained for classification, not similarity |
| x-vector | 2018 | TDNN with statistics pooling over time | 512 | Strong baseline; superseded on accuracy |
| GE2E / ECAPA-TDNN | 2018–2020 | Metric learning with attentive pooling and channel attention | 192–256 | Current default; needs large multi-speaker data |
The jump from d-vector to GE2E is a lesson in objective design. A d-vector network is trained to answer "which of these 5,000 speakers is this?" — so it learns features that separate those 5,000, with no pressure to generalise to a speaker it has never seen. GE2E is trained to answer "are these two clips the same person?", which is exactly the question you ask at cloning time. Same data, same architecture, radically different generalisation, because the training question matched the deployment question.
Train for the question you will actually ask at inference time. A network trained to classify known speakers will not reliably embed unknown ones.
1import torch2from speechbrain.inference import EncoderClassifier34encoder = EncoderClassifier.from_hparams(5 source="speechbrain/spkrec-ecapa-voxceleb", savedir="/tmp/ecapa"6)78def embed(wav_path):9 signal = encoder.load_audio(wav_path) # expects 16 kHz mono10 with torch.no_grad():11 e = encoder.encode_batch(signal.unsqueeze(0)).squeeze()12 return torch.nn.functional.normalize(e, dim=0) # unit norm: cosine == dot1314a1, a2, b1 = embed("alice_1.wav"), embed("alice_2.wav"), embed("bob_1.wav")1516print("alice vs alice:", float(a1 @ a2)) # typically 0.75 - 0.9017print("alice vs bob :", float(a1 @ b1)) # typically 0.00 - 0.30Normalising to unit length is not cosmetic. Once every embedding has norm 1, the dot product is the cosine similarity, and distances become comparable across clips of different loudness and length. Skip it and your similarity scores drift with recording volume.
Zero-shot cloning
With an encoder in hand, cloning is assembly rather than training:
reference audio -> speaker encoder -> e (192-dim, unit norm) |target text -> text encoder -> content representation | both conditioned v acoustic model + vocoder v new speech, target voiceNo gradient step is taken. The TTS model was trained once across thousands of speakers to accept an identity vector as a conditioning input, so a new speaker is simply a vector it has not seen before — interpolated, in effect, from the region of voice-space it already knows.
1# pip install coqui-tts (the maintained fork; Coqui the company closed in 2024)2from TTS.api import TTS34tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2").to("cuda")56tts.tts_to_file(7 text="Your appointment has been moved to Thursday at ten.",8 speaker_wav="reference_30s.wav", # clean, single speaker, no music9 language="en",10 file_path="cloned.wav",11)The toolkit is open source, but the XTTS v2 weights are released under the Coqui Public Model License, which does not allow commercial use — check the licence of any cloning model before you build a product on it.
How much reference audio you actually need
Quality rises steeply and then flattens. The figures below are illustrative — the exact values depend on the cloning model and on which speaker encoder does the measuring — for cosine similarity between the reference embedding and an embedding of the generated audio:
| Reference length | Typical similarity | What it sounds like |
|---|---|---|
| 3 s | ~0.65 | Right gender and rough pitch; not identifiable as the person |
| 6 s | ~0.72 | Recognisable to someone who knows them |
| 10 s | ~0.78 | Convincing on a phone line |
| 30 s | ~0.85 | Convincing to most listeners; accent and rhythm carry |
| 60 s | ~0.87 | Marginal gain over 30 s |
| 5 min | ~0.88 | Effectively the zero-shot ceiling |
The important shape here is the flattening. Going from 3 s to 30 s buys 0.20 of similarity; going from 30 s to 5 minutes buys 0.03. Zero-shot cloning has a ceiling determined by what the conditioning vector can express, and no amount of extra reference audio pushes past it.
Audio quality matters far more than audio quantity below that ceiling. Ten seconds of clean, close-microphone, single-speaker speech beats two minutes of a video call with room echo and a second person interrupting. The encoder was trained on speech; background music, reverb and overlapping speakers all get partly absorbed into the embedding as if they were characteristics of the voice, and then the clone inherits them.
Few-shot cloning: when zero-shot is not enough
Fine-tuning the whole TTS model on a specific speaker breaks through the zero-shot ceiling, at real cost.
| Zero-shot | Few-shot (fine-tuned) | |
|---|---|---|
| Reference audio needed | 6–30 seconds | 10 minutes to several hours |
| Setup time | ~300 ms per clone | 30 minutes to 12 GPU-hours |
| Typical similarity | 0.80–0.88 | 0.90–0.95 |
| Prosody and idiolect | Generic; timbre only | Captures habitual rhythm, filler patterns, laugh |
| Storage per voice | 2 KB (one vector) | Hundreds of MB to GB (model weights or adapter) |
| Scales to | Millions of voices | Dozens, realistically |
| Right for | User-facing personalisation, one-off use | An audiobook narrator, a brand voice, accessibility |
The storage row is the one that decides architecture. If your product lets every user clone their own voice, few-shot is impossible — a million users at 400 MB each is 400 TB and a million fine-tuning jobs. If your product has four brand voices, few-shot is obviously correct, and the extra 0.07 of similarity is exactly what separates "an impression of them" from "them".
Parameter-efficient fine-tuning sits in between: freeze the base model and train a small adapter (a few MB) per speaker. You keep most of the quality gain and the per-voice storage stays manageable.
One failure mode dominates fine-tuning: catastrophic forgetting. Train too long on one speaker's ten minutes and the model stops being a general TTS system. It gets better at that speaker's typical sentences and worse at everything else — unusual words degrade, other languages break, prosody collapses to whatever was in the ten minutes. The mitigations are the usual ones: low learning rate, few epochs, a held-out set of general sentences you evaluate on every epoch, and stopping when general quality starts to fall even though speaker similarity is still rising.
Style transfer and voice morphing
Identity is not the only thing you might want to copy. Style transfer conditions generation on a second reference clip that supplies delivery — energy, pace, emotional colour — while identity comes from the first. Architecturally this means two conditioning vectors instead of one, from two separately-trained encoders: a speaker encoder trained to be style-invariant, and a style encoder trained to be speaker-invariant. Getting that disentanglement right is the hard part, and imperfect disentanglement is why "say it in an angry voice" sometimes also shifts who it sounds like.
Voice morphing interpolates between two identity vectors to produce a voice that is neither person — a legitimate technique for generating synthetic voices that do not belong to anybody real.
There is one piece of maths people get wrong here. Take unit-norm embeddings eA and eB with cosine similarity 0.3, and take their midpoint:
The midpoint has norm 0.806, not 1. It has fallen off the unit sphere the model was trained on, into a lower-magnitude region it has essentially never seen, and the usual symptom is a thin, under-energised voice. Renormalising to unit length fixes it — or use spherical interpolation, which travels along the sphere's surface and keeps the norm correct throughout.
1import numpy as np23def slerp(a, b, alpha):4 """Spherical interpolation between two unit-norm embeddings."""5 a = a / np.linalg.norm(a)6 b = b / np.linalg.norm(b)7 dot = float(np.clip(a @ b, -1.0, 1.0))8 theta = np.arccos(dot)9 if theta < 1e-6: # nearly identical: linear is fine10 return a11 s = np.sin(theta)12 return (np.sin((1 - alpha) * theta) / s) * a + (np.sin(alpha * theta) / s) * b1314blended = slerp(embed_a, embed_b, 0.5) # norm stays 1.0Interpolating naively between two unit vectors takes you off the sphere the model was trained on; renormalise, or interpolate along the sphere.
Measuring a clone
Four axes, and they trade against each other.
| Axis | Metric | Good value | Failure it catches |
|---|---|---|---|
| Speaker similarity | Cosine between reference and output embeddings | > 0.80 | Sounds like a different person |
| Naturalness | Human MOS, or predicted MOS | > 4.0 | Right voice, robotic delivery |
| Intelligibility | WER from ASR over the output | < 5% | Skipped or mangled words |
| Prosody transfer | Pitch-contour correlation with reference style | > 0.7 | Flat, affectless delivery |
There is a trap in the first row that catches almost everyone. If you condition the TTS model on embeddings from encoder E, and then evaluate similarity using encoder E, you are measuring how well the model reproduced the thing you handed it — not whether a human would recognise the voice. The metric is circular and it will read high while the audio sounds wrong. Always evaluate with a different speaker encoder from the one used for conditioning, and confirm with human similarity ratings on a sample.
The second trap: high similarity with low naturalness is common and often worse than the reverse. A clone that nails the timbre but delivers every sentence with identical flat intonation reads to listeners as uncanny rather than impressive, and uncanny is a worse product outcome than "clearly synthetic".
The ethics are the engineering
Voice cloning is not a technology with an ethics appendix. The capability and the harm are the same capability, and the mitigations have to be built into the system, not written into a policy document.
What has actually happened
These are documented cases, not hypotheticals.
In 2019, the chief executive of a UK energy firm was called by someone with the voice, accent and cadence of his German parent company's chief executive, and instructed to make an urgent transfer of €220,000. He did. The insurer that handled the claim attributed it to voice-synthesis software.
In 2023, Jennifer DeStefano testified before the US Senate Judiciary Committee about receiving a call in which her fifteen-year-old daughter's voice sobbed and pleaded while a man demanded a ransom. Her daughter was safe on a school trip. The voice had been cloned from social media audio.
In 2024, following robocalls that used a synthesised version of a sitting president's voice to discourage voting in a primary election, the US Federal Communications Commission ruled that AI-generated voices in robocalls are illegal under existing telephone consumer protection law.
Also in 2023, a journalist demonstrated that a clone of their own voice, built from publicly available recordings, could pass a major bank's voice-authentication system and reach their account balance.
Note the pattern. Three of the four are not attacks on the technology; they are attacks that use it against systems built on the assumption that a voice proves identity. That assumption is now false, and any system still relying on it is the vulnerability.
A voice is a biometric you cannot rotate. When a password leaks you change it; when your voice is cloned you have no equivalent move for the rest of your life.
Consent that means something
"They agreed" is not consent for a voice. Real consent for cloning has five properties, and a system that omits any of them will eventually produce an incident.
| Property | What it requires | What its absence looks like |
|---|---|---|
| Specific | Consent to voice cloning by name, not buried in general terms | A clause in a 40-page ToS nobody read |
| Scoped | Named permitted uses; everything else prohibited | A voice recorded for an advert appears in a political message |
| Time-bounded | An expiry date, renewed deliberately | A voice actor's clone still shipping ten years later |
| Revocable | Withdrawal deletes the embedding and stops generation | "We cannot un-train the model" as a permanent excuse |
| Verified | Proof that the consenter is the voice owner | Anyone uploading anyone's voice |
The verification row is the one most systems skip and the one that turns a product into a fraud tool. If a user can upload arbitrary audio and receive a clone, you have built a service whose primary competitive advantage is impersonation. The standard mitigation is a spoken consent phrase: the enrolling person must record a specific, freshly-generated sentence (containing a nonce — a random word or number issued at enrolment time), which cannot be satisfied by any pre-existing recording of the target.
The revocability row has real teeth in law. In many jurisdictions a voiceprint used to identify a person is biometric data: special-category personal data under the GDPR, requiring explicit consent and carrying deletion rights. Illinois' Biometric Information Privacy Act attaches statutory damages per violation, and Tennessee's ELVIS Act (2024) extended right-of-publicity protection explicitly to voice. The EU AI Act's transparency rules (Article 50, applying from 2 August 2026) require providers to mark synthetic audio as AI-generated in a machine-readable way, and require whoever publishes a deepfake of a real person to disclose it. The engineering consequence is concrete: you need per-voice deletion that actually works, which means storing embeddings in a way you can delete, and it means preferring zero-shot embeddings over fine-tuned weights whenever the quality difference permits — because you can delete a 2 KB vector, and you cannot cleanly remove a speaker from a fine-tuned checkpoint.
Logging, watermarking, provenance
Three technical controls, each covering something the others do not.
Usage logging answers "what was generated, for whom, when". Every generation should record the voice identifier, the consent record it relied on, the text, the requesting account and a hash of the output. This is what lets you answer an abuse report, and what lets a voice owner audit how their voice was used. Without it you cannot even establish whether a disputed clip came from your system.
Watermarking embeds an imperceptible signal in generated audio that survives compression, resampling and re-recording, and can be detected later. Modern neural watermarks (Meta's AudioSeal and similar) can say which parts of a clip are watermarked down to the individual sample, and detect reliably at very low false-positive rates through MP3 encoding and moderate noise. Watermarking is not a defence against a determined adversary running their own model; it is a defence against your output being used to deceive, and it lets platforms label content at scale.
Provenance metadata — C2PA-style signed manifests attached to the file — is cryptographically strong but trivially stripped, since removing metadata is a one-line operation. Watermarks survive stripping; provenance survives forgery. Use both.
1import hashlib, json, datetime, uuid23def record_generation(consent_id, voice_id, text, audio_bytes, account_id):4 """Write one immutable audit row per generation. Store the hash, not the audio."""5 return {6 "generation_id": str(uuid.uuid4()),7 "timestamp": datetime.datetime.now(datetime.timezone.utc).isoformat(),8 "consent_id": consent_id, # must be active and unexpired9 "voice_id": voice_id,10 "account_id": account_id,11 "text_hash": hashlib.sha256(text.encode()).hexdigest(),12 "audio_hash": hashlib.sha256(audio_bytes).hexdigest(),13 "watermarked": True,14 }1516def check_consent(consent, requested_use, now):17 if consent["revoked_at"] is not None:18 raise PermissionError("consent revoked")19 if now > consent["expires_at"]:20 raise PermissionError("consent expired")21 if requested_use not in consent["permitted_uses"]:22 raise PermissionError(f"use '{requested_use}' not permitted")23 if not consent["identity_verified"]:24 raise PermissionError("voice ownership not verified")25 return TrueStoring the hash rather than the audio is deliberate. It proves that a disputed clip did or did not come from your system, without your building a searchable archive of everything anyone ever synthesised — which would itself be a serious liability.
Building with this
Two design decisions carry almost all the risk, and both are made early.
The first is where the reference audio comes from. A system that clones from audio the user records live, with a nonce phrase, in a session tied to a verified account, is fundamentally different from a system that clones from an uploaded file. Same model, same quality, entirely different threat profile. If you accept uploads, you have accepted that your service will be used to impersonate people, and you need detection, human review and rate limiting as first-class components rather than later additions.
The second is zero-shot versus fine-tuned, and the deciding factor is usually not quality. A 2 KB embedding can be deleted on request in milliseconds, is cheap to store per user, and scales to millions of voices. Fine-tuned weights give you a genuine quality gain and give you a deletion problem that has no clean solution. Choose fine-tuning when the voice count is small and the relationship with the speaker is contractual — a narrator, a brand voice. Choose zero-shot when users bring their own voices, and accept the ceiling.
Then assume your output will be used to deceive someone, and build accordingly: watermark every generation, log every generation against a live consent record, enforce expiry rather than trusting a spreadsheet, and refuse to synthesise a voice you cannot prove the requester is entitled to use. These are not costs imposed by regulation. They are the parts of the system that let you keep operating after the first time someone tries to misuse it — and someone will.