Voice and Speech AI

Course Content

Voice Cloning & Advanced Audio Synthesis


You have thirty seconds of a colleague's voice from a voicemail. You want a system that can say any new sentence in that voice. Take the obvious route: fine-tune a text-to-speech model on the recording.

It fails before it starts. Thirty seconds at 22,050 Hz with a 256-sample hop is about 2,580 mel frames of training signal. A modest TTS decoder has tens of millions of parameters. The model does not learn the voice; it memorises the clip, and asked to say anything new it produces noise, or it produces the original sentence regardless of what you typed. The classical answer was to record twenty hours in a studio and fine-tune for a day on a GPU — which is why, until about 2018, having a synthetic version of your voice was a thing that happened to celebrities and nobody else.

The trick that broke this open is a change in what you are learning. Do not learn the voice. Learn, once, across thousands of speakers, what makes voices different from each other — and then represent any particular voice as a short vector in that space. The TTS model is trained once, on many speakers, to be conditioned on such a vector. Cloning a new voice then requires no training at all: extract the vector from a few seconds of audio and pass it in.

That works. Six seconds of reference audio is now enough to produce a recognisable clone in roughly 300 milliseconds. It works so well that in 2019 a UK energy company's chief executive transferred €220,000 to a fraudulent account because the voice on the phone was, unmistakably, his German parent-company boss. Both facts are the same fact. This lesson covers the mechanism and it covers what the mechanism obliges you to do, because the two are not separable.

Thirty seconds of voicemail into any new sentenceReferenceclip, sixseconds or moreSpeaker encoder,trained on thousandsOne 192-dimspeaker embeddingTTS conditionedon that vectorNew text spokenin that voiceNo fine-tuning happens: the voice is a vector the synthesiser is conditioned on, computed in under a second.
Because cloning now costs seconds rather than a training run, consent and provenance are engineering requirements, not policy notes.

Speaker embeddings: compressing "who" out of audio

A speaker embedding is a fixed-length vector — commonly 192 or 512 dimensions — that captures speaker identity and discards content. Feed it audio of Alice saying "hello" and audio of Alice reading a legal contract, and you should get nearly the same vector. Feed it Bob saying "hello" and you should get a distant one.

The compression is severe and that severity is the point. A 30-second reference clip at 16 kHz in float32 is 1,920,000 bytes. A 512-dimensional float32 embedding is 2,048 bytes. That is a 937× reduction, and everything that survives it is, by construction, speaker identity: vocal tract length, glottal characteristics, habitual pitch range, accent, speaking-rate tendencies. The words are gone.

How the encoder learns to do this

You cannot supervise this directly, because there is no ground-truth "identity vector" for a person. Instead you train with a metric-learning objective: sample audio clips, and push the network to make two clips from the same speaker land close together while two clips from different speakers land far apart.

The GE2E (Generalised End-to-End) loss does this in batches structured as NN speakers × MM utterances each. For every utterance, compute its similarity to the centroid of its own speaker and to the centroids of all other speakers, then apply a softmax that rewards the correct centroid. Training on thousands of speakers — VoxCeleb2 has 6,112 — forces the network to discover the axes along which human voices actually vary, rather than memorising any particular person.

MethodEraHow it worksDimensionsLimitation
i-vector2011Factor analysis over GMM supervectors; entirely statistical400–600Needs long utterances; degrades badly under noise
d-vector2014Average the hidden layer of a speaker-classification DNN256Trained for classification, not similarity
x-vector2018TDNN with statistics pooling over time512Strong baseline; superseded on accuracy
GE2E / ECAPA-TDNN2018–2020Metric learning with attentive pooling and channel attention192–256Current default; needs large multi-speaker data

The jump from d-vector to GE2E is a lesson in objective design. A d-vector network is trained to answer "which of these 5,000 speakers is this?" — so it learns features that separate those 5,000, with no pressure to generalise to a speaker it has never seen. GE2E is trained to answer "are these two clips the same person?", which is exactly the question you ask at cloning time. Same data, same architecture, radically different generalisation, because the training question matched the deployment question.

Train for the question you will actually ask at inference time. A network trained to classify known speakers will not reliably embed unknown ones.

Python
import torchfrom speechbrain.inference import EncoderClassifierencoder = EncoderClassifier.from_hparams(    source="speechbrain/spkrec-ecapa-voxceleb", savedir="/tmp/ecapa")def embed(wav_path):    signal = encoder.load_audio(wav_path)          # expects 16 kHz mono    with torch.no_grad():        e = encoder.encode_batch(signal.unsqueeze(0)).squeeze()    return torch.nn.functional.normalize(e, dim=0)  # unit norm: cosine == dota1, a2, b1 = embed("alice_1.wav"), embed("alice_2.wav"), embed("bob_1.wav")print("alice vs alice:", float(a1 @ a2))   # typically 0.75 - 0.90print("alice vs bob  :", float(a1 @ b1))   # typically 0.00 - 0.30

Normalising to unit length is not cosmetic. Once every embedding has norm 1, the dot product is the cosine similarity, and distances become comparable across clips of different loudness and length. Skip it and your similarity scores drift with recording volume.

Zero-shot cloning

With an encoder in hand, cloning is assembly rather than training:

Text
reference audio  ->  speaker encoder  ->  e  (192-dim, unit norm)                                           |target text      ->  text encoder    ->  content representation                                           |                                    both conditioned                                           v                             acoustic model + vocoder                                           v                              new speech, target voice

No gradient step is taken. The TTS model was trained once across thousands of speakers to accept an identity vector as a conditioning input, so a new speaker is simply a vector it has not seen before — interpolated, in effect, from the region of voice-space it already knows.

Python
# pip install coqui-tts  (the maintained fork; Coqui the company closed in 2024)from TTS.api import TTStts = TTS("tts_models/multilingual/multi-dataset/xtts_v2").to("cuda")tts.tts_to_file(    text="Your appointment has been moved to Thursday at ten.",    speaker_wav="reference_30s.wav",   # clean, single speaker, no music    language="en",    file_path="cloned.wav",)

The toolkit is open source, but the XTTS v2 weights are released under the Coqui Public Model License, which does not allow commercial use — check the licence of any cloning model before you build a product on it.

How much reference audio you actually need

Quality rises steeply and then flattens. The figures below are illustrative — the exact values depend on the cloning model and on which speaker encoder does the measuring — for cosine similarity between the reference embedding and an embedding of the generated audio:

Reference lengthTypical similarityWhat it sounds like
3 s~0.65Right gender and rough pitch; not identifiable as the person
6 s~0.72Recognisable to someone who knows them
10 s~0.78Convincing on a phone line
30 s~0.85Convincing to most listeners; accent and rhythm carry
60 s~0.87Marginal gain over 30 s
5 min~0.88Effectively the zero-shot ceiling

The important shape here is the flattening. Going from 3 s to 30 s buys 0.20 of similarity; going from 30 s to 5 minutes buys 0.03. Zero-shot cloning has a ceiling determined by what the conditioning vector can express, and no amount of extra reference audio pushes past it.

Audio quality matters far more than audio quantity below that ceiling. Ten seconds of clean, close-microphone, single-speaker speech beats two minutes of a video call with room echo and a second person interrupting. The encoder was trained on speech; background music, reverb and overlapping speakers all get partly absorbed into the embedding as if they were characteristics of the voice, and then the clone inherits them.

Few-shot cloning: when zero-shot is not enough

Fine-tuning the whole TTS model on a specific speaker breaks through the zero-shot ceiling, at real cost.

Zero-shotFew-shot (fine-tuned)
Reference audio needed6–30 seconds10 minutes to several hours
Setup time~300 ms per clone30 minutes to 12 GPU-hours
Typical similarity0.80–0.880.90–0.95
Prosody and idiolectGeneric; timbre onlyCaptures habitual rhythm, filler patterns, laugh
Storage per voice2 KB (one vector)Hundreds of MB to GB (model weights or adapter)
Scales toMillions of voicesDozens, realistically
Right forUser-facing personalisation, one-off useAn audiobook narrator, a brand voice, accessibility

The storage row is the one that decides architecture. If your product lets every user clone their own voice, few-shot is impossible — a million users at 400 MB each is 400 TB and a million fine-tuning jobs. If your product has four brand voices, few-shot is obviously correct, and the extra 0.07 of similarity is exactly what separates "an impression of them" from "them".

Parameter-efficient fine-tuning sits in between: freeze the base model and train a small adapter (a few MB) per speaker. You keep most of the quality gain and the per-voice storage stays manageable.

One failure mode dominates fine-tuning: catastrophic forgetting. Train too long on one speaker's ten minutes and the model stops being a general TTS system. It gets better at that speaker's typical sentences and worse at everything else — unusual words degrade, other languages break, prosody collapses to whatever was in the ten minutes. The mitigations are the usual ones: low learning rate, few epochs, a held-out set of general sentences you evaluate on every epoch, and stopping when general quality starts to fall even though speaker similarity is still rising.

Style transfer and voice morphing

Identity is not the only thing you might want to copy. Style transfer conditions generation on a second reference clip that supplies delivery — energy, pace, emotional colour — while identity comes from the first. Architecturally this means two conditioning vectors instead of one, from two separately-trained encoders: a speaker encoder trained to be style-invariant, and a style encoder trained to be speaker-invariant. Getting that disentanglement right is the hard part, and imperfect disentanglement is why "say it in an angry voice" sometimes also shifts who it sounds like.

Voice morphing interpolates between two identity vectors to produce a voice that is neither person — a legitimate technique for generating synthetic voices that do not belong to anybody real.

There is one piece of maths people get wrong here. Take unit-norm embeddings eAe_A and eBe_B with cosine similarity 0.3, and take their midpoint:

∥0.5 eA+0.5 eB∥=0.5∥eA∥2+2 eA ⁣⋅ ⁣eB+∥eB∥2=0.51+0.6+1=0.52.6=0.806\lVert 0.5\,e_A + 0.5\,e_B \rVert = 0.5\sqrt{\lVert e_A\rVert^2 + 2\,e_A\!\cdot\!e_B + \lVert e_B\rVert^2} = 0.5\sqrt{1 + 0.6 + 1} = 0.5\sqrt{2.6} = 0.806

The midpoint has norm 0.806, not 1. It has fallen off the unit sphere the model was trained on, into a lower-magnitude region it has essentially never seen, and the usual symptom is a thin, under-energised voice. Renormalising to unit length fixes it — or use spherical interpolation, which travels along the sphere's surface and keeps the norm correct throughout.

Python
import numpy as npdef slerp(a, b, alpha):    """Spherical interpolation between two unit-norm embeddings."""    a = a / np.linalg.norm(a)    b = b / np.linalg.norm(b)    dot = float(np.clip(a @ b, -1.0, 1.0))    theta = np.arccos(dot)    if theta < 1e-6:                      # nearly identical: linear is fine        return a    s = np.sin(theta)    return (np.sin((1 - alpha) * theta) / s) * a + (np.sin(alpha * theta) / s) * bblended = slerp(embed_a, embed_b, 0.5)    # norm stays 1.0

Interpolating naively between two unit vectors takes you off the sphere the model was trained on; renormalise, or interpolate along the sphere.

Measuring a clone

Four axes, and they trade against each other.

AxisMetricGood valueFailure it catches
Speaker similarityCosine between reference and output embeddings> 0.80Sounds like a different person
NaturalnessHuman MOS, or predicted MOS> 4.0Right voice, robotic delivery
IntelligibilityWER from ASR over the output< 5%Skipped or mangled words
Prosody transferPitch-contour correlation with reference style> 0.7Flat, affectless delivery

There is a trap in the first row that catches almost everyone. If you condition the TTS model on embeddings from encoder EE, and then evaluate similarity using encoder EE, you are measuring how well the model reproduced the thing you handed it — not whether a human would recognise the voice. The metric is circular and it will read high while the audio sounds wrong. Always evaluate with a different speaker encoder from the one used for conditioning, and confirm with human similarity ratings on a sample.

The second trap: high similarity with low naturalness is common and often worse than the reverse. A clone that nails the timbre but delivers every sentence with identical flat intonation reads to listeners as uncanny rather than impressive, and uncanny is a worse product outcome than "clearly synthetic".

The ethics are the engineering

Voice cloning is not a technology with an ethics appendix. The capability and the harm are the same capability, and the mitigations have to be built into the system, not written into a policy document.

What has actually happened

These are documented cases, not hypotheticals.

In 2019, the chief executive of a UK energy firm was called by someone with the voice, accent and cadence of his German parent company's chief executive, and instructed to make an urgent transfer of €220,000. He did. The insurer that handled the claim attributed it to voice-synthesis software.

In 2023, Jennifer DeStefano testified before the US Senate Judiciary Committee about receiving a call in which her fifteen-year-old daughter's voice sobbed and pleaded while a man demanded a ransom. Her daughter was safe on a school trip. The voice had been cloned from social media audio.

In 2024, following robocalls that used a synthesised version of a sitting president's voice to discourage voting in a primary election, the US Federal Communications Commission ruled that AI-generated voices in robocalls are illegal under existing telephone consumer protection law.

Also in 2023, a journalist demonstrated that a clone of their own voice, built from publicly available recordings, could pass a major bank's voice-authentication system and reach their account balance.

Note the pattern. Three of the four are not attacks on the technology; they are attacks that use it against systems built on the assumption that a voice proves identity. That assumption is now false, and any system still relying on it is the vulnerability.

A voice is a biometric you cannot rotate. When a password leaks you change it; when your voice is cloned you have no equivalent move for the rest of your life.

Consent that means something

"They agreed" is not consent for a voice. Real consent for cloning has five properties, and a system that omits any of them will eventually produce an incident.

PropertyWhat it requiresWhat its absence looks like
SpecificConsent to voice cloning by name, not buried in general termsA clause in a 40-page ToS nobody read
ScopedNamed permitted uses; everything else prohibitedA voice recorded for an advert appears in a political message
Time-boundedAn expiry date, renewed deliberatelyA voice actor's clone still shipping ten years later
RevocableWithdrawal deletes the embedding and stops generation"We cannot un-train the model" as a permanent excuse
VerifiedProof that the consenter is the voice ownerAnyone uploading anyone's voice

The verification row is the one most systems skip and the one that turns a product into a fraud tool. If a user can upload arbitrary audio and receive a clone, you have built a service whose primary competitive advantage is impersonation. The standard mitigation is a spoken consent phrase: the enrolling person must record a specific, freshly-generated sentence (containing a nonce — a random word or number issued at enrolment time), which cannot be satisfied by any pre-existing recording of the target.

The revocability row has real teeth in law. In many jurisdictions a voiceprint used to identify a person is biometric data: special-category personal data under the GDPR, requiring explicit consent and carrying deletion rights. Illinois' Biometric Information Privacy Act attaches statutory damages per violation, and Tennessee's ELVIS Act (2024) extended right-of-publicity protection explicitly to voice. The EU AI Act's transparency rules (Article 50, applying from 2 August 2026) require providers to mark synthetic audio as AI-generated in a machine-readable way, and require whoever publishes a deepfake of a real person to disclose it. The engineering consequence is concrete: you need per-voice deletion that actually works, which means storing embeddings in a way you can delete, and it means preferring zero-shot embeddings over fine-tuned weights whenever the quality difference permits — because you can delete a 2 KB vector, and you cannot cleanly remove a speaker from a fine-tuned checkpoint.

Logging, watermarking, provenance

Three technical controls, each covering something the others do not.

Usage logging answers "what was generated, for whom, when". Every generation should record the voice identifier, the consent record it relied on, the text, the requesting account and a hash of the output. This is what lets you answer an abuse report, and what lets a voice owner audit how their voice was used. Without it you cannot even establish whether a disputed clip came from your system.

Watermarking embeds an imperceptible signal in generated audio that survives compression, resampling and re-recording, and can be detected later. Modern neural watermarks (Meta's AudioSeal and similar) can say which parts of a clip are watermarked down to the individual sample, and detect reliably at very low false-positive rates through MP3 encoding and moderate noise. Watermarking is not a defence against a determined adversary running their own model; it is a defence against your output being used to deceive, and it lets platforms label content at scale.

Provenance metadata — C2PA-style signed manifests attached to the file — is cryptographically strong but trivially stripped, since removing metadata is a one-line operation. Watermarks survive stripping; provenance survives forgery. Use both.

Python
import hashlib, json, datetime, uuiddef record_generation(consent_id, voice_id, text, audio_bytes, account_id):    """Write one immutable audit row per generation. Store the hash, not the audio."""    return {        "generation_id": str(uuid.uuid4()),        "timestamp": datetime.datetime.now(datetime.timezone.utc).isoformat(),        "consent_id": consent_id,              # must be active and unexpired        "voice_id": voice_id,        "account_id": account_id,        "text_hash": hashlib.sha256(text.encode()).hexdigest(),        "audio_hash": hashlib.sha256(audio_bytes).hexdigest(),        "watermarked": True,    }def check_consent(consent, requested_use, now):    if consent["revoked_at"] is not None:        raise PermissionError("consent revoked")    if now > consent["expires_at"]:        raise PermissionError("consent expired")    if requested_use not in consent["permitted_uses"]:        raise PermissionError(f"use '{requested_use}' not permitted")    if not consent["identity_verified"]:        raise PermissionError("voice ownership not verified")    return True

Storing the hash rather than the audio is deliberate. It proves that a disputed clip did or did not come from your system, without your building a searchable archive of everything anyone ever synthesised — which would itself be a serious liability.

Building with this

Two design decisions carry almost all the risk, and both are made early.

The first is where the reference audio comes from. A system that clones from audio the user records live, with a nonce phrase, in a session tied to a verified account, is fundamentally different from a system that clones from an uploaded file. Same model, same quality, entirely different threat profile. If you accept uploads, you have accepted that your service will be used to impersonate people, and you need detection, human review and rate limiting as first-class components rather than later additions.

The second is zero-shot versus fine-tuned, and the deciding factor is usually not quality. A 2 KB embedding can be deleted on request in milliseconds, is cheap to store per user, and scales to millions of voices. Fine-tuned weights give you a genuine quality gain and give you a deletion problem that has no clean solution. Choose fine-tuning when the voice count is small and the relationship with the speaker is contractual — a narrator, a brand voice. Choose zero-shot when users bring their own voices, and accept the ceiling.

Then assume your output will be used to deceive someone, and build accordingly: watermark every generation, log every generation against a live consent record, enforce expiry rather than trusting a spreadsheet, and refuse to synthesise a voice you cannot prove the requester is entitled to use. These are not costs imposed by regulation. They are the parts of the system that let you keep operating after the first time someone tries to misuse it — and someone will.