Course Content
Synthetic Data Generation
3 sections · 7 lessons
Data Augmentation Strategies
A team had 20,000 labelled product reviews and a sentiment classifier scoring 0.883 on a real held-out set. They wanted more data, so they applied a standard recipe: four augmented copies of each review, each produced by synonym replacement, random insertion, random swap and random deletion at probability 0.1 per token. Eighty thousand extra rows for the price of a CPU afternoon.
Accuracy fell to 0.861.
The cause was one operation. Random deletion at p = 0.1 on a 14-token review deletes any given token with probability 0.10. About 31 percent of reviews in this corpus contain a negation — not, never, no, n't. So the probability that an augmented copy loses its negation is 0.31 × 0.10 = 0.031. Across 80,000 augmented rows that is roughly 2,480 rows carrying a label that is now the opposite of what the text says. "The delivery was not late" became "The delivery was late", and kept its positive label.
Worse than the count is the placement. That 3.1 percent is not spread evenly — it lands entirely on reviews containing negation, which are the hardest and most linguistically informative examples in the set. The augmentation deleted signal precisely where the model needed it most.
Two different contracts
Augmentation and generation get grouped together as "making more data", but they make different promises, and confusing them is the source of most augmentation disasters.
| Augmentation | Generation | |
|---|---|---|
| Starts from | A specific real, labelled example | A prompt, a distribution, or a simulator |
| The promise | The label is unchanged under this transform | The row is drawn from a plausible distribution |
| What it adds | Robustness — invariance to surface variation | Coverage — regions of the space with no real examples |
| Cost per row | Microseconds to milliseconds, CPU only | Milliseconds to seconds, usually a paid API call |
| Main failure | Silent label corruption | Hallucination, mode collapse, distribution drift |
| Privacy status | Derived from one real record — no anonymity gain | Not automatically anonymous either, but no one-to-one link |
| Ceiling | Bounded by the diversity already in your real data | Bounded by the generator's knowledge and your prompt design |
The row that matters most is "what it adds". If your model fails because it has never seen a complaint about a subscription cancellation, no amount of paraphrasing the complaints you do have will help — augmentation cannot invent a topic. If your model fails because it breaks on lowercase input or a slightly rotated photo, generation is a wasteful way to fix something a transform handles for free.
Augmentation teaches a model what to ignore. Generation teaches it what exists. Diagnose which one you are missing before choosing a technique.
The invariance question you must answer first
Every augmentation is a claim: the label of this example is invariant under this transformation. The claim is sometimes true, often approximately true, and occasionally catastrophically false. It is domain-specific, and it is your job to check it, because the library will not.
| Transform | Claimed invariance | Where the claim is false |
|---|---|---|
| Horizontal flip | Class survives mirroring | Chest X-rays (heart position), text in images, left/right traffic signs, handedness tasks |
| Rotation ±30° | Class survives orientation change | Digits (6 and 9), dermatology where lesion orientation to skin lines matters, document layout |
| Colour jitter | Class survives hue and saturation shift | Anything where colour is the label: ripeness, corrosion, inflammation, wire coding |
| Random deletion (text) | Sentiment survives dropping a word | Negation, quantifiers, "only", "except", any clause-scoping word |
| Synonym replacement | Meaning survives a WordNet swap | Named entities, technical terms, idioms, polarity-carrying intensifiers |
| Back-translation | Meaning survives a round trip | Sarcasm, register-dependent politeness, domain jargon, exact numbers |
| Speed perturbation (audio) | Transcript survives tempo change | Speaker-identification tasks, emotion recognition, prosody-dependent labels |
The practical rule: for each transform you enable, hand-check 50 augmented examples against their inherited labels. Fifty examples takes twenty minutes and catches the negation bug before it costs you a week.
Text augmentation
Back-translation
Translate to a pivot language and back. The round trip forces a rephrasing that preserves meaning better than word-level tricks, because the intermediate representation is semantic.
The pivot language is the tuning knob, and the trade-off is direct. On 1,000 English reviews:
| Pivot | Output identical to input | Changed by 3+ tokens | Label flips (200 checked) |
|---|---|---|---|
| German | 41% | 34% | 1.8% |
| French | 38% | 37% | 2.0% |
| Japanese | 9% | 71% | 6.5% |
| English → German → French → English | 12% | 66% | 4.1% |
A close pivot gives safe but weak augmentation — 41 percent of German round trips return the input unchanged, which is 410 exact duplicates you must then remove. A distant pivot gives strong variation and 3.6× the label corruption. Two-hop pivoting sits between them and is usually the best value. Always deduplicate after back-translation; the identical-output rate is the single biggest waste in the technique.
LLM paraphrasing and style rewriting
A language model can rewrite with explicit control that back-translation cannot offer: change the register, add typos, make it terser, make it angrier, keep every number identical.
1REWRITE = """Rewrite the customer message below.23KEEP EXACTLY THE SAME: the problem being reported, the sentiment,4every number, every product name, every date.5CHANGE: {axis}67Return only the rewritten message.89Message: {text}"""1011AXES = ["make it two sentences shorter",12 "rewrite as if typed quickly on a phone, with 2-3 typos",13 "make the tone more formal and restrained",14 "make the tone noticeably more frustrated, without new facts",15 "rewrite in the first person plural, as a small business"]The "KEEP EXACTLY THE SAME" block is not decoration. Without it, a rewrite asked to be "more frustrated" will invent a second problem, because escalation is what frustrated messages do — and your label, which describes one problem, becomes wrong. Verification is cheap: extract numbers and product names from both versions and assert set equality. On one run this check rejected 7.2 percent of rewrites, nearly all of them cases where the model had helpfully added a detail.
Synonym and entity replacement
The cheapest technique and the most dangerous. Two safeguards make it usable:
- Protect a stop-list. Never touch negations, quantifiers (all, some, only, never), modal verbs, or any token inside a labelled span. Never delete tokens at all in a task where scoping words carry the label.
- Replace entities from typed pools, not from a thesaurus. Swap a person name for another person name, a city for another city, an order ID for another order ID matching your format regex. This is genuinely useful — it stops the model memorising that "Acme Router" implies a technical intent — and it is safe because the entity type is preserved.
Entity replacement has a second benefit that is easy to miss: it breaks the one-to-one link between an augmented row and the real person in the original record, which slightly reduces (but does not eliminate) re-identification risk.
Template slot-filling
For narrow, structured domains — intent utterances, command phrasings, form-like queries — write templates with typed slots and enumerate.
TEMPLATE "I need to {verb} my {item} for order {order_id}" verb cancel | return | exchange | get a refund for item order | package | subscription | last purchase order_id drawn from the real ID format regex4 verbs x 4 items x N ids = 16 distinct surface forms per template12 templates = 192 surface formsThis gives perfect label control — the slot values determine the label — and terrible naturalness. Templated data has a distinctive uniformity that a discriminator spots instantly. Use it to guarantee coverage of every intent, then mix it at 10–20 percent with organic data rather than letting it dominate.
Image augmentation
Geometric and photometric transforms
Crops, flips, rotations, scaling, brightness, contrast, saturation, blur. These are the workhorses and the effect size is real: on a fine-grained classification task, going from no augmentation to random-resized-crop plus horizontal flip plus colour jitter typically moves top-1 accuracy by 3–6 points.
The magnitude question is what teams get wrong. Rotating MNIST digits by up to ±30° during training dropped accuracy on the standard upright test set from 0.991 to 0.984 while lifting accuracy on a rotated test set from 0.42 to 0.96. Neither number is "better" in the abstract — the question is whether your production images arrive rotated. Augmentation magnitude should be set by the variation you expect at inference, not by what looks impressive in a notebook.
Two structural traps. If your task has bounding boxes or masks, every geometric transform must be applied to the annotation as well — a rotated image with an unrotated box is a labelled error. And augmentation belongs in the training dataloader only. Augmenting before the train/test split leaks: a rotated copy of a test image sitting in the training set inflates your metric by several points and you will not find out until production.
Noise injection
Gaussian noise, salt-and-pepper, JPEG compression artefacts, sensor-style read noise. Effective for robustness when your real inputs are noisy, and pointless when they are not. Calibrate to measured conditions: if your production photos have a signal-to-noise ratio between 22 and 35 dB, augment across that band, not at 5 dB.
Mixup and CutMix
Mixup blends two examples and their labels linearly. Draw λ∼Beta(α,α) and form:
x~ = lambda * x_i + (1 - lambda) * x_jy~ = lambda * y_i + (1 - lambda) * y_jwith lambda = 0.7, y_i = [1, 0, 0] (cat), y_j = [0, 1, 0] (dog): y~ = 0.7 * [1,0,0] + 0.3 * [0,1,0] = [0.7, 0.3, 0.0]loss = 0.7 * CrossEntropy(pred, cat) + 0.3 * CrossEntropy(pred, dog)With the usual α = 0.2, the Beta density is U-shaped — most draws land near 0 or 1, giving mild blends, with occasional near-even mixes. CutMix does the same idea spatially: paste a rectangular patch from image B into image A and set the label weights to the pixel-area proportions, so a patch covering 30 percent of the image gives label weights 0.7 and 0.3.
These are unusual among augmentations in that they deliberately break the label-invariance contract — they produce examples that are genuinely both classes. That is the point: they smooth the decision boundary and improve calibration. They are also inappropriate for tasks where a blended input is meaningless, such as detection with hard box assignments or any task where the model must output a discrete parse.
Diffusion-based style transfer and inpainting
Image-to-image with a low strength setting (0.2–0.35) re-renders lighting, texture and background while preserving composition — so bounding boxes stay valid. Inpainting regenerates a masked region only, which is how you place a rare defect onto 400 photographs of a real product: the background remains genuinely real, the defect varies, and because you drew the mask you get the segmentation label for free.
This blurs the augmentation/generation line and inherits risks from both sides. Verify that the class is still present and correct in a sample — a diffusion pass asked to restyle a photograph of a cracked housing will sometimes helpfully repair the crack.
Audio augmentation
| Technique | What it simulates | Typical setting | Breaks |
|---|---|---|---|
| Speed perturbation | Speaking rate variation | 0.9×, 1.0×, 1.1× — triples the corpus | Speaker ID, emotion, age estimation |
| SpecAugment | Occlusion and channel dropout | Mask 2 frequency bands and 2 time spans on the spectrogram | Very short utterances (wake words) |
| Room impulse response | Reverberation of real rooms | Convolve with measured RIRs | Tasks that classify the acoustic environment itself |
| Additive noise | Background sound | Mix at SNR sampled from 10–20 dB | Below 5 dB, human transcribers disagree — so does your label |
SpecAugment is notable because it operates on the spectrogram rather than the waveform, which makes it essentially free and lets it run inside the training loop with a different mask every epoch.
Combining augmentation with generation
The two techniques compose, and the order is not arbitrary.
1. Split real data. Holdout is sealed before anything else happens.2. Diagnose the gap: missing topics/classes -> generation brittleness to surface -> augmentation3. Generate for coverage. Filter: schema, rules, dedup, judge.4. Merge real + accepted synthetic. Deduplicate across the union.5. Augment the merged training set inside the dataloader, at train time.6. Evaluate on the sealed real holdout. Never on augmented data.Step 4's cross-set deduplication is routinely skipped and routinely bites: a generator seeded with few-shot examples from your real data will sometimes reproduce them almost verbatim, and you end up with the same example three times at three different weights.
Step 5 says at train time for a reason. Materialising augmented copies to disk fixes the augmentation, so the model sees the same 4 variants every epoch and memorises them. Applying transforms in the dataloader means a different variant each epoch — strictly more effective, and it costs no storage.
Measuring whether it helped
Diversity: before and after
| Metric | Real only | + augmentation (4×) | + generation |
|---|---|---|---|
| Distinct-2 | 0.70 | 0.74 | 0.71 |
| Embedding dispersion | 0.71 | 0.72 | 0.79 |
| Cluster coverage (k = 50) | 0.62 | 0.64 | 0.86 |
Read the coverage row. Quadrupling the dataset by augmentation moved cluster coverage by 0.02 — from 31 of 50 real behaviour clusters to 32. Augmentation produced four times the rows and essentially zero new semantic territory, exactly as the contract predicts. Generation moved it to 43 of 50. If someone claims augmentation fixed a coverage problem, this is the table that disproves it.
Downstream lift, with a significance test
The only number that decides anything. Train with and without, five seeds each, evaluate on the sealed real holdout:
baseline (1x): macro-F1 = 0.742, sd = 0.011 (5 seeds)augmented (2x): macro-F1 = 0.769, sd = 0.014 (5 seeds)pooled sd = sqrt((0.011^2 + 0.014^2) / 2) = sqrt(0.00015850) = 0.0126SE of difference = 0.0126 * sqrt(2/5) = 0.0126 * 0.6325 = 0.00796t = (0.769 - 0.742) / 0.00796 = 3.39, df = 8, p = 0.0095A 0.027 gain with a standard deviation of 0.011 across seeds is real. A 0.008 gain with the same spread would not be, and reporting it as an improvement is how teams accumulate augmentation steps that do nothing but slow training down. Run the sweep:
| Augmentation multiplier | Macro-F1 (5-seed mean) | Change vs baseline |
|---|---|---|
| 1× (none) | 0.742 | — |
| 2× | 0.769 | +0.027 |
| 4× | 0.771 | +0.029 |
| 8× | 0.764 | +0.022 |
| 16× | 0.751 | +0.009 |
The curve rises, plateaus at 4×, then declines. Beyond the plateau the augmented copies increasingly dominate the loss while carrying accumulated label noise, and training time grows linearly for nothing. There is always a plateau. Find it once, with a sweep, rather than assuming more is better.
Auditing for label drift
Train a classifier on the original data only, run it over the augmented copies, and compare its prediction with the inherited label. Overall agreement is uninformative; the breakdown is where the bug lives.
agreement, all augmented rows 97.1%agreement, rows containing a negation 89.4%agreement, rows from random-deletion only 92.6%agreement, negation AND random-deletion 78.2% <-- the bugThat last cell names the failure exactly: random deletion applied to negation-bearing sentences. Disabling deletion for those rows recovered the accuracy the team had lost, and the audit took an hour.
Any augmentation that cannot survive a per-technique, per-slice agreement audit is a source of label noise wearing the costume of extra data.
What this means when you build
Make augmentation configurable per technique and per slice, not a single on/off switch. The negation failure could not be fixed by "less augmentation" — the right fix was disabling one operation for 31 percent of rows. A pipeline that only offers a global magnitude parameter forces you to weaken everything to fix one thing.
Sweep the multiplier once and record the plateau in your repository. It is a property of your dataset and it does not change often, but every new engineer will otherwise assume that 10× is better than 4× and quietly reintroduce the decline.
Keep the two contracts separate in your head and in your metrics. Report coverage numbers to justify generation and robustness numbers to justify augmentation. When a model is failing, the first question is which of the two is missing, and the answer is almost always visible in cluster coverage against real data: low coverage means you need new territory, and no amount of paraphrasing will supply it.