Synthetic Data Generation

Data Augmentation Strategies


A team had 20,000 labelled product reviews and a sentiment classifier scoring 0.883 on a real held-out set. They wanted more data, so they applied a standard recipe: four augmented copies of each review, each produced by synonym replacement, random insertion, random swap and random deletion at probability 0.1 per token. Eighty thousand extra rows for the price of a CPU afternoon.

Accuracy fell to 0.861.

The cause was one operation. Random deletion at p = 0.1 on a 14-token review deletes any given token with probability 0.10. About 31 percent of reviews in this corpus contain a negation — not, never, no, n't. So the probability that an augmented copy loses its negation is 0.31 × 0.10 = 0.031. Across 80,000 augmented rows that is roughly 2,480 rows carrying a label that is now the opposite of what the text says. "The delivery was not late" became "The delivery was late", and kept its positive label.

Worse than the count is the placement. That 3.1 percent is not spread evenly — it lands entirely on reviews containing negation, which are the hardest and most linguistically informative examples in the set. The augmentation deleted signal precisely where the model needed it most.

Augmentation is a claim about invarianceWhere the invariance holds• Back-translating a product review• Mirroring a photo of a cat• Time-shifting a spoken command• Synonyms that preserve polarityWhere it silently fails• Random deletion removing "not"• Mirroring a road sign or a digit• Word swaps that break a negation• Mixup across incompatible labels
Four copies of every review multiply the data by four and the mislabelled negations by four as well, which is how a 0.883 baseline gets worse after augmentation.

Two different contracts

Augmentation and generation get grouped together as "making more data", but they make different promises, and confusing them is the source of most augmentation disasters.

AugmentationGeneration
Starts fromA specific real, labelled exampleA prompt, a distribution, or a simulator
The promiseThe label is unchanged under this transformThe row is drawn from a plausible distribution
What it addsRobustness — invariance to surface variationCoverage — regions of the space with no real examples
Cost per rowMicroseconds to milliseconds, CPU onlyMilliseconds to seconds, usually a paid API call
Main failureSilent label corruptionHallucination, mode collapse, distribution drift
Privacy statusDerived from one real record — no anonymity gainNot automatically anonymous either, but no one-to-one link
CeilingBounded by the diversity already in your real dataBounded by the generator's knowledge and your prompt design

The row that matters most is "what it adds". If your model fails because it has never seen a complaint about a subscription cancellation, no amount of paraphrasing the complaints you do have will help — augmentation cannot invent a topic. If your model fails because it breaks on lowercase input or a slightly rotated photo, generation is a wasteful way to fix something a transform handles for free.

Augmentation teaches a model what to ignore. Generation teaches it what exists. Diagnose which one you are missing before choosing a technique.

The invariance question you must answer first

Every augmentation is a claim: the label of this example is invariant under this transformation. The claim is sometimes true, often approximately true, and occasionally catastrophically false. It is domain-specific, and it is your job to check it, because the library will not.

TransformClaimed invarianceWhere the claim is false
Horizontal flipClass survives mirroringChest X-rays (heart position), text in images, left/right traffic signs, handedness tasks
Rotation ±30°Class survives orientation changeDigits (6 and 9), dermatology where lesion orientation to skin lines matters, document layout
Colour jitterClass survives hue and saturation shiftAnything where colour is the label: ripeness, corrosion, inflammation, wire coding
Random deletion (text)Sentiment survives dropping a wordNegation, quantifiers, "only", "except", any clause-scoping word
Synonym replacementMeaning survives a WordNet swapNamed entities, technical terms, idioms, polarity-carrying intensifiers
Back-translationMeaning survives a round tripSarcasm, register-dependent politeness, domain jargon, exact numbers
Speed perturbation (audio)Transcript survives tempo changeSpeaker-identification tasks, emotion recognition, prosody-dependent labels

The practical rule: for each transform you enable, hand-check 50 augmented examples against their inherited labels. Fifty examples takes twenty minutes and catches the negation bug before it costs you a week.

Text augmentation

Back-translation

Translate to a pivot language and back. The round trip forces a rephrasing that preserves meaning better than word-level tricks, because the intermediate representation is semantic.

The pivot language is the tuning knob, and the trade-off is direct. On 1,000 English reviews:

PivotOutput identical to inputChanged by 3+ tokensLabel flips (200 checked)
German41%34%1.8%
French38%37%2.0%
Japanese9%71%6.5%
English → German → French → English12%66%4.1%

A close pivot gives safe but weak augmentation — 41 percent of German round trips return the input unchanged, which is 410 exact duplicates you must then remove. A distant pivot gives strong variation and 3.6× the label corruption. Two-hop pivoting sits between them and is usually the best value. Always deduplicate after back-translation; the identical-output rate is the single biggest waste in the technique.

LLM paraphrasing and style rewriting

A language model can rewrite with explicit control that back-translation cannot offer: change the register, add typos, make it terser, make it angrier, keep every number identical.

Python
REWRITE = """Rewrite the customer message below.KEEP EXACTLY THE SAME: the problem being reported, the sentiment,every number, every product name, every date.CHANGE: {axis}Return only the rewritten message.Message: {text}"""AXES = ["make it two sentences shorter",        "rewrite as if typed quickly on a phone, with 2-3 typos",        "make the tone more formal and restrained",        "make the tone noticeably more frustrated, without new facts",        "rewrite in the first person plural, as a small business"]

The "KEEP EXACTLY THE SAME" block is not decoration. Without it, a rewrite asked to be "more frustrated" will invent a second problem, because escalation is what frustrated messages do — and your label, which describes one problem, becomes wrong. Verification is cheap: extract numbers and product names from both versions and assert set equality. On one run this check rejected 7.2 percent of rewrites, nearly all of them cases where the model had helpfully added a detail.

Synonym and entity replacement

The cheapest technique and the most dangerous. Two safeguards make it usable:

  • Protect a stop-list. Never touch negations, quantifiers (all, some, only, never), modal verbs, or any token inside a labelled span. Never delete tokens at all in a task where scoping words carry the label.
  • Replace entities from typed pools, not from a thesaurus. Swap a person name for another person name, a city for another city, an order ID for another order ID matching your format regex. This is genuinely useful — it stops the model memorising that "Acme Router" implies a technical intent — and it is safe because the entity type is preserved.

Entity replacement has a second benefit that is easy to miss: it breaks the one-to-one link between an augmented row and the real person in the original record, which slightly reduces (but does not eliminate) re-identification risk.

Template slot-filling

For narrow, structured domains — intent utterances, command phrasings, form-like queries — write templates with typed slots and enumerate.

Text
TEMPLATE  "I need to {verb} my {item} for order {order_id}"  verb      cancel | return | exchange | get a refund for  item      order | package | subscription | last purchase  order_id  drawn from the real ID format regex4 verbs x 4 items x N ids  =  16 distinct surface forms per template12 templates               =  192 surface forms

This gives perfect label control — the slot values determine the label — and terrible naturalness. Templated data has a distinctive uniformity that a discriminator spots instantly. Use it to guarantee coverage of every intent, then mix it at 10–20 percent with organic data rather than letting it dominate.

Image augmentation

Geometric and photometric transforms

Crops, flips, rotations, scaling, brightness, contrast, saturation, blur. These are the workhorses and the effect size is real: on a fine-grained classification task, going from no augmentation to random-resized-crop plus horizontal flip plus colour jitter typically moves top-1 accuracy by 3–6 points.

The magnitude question is what teams get wrong. Rotating MNIST digits by up to ±30° during training dropped accuracy on the standard upright test set from 0.991 to 0.984 while lifting accuracy on a rotated test set from 0.42 to 0.96. Neither number is "better" in the abstract — the question is whether your production images arrive rotated. Augmentation magnitude should be set by the variation you expect at inference, not by what looks impressive in a notebook.

Two structural traps. If your task has bounding boxes or masks, every geometric transform must be applied to the annotation as well — a rotated image with an unrotated box is a labelled error. And augmentation belongs in the training dataloader only. Augmenting before the train/test split leaks: a rotated copy of a test image sitting in the training set inflates your metric by several points and you will not find out until production.

Noise injection

Gaussian noise, salt-and-pepper, JPEG compression artefacts, sensor-style read noise. Effective for robustness when your real inputs are noisy, and pointless when they are not. Calibrate to measured conditions: if your production photos have a signal-to-noise ratio between 22 and 35 dB, augment across that band, not at 5 dB.

Mixup and CutMix

Mixup blends two examples and their labels linearly. Draw λ∼Beta(α,α)\lambda \sim \text{Beta}(\alpha, \alpha) and form:

Text
x~ = lambda * x_i + (1 - lambda) * x_jy~ = lambda * y_i + (1 - lambda) * y_jwith lambda = 0.7, y_i = [1, 0, 0] (cat), y_j = [0, 1, 0] (dog):  y~ = 0.7 * [1,0,0] + 0.3 * [0,1,0] = [0.7, 0.3, 0.0]loss = 0.7 * CrossEntropy(pred, cat) + 0.3 * CrossEntropy(pred, dog)

With the usual α = 0.2, the Beta density is U-shaped — most draws land near 0 or 1, giving mild blends, with occasional near-even mixes. CutMix does the same idea spatially: paste a rectangular patch from image B into image A and set the label weights to the pixel-area proportions, so a patch covering 30 percent of the image gives label weights 0.7 and 0.3.

These are unusual among augmentations in that they deliberately break the label-invariance contract — they produce examples that are genuinely both classes. That is the point: they smooth the decision boundary and improve calibration. They are also inappropriate for tasks where a blended input is meaningless, such as detection with hard box assignments or any task where the model must output a discrete parse.

Diffusion-based style transfer and inpainting

Image-to-image with a low strength setting (0.2–0.35) re-renders lighting, texture and background while preserving composition — so bounding boxes stay valid. Inpainting regenerates a masked region only, which is how you place a rare defect onto 400 photographs of a real product: the background remains genuinely real, the defect varies, and because you drew the mask you get the segmentation label for free.

This blurs the augmentation/generation line and inherits risks from both sides. Verify that the class is still present and correct in a sample — a diffusion pass asked to restyle a photograph of a cracked housing will sometimes helpfully repair the crack.

Audio augmentation

TechniqueWhat it simulatesTypical settingBreaks
Speed perturbationSpeaking rate variation0.9×, 1.0×, 1.1× — triples the corpusSpeaker ID, emotion, age estimation
SpecAugmentOcclusion and channel dropoutMask 2 frequency bands and 2 time spans on the spectrogramVery short utterances (wake words)
Room impulse responseReverberation of real roomsConvolve with measured RIRsTasks that classify the acoustic environment itself
Additive noiseBackground soundMix at SNR sampled from 10–20 dBBelow 5 dB, human transcribers disagree — so does your label

SpecAugment is notable because it operates on the spectrogram rather than the waveform, which makes it essentially free and lets it run inside the training loop with a different mask every epoch.

Combining augmentation with generation

The two techniques compose, and the order is not arbitrary.

Text
1. Split real data. Holdout is sealed before anything else happens.2. Diagnose the gap:      missing topics/classes   -> generation      brittleness to surface   -> augmentation3. Generate for coverage. Filter: schema, rules, dedup, judge.4. Merge real + accepted synthetic. Deduplicate across the union.5. Augment the merged training set inside the dataloader, at train time.6. Evaluate on the sealed real holdout. Never on augmented data.

Step 4's cross-set deduplication is routinely skipped and routinely bites: a generator seeded with few-shot examples from your real data will sometimes reproduce them almost verbatim, and you end up with the same example three times at three different weights.

Step 5 says at train time for a reason. Materialising augmented copies to disk fixes the augmentation, so the model sees the same 4 variants every epoch and memorises them. Applying transforms in the dataloader means a different variant each epoch — strictly more effective, and it costs no storage.

Measuring whether it helped

Diversity: before and after

MetricReal only+ augmentation (4×)+ generation
Distinct-20.700.740.71
Embedding dispersion0.710.720.79
Cluster coverage (k = 50)0.620.640.86

Read the coverage row. Quadrupling the dataset by augmentation moved cluster coverage by 0.02 — from 31 of 50 real behaviour clusters to 32. Augmentation produced four times the rows and essentially zero new semantic territory, exactly as the contract predicts. Generation moved it to 43 of 50. If someone claims augmentation fixed a coverage problem, this is the table that disproves it.

Downstream lift, with a significance test

The only number that decides anything. Train with and without, five seeds each, evaluate on the sealed real holdout:

Text
baseline  (1x):  macro-F1 = 0.742,  sd = 0.011  (5 seeds)augmented (2x):  macro-F1 = 0.769,  sd = 0.014  (5 seeds)pooled sd = sqrt((0.011^2 + 0.014^2) / 2) = sqrt(0.00015850) = 0.0126SE of difference = 0.0126 * sqrt(2/5) = 0.0126 * 0.6325 = 0.00796t = (0.769 - 0.742) / 0.00796 = 3.39,  df = 8,  p = 0.0095

A 0.027 gain with a standard deviation of 0.011 across seeds is real. A 0.008 gain with the same spread would not be, and reporting it as an improvement is how teams accumulate augmentation steps that do nothing but slow training down. Run the sweep:

Augmentation multiplierMacro-F1 (5-seed mean)Change vs baseline
1× (none)0.742—
2×0.769+0.027
4×0.771+0.029
8×0.764+0.022
16×0.751+0.009

The curve rises, plateaus at 4×, then declines. Beyond the plateau the augmented copies increasingly dominate the loss while carrying accumulated label noise, and training time grows linearly for nothing. There is always a plateau. Find it once, with a sweep, rather than assuming more is better.

Auditing for label drift

Train a classifier on the original data only, run it over the augmented copies, and compare its prediction with the inherited label. Overall agreement is uninformative; the breakdown is where the bug lives.

Text
agreement, all augmented rows                     97.1%agreement, rows containing a negation             89.4%agreement, rows from random-deletion only         92.6%agreement, negation AND random-deletion           78.2%   <-- the bug

That last cell names the failure exactly: random deletion applied to negation-bearing sentences. Disabling deletion for those rows recovered the accuracy the team had lost, and the audit took an hour.

Any augmentation that cannot survive a per-technique, per-slice agreement audit is a source of label noise wearing the costume of extra data.

What this means when you build

Make augmentation configurable per technique and per slice, not a single on/off switch. The negation failure could not be fixed by "less augmentation" — the right fix was disabling one operation for 31 percent of rows. A pipeline that only offers a global magnitude parameter forces you to weaken everything to fix one thing.

Sweep the multiplier once and record the plateau in your repository. It is a property of your dataset and it does not change often, but every new engineer will otherwise assume that 10× is better than 4× and quietly reintroduce the decline.

Keep the two contracts separate in your head and in your metrics. Report coverage numbers to justify generation and robustness numbers to justify augmentation. When a model is failing, the first question is which of the two is missing, and the answer is almost always visible in cluster coverage against real data: low coverage means you need new territory, and no amount of paraphrasing will supply it.