Synthetic Data Generation

Why Synthetic Data Matters


Priya's team at a payments company has 4.2 million card transactions from the last eighteen months. Every one of them is labelled. Of those 4.2 million, 812 are confirmed fraud. That is 0.019 percent — roughly one fraudulent transaction in every 5,170.

That sounds workable until you break the 812 down by fraud type. Card-testing attacks: 604 cases. Stolen-card purchases: 141. Rarer patterns: 44. Account-takeover-then-refund, the pattern that cost the company USD 380,000 in losses last quarter: 23 cases. Twenty-three. You cannot train a classifier on twenty-three examples of anything, and you certainly cannot validate one, because a held-out test set would contain about five of them and a single misclassification would swing your recall by twenty percentage points.

Priya has three options. Wait — at the current rate she will have roughly 30 more examples in eighteen months, and the attack pattern will have changed by then. Buy data from a partner — which means moving customers' card records across a corporate boundary, which means a data protection impact assessment, a legal review, and a contract that her company's counsel will spend four months negotiating. Or manufacture the examples: write something that produces transaction sequences which behave the way account-takeover-then-refund fraud behaves, in enough volume to train on.

That third option is synthetic data. It is not a trick and it is not free, and the rest of this is about knowing precisely when it is the right call.

The loop that eats itself812 realfraud rowsGeneratorSynthetic rowsModeltrained on bothIts outputbecomes datatails thin outvariancecollapsesEach pass resamples the centre of the previous pass and loses its rare cases.
Synthesis can only redistribute the information in those 812 rows, so every turn of the loop buys volume by spending diversity.

What synthetic data actually is

Synthetic data is data that was produced by a process you control, rather than recorded from a real-world event. No real customer placed the order. No real patient had the scan. The record exists because a generator — a statistical model, a simulator, a large language model, a rules engine — emitted it.

The word gets used loosely, and the loose usage causes expensive mistakes, particularly around privacy. These are five genuinely different things:

TermWhat it isLink to real people
Synthetic dataNew records emitted by a generator fitted to, or prompted about, a real distributionNo one-to-one link by construction — but leakage is possible and must be measured
Anonymised / de-identified dataReal records with identifiers stripped, masked, or generalisedEvery row is a real person; re-identification risk is direct
Augmented dataReal records transformed — rotated, paraphrased, noised — with the label preservedEvery row is derived from one specific real record
Simulated dataOutput of a mechanistic model of a process (physics engine, queueing model, market simulator)None — no real dataset was involved at all
Mock / fake dataRandom plausible-looking values from a library like FakerNone — and no statistical fidelity either

The distinction between the last two rows and the first matters more than it looks. Mock data has correct types — an email column contains things shaped like emails — but no structure. In mock data, order value and refund amount are independent random draws. In real data they correlate at 0.62. A model trained on mock data learns nothing, and worse, it appears to train fine: loss goes down, accuracy on the mock test set is 0.94, and the thing falls over in production. Synthetic data is only worth the name when it preserves the joint structure, not just the column types.

Mock data tests whether your code runs. Synthetic data tests whether your model learns. Confusing the two produces systems that pass every test and fail every customer.

The four pressures that push teams towards synthesis

Privacy and regulation

Under GDPR, personal data cannot be freely moved to a vendor, a test environment, or an offshore annotation team. Under HIPAA, a hospital cannot hand a research group its patient records without a limited-data-set agreement or de-identification to the Safe Harbor standard. Neither regime forbids the work; both add months.

The real cost is usually not the fine — it is the environment. Engineers who cannot get production-like data into a development environment write code against Faker output, and then discover in staging that 4 percent of real names contain characters their validation regex rejects.

The cost of annotation

Raw data is often cheap; labelled data almost never is. Suppose you need 50,000 support tickets labelled into five intent classes. A vendor annotator handles about 60 tickets an hour at a fully-loaded rate of USD 18 per hour, so USD 0.30 per ticket, so USD 15,000 for the batch. Add a 15 percent gold set double-labelled for agreement measurement and you are at roughly USD 17,250. With four annotators at 30 productive hours a week you get 7,200 tickets a week, so 50,000 takes 6.9 weeks of calendar time before anyone trains a model.

Scarcity

Some data does not exist in quantity at any price. A new drug has 40 patients in trial. A newly-launched product has three weeks of clickstream. A rare manufacturing defect appears twice per million units. No budget converts into examples that were never recorded.

Class imbalance

Priya's 0.019 percent is the canonical case. A classifier that predicts "not fraud" for every input scores 99.981 percent accuracy and catches zero fraud. Standard fixes — class weights, oversampling the minority, SMOTE-style interpolation — help, but they all reuse the same 812 records. Oversampling the account-takeover class 50 times does not create 50 kinds of account takeover; it creates 1,150 copies of 23 events, and the model memorises them.

Where synthetic data earns its keep

Use caseWhat you generateThe thing that makes it work
Training-set augmentationExtra examples in classes the model is weak onYou have enough real data to validate against; synthetic only fills gaps
Privacy-preserving sharingA statistically faithful stand-in for a sensitive tableMeasured leakage, not assumed leakage — see below
Rare-event simulationFraud patterns, equipment failures, edge-case driving scenesA causal or behavioural model of the event, not just resampling
Cold startSeed data for a product with no users yetYou replace it with real data as it arrives, and you track the mix ratio
Test and load dataRealistic volume with realistic distributions and edge casesIncludes deliberately nasty values: empty strings, four-byte emoji, 200-character names
Rebalancing skewed dataMinority-class examples that are genuinely different from each otherDiversity measurement — otherwise you have oversampling with extra steps

The trade-offs, priced out

Return to the 50,000 tickets. The synthetic route costs something too, and it is worth doing the arithmetic properly rather than waving at "cheaper".

Assume each generated row costs 850 input tokens (system prompt plus few-shot examples) and 260 output tokens. Assume a post-generation filtering yield of 62 percent — schema failures, duplicates, and quality rejections together throw away 38 percent — so to end with 50,000 you must generate 50,000 ÷ 0.62 = 80,645 rows.

Token prices change often, so treat these as example rates (typical of a capable mid-price model at the time of writing) and check your provider's pricing page before budgeting:

  • Input: 80,645 × 850 = 68.5 million tokens. At USD 3 per million: USD 205
  • Output: 80,645 × 260 = 21.0 million tokens. At USD 15 per million: USD 315
  • Automated quality judging (a second, cheaper model pass): USD 180
  • Engineering: three days of prompt design, schema work and filter tuning at USD 600 per day: USD 1,800
  • Human spot-review of a 3 percent sample — 1,500 rows at 40 per hour, 37.5 hours at USD 45 per hour: USD 1,690

Total: about USD 4,190, over roughly five working days. Against USD 17,250 and 6.9 weeks for annotation. But cost is only half the comparison. Here is what the same team measured on a held-out set of real tickets:

Training setMacro-F1 on real held-out data
50,000 real, human-labelled0.867
50,000 synthetic only0.781
5,000 real only0.742
5,000 real + 45,000 synthetic0.858

Read that table carefully, because it contains the whole argument. Synthetic-only lost 0.086 F1 against real data — a real and permanent gap. But the mixed row is the interesting one. Starting from 5,000 real tickets, adding synthetic data closed 0.116 of the available 0.125 gap, which is 93 percent of the achievable gain, for USD 4,190 instead of the roughly USD 15,500 it would have cost to label the other 45,000 tickets. Twenty-seven percent of the cost for ninety-three percent of the gain — and none of it available if you had zero real tickets, because without real data you cannot even build the test set that produced this table.

Synthetic data is an amplifier of the real data you already have, not a substitute for having any.

Synthetic is not automatically anonymous

This is the single most commonly repeated falsehood about synthetic data, and it appears in vendor marketing constantly: "synthetic data contains no personal information, therefore GDPR does not apply." That claim is not true by construction. It is a hypothesis that must be tested per dataset.

The reason is memorisation. A generator fitted to 10,000 patient records has, in its parameters, information about those 10,000 records. If it is expressive enough and trained long enough, it will reproduce some of them.

Distance to closest record

The first check is blunt and effective. For every synthetic row, compute the distance to its nearest neighbour in the real training set — the distance to closest record, DCR. Then compute the same thing against a real holdout set the generator never saw.

Text
Synthetic row --> nearest TRAINING row:   5th percentile DCR = 0.014Synthetic row --> nearest HOLDOUT row:    5th percentile DCR = 0.041Ratio = 0.014 / 0.041 = 0.34Exact duplicates (DCR = 0) = 312 of 10,000 rows = 3.1%

If the generator had learned the distribution, synthetic rows would sit about as close to training rows as to holdout rows and the ratio would be near 1.0. A ratio of 0.34 says the generator is hugging its training set. And 312 rows are verbatim copies of real patients' records, republished under the label "synthetic".

Membership inference

The sharper attack: an adversary holds one candidate record and wants to know whether that person was in the training set. In a medical context, mere membership is disclosure — being in the oncology cohort tells you the person has cancer.

The attack works by exploiting the same memorisation. Records that were in training tend to sit in denser regions of the synthetic data than records that were not. The attacker scores each candidate by, say, its distance to the nearest synthetic row, and thresholds. Random guessing gives AUC 0.50. If an attack against an unprotected tabular generator reaches AUC 0.71, the attacker is meaningfully better than chance, and if they hold 1,000 candidates of whom 100 were in training, they can identify a subset where their precision is far above the 10 percent base rate.

Defences exist — differential privacy in training, which bounds the influence of any single record and gives a formal guarantee, at a measurable cost in fidelity — but "we generated it, so it is anonymous" is not one of them.

Treat the privacy of a synthetic dataset as a measured property with a number attached, exactly as you would treat its accuracy.

The feedback loop that eats itself

There is a second failure that only shows up over time: train a generator on real data, generate, train the next generator partly on that output, and repeat. Each round loses information, and the loss compounds.

The mechanism is easiest to see with a single number. Suppose the real quantity is normally distributed with variance 1.0. Each generation, you draw n = 1,000 samples from the current model and fit the next model by maximum likelihood. The MLE variance estimator has expectation n−1nσ2\frac{n-1}{n}\sigma^2, so each round the expected variance shrinks by a factor of 0.999:

GenerationExpected varianceWhat it means
01.000Real data
1000.905Barely noticeable
1,0000.368Distribution visibly narrower
5,0000.007Collapsed to a point

The tails go first, and they go much faster than that table suggests. An event with probability 0.001 appears in a 1,000-sample draw with probability 1−0.9991000=0.6321 - 0.999^{1000} = 0.632 — so about 37 percent of the time it does not appear at all, and once it is missing from one generation's training data, the next generator assigns it probability zero and it never returns. Rare classes, minority dialects, unusual clinical presentations: exactly the things you built synthetic data to preserve are the first things it destroys.

The practical rule that follows: keep a fixed anchor of real data in every training mix, and track the synthetic fraction as a first-class number in your dataset metadata.

A decision framework

Before generating anything, answer these five questions. If you cannot answer the first, do not start.

  1. Do I have a real held-out set to validate against? Even 500 real examples. Without them you have no way to know whether the synthetic data is any good, and every downstream number is self-referential.
  2. What specifically is scarce — volume, labels, or a particular slice? "We need more data" is not a spec. "We need 3,000 account-takeover sequences that differ from each other in the timing of the refund request" is.
  3. Do I understand the generating process well enough to encode it? If nobody on the team can describe what makes an account-takeover sequence look different from a legitimate refund, the generator will not discover it for you.
  4. What is the acceptance threshold? Decide in advance: total variation distance below 0.05 on key categoricals, correlation differences below 0.10, DCR ratio above 0.8, downstream F1 lift of at least 0.02. Deciding afterwards means deciding in favour of whatever you produced.
  5. Who signs off, and on what evidence? Someone must own the claim that this data is fit to train on and safe to share.

When not to reach for synthetic data

SituationWhy synthesis fails here
Zero real data in the domainYou are encoding your assumptions and then measuring them. The model learns your priors, not the world.
Regulated decisions on individuals — credit, hiring, diagnosisRegulators want provenance for the training data behind a decision affecting a person. Synthetic provenance is hard to defend and easy to challenge.
The real bottleneck is a modelling bugMore data will not fix a leaking feature, a broken tokeniser, or a label-shuffled join.
You need to estimate a population statisticThe synthetic mean is your generator's mean. Reporting it as the population mean is circular.
Real data is available and cheapA week of engineering costs more than a week of annotation for small volumes. Do the arithmetic before assuming.

What this means when you build

Treat a synthetic dataset as an engineering artefact with a spec, not as a pile of rows. Concretely, three habits separate teams that get value from this from teams that quietly poison their models.

Split your real data before you generate, not after. The moment a real record influences a prompt, a few-shot example, or a generator's weights, it is contaminated — it can no longer serve as a test example. Carve out the holdout first and put it somewhere the generation pipeline cannot read.

Ship the provenance with the data. Every row should carry which generator produced it, which seed and parameters, which real records seeded it, and when. When a model behaves strangely in six months, the question "which parts of this training set were synthetic?" needs an answer that takes thirty seconds, not three days.

Report the mixed number, never the synthetic-only number. The measurement that matters is performance on real held-out data. A model that scores 0.94 on synthetic test data and 0.71 on real test data has told you something important about your generator and nothing at all about your model.

For Priya, the honest answer is not "generate 3,000 account-takeover cases and train on them". It is: use the 23 real cases plus domain knowledge of the attack to build a sequence generator; produce 3,000 varied cases; train on those plus all 4.2 million real transactions; and evaluate strictly on real fraud that the generator never saw, accepting the result only if recall on those held-out real cases improves. If it does not, the generator encoded her assumptions rather than the attack, and the correct response is to fix the generator — not to lower the bar.