Course Content
Synthetic Data Generation
3 sections · 7 lessons
Quality and Bias Considerations
A team generated 40,000 synthetic e-commerce orders to test a refund-risk model. They eyeballed 30 rows. Order values looked plausible, dates were in range, customer names looked like names. Sign-off took ten minutes.
Six weeks later the model was flagging 31 percent of orders as refund risks in production against an expected 4 percent. The post-mortem found it in one line: in the real data, order value and refund amount correlate at 0.62 — expensive orders get refunded more, and for more. In the synthetic data the correlation was 0.11. The generator had learned both marginal distributions beautifully and the relationship between them not at all. Every single column passed inspection. The dataset was still useless, because the thing the model needed to learn lived in the joint distribution, and nobody had looked at the joint distribution.
Eyeballing rows tells you whether values are plausible. It cannot tell you whether the dataset is faithful. Those are different properties and only one of them can be checked by reading.
Why synthetic data needs a QA process of its own
Real data has errors, but they are honest errors: a sensor drifts, a form field is mistyped, a join drops rows. They are usually visible and usually local. Synthetic data has a different pathology — the errors are systematic and confident. A generator does not fail on 3 percent of rows; it applies a wrong assumption to 100 percent of rows, consistently, in a format that passes every schema check you have.
There is also a circularity trap. If you evaluate a model trained on synthetic data using a synthetic test set, you are measuring how well the model reproduces the generator. That number can be 0.96 while the real-world number is 0.58. Every validation gate below assumes you have a real held-out set that the generation pipeline has never touched.
The only number that means anything is performance on real data the generator never saw. Everything else is diagnostics.
Four failure modes, and how each one looks
Mode collapse and low diversity
The generator finds a small number of high-probability outputs and returns them repeatedly with cosmetic variation. In text: 8,000 tickets, of which 1,470 open with "I am writing to". In images: 2,000 faces, all frontal, all with the same three-point studio lighting. In tabular data: an age column whose real distribution has a mode at 34 and a long tail to 91, reproduced with 78 percent of mass between 28 and 42.
Symptom: per-row quality is high and set-level utility is low. Models trained on it overfit fast and generalise badly.
Unrealistic artefacts
Things no real record would contain. Six-fingered hands and unreadable pseudo-text in images. In tabular data: an account created on 2024-03-14 with a first order dated 2023-11-02, or a customer whose total_spend is less than the sum of their orders. These are violations of constraints the generator was never told about, because nobody wrote them down.
Symptom: a rules check catches them instantly, if you write the rules. In practice a dozen or so domain assertions catch most tabular artefacts.
Hallucinated facts
Text that is fluent, well-formed and false. A synthetic support dialogue in which the agent quotes "our standard 45-day return window" when the company's window is 30 days. This is the most dangerous mode, because it is invisible to every automated check — the text is grammatical, on-topic, correctly labelled, and not a duplicate. Train a support assistant on 5,000 such dialogues and you have taught it a policy that does not exist.
Symptom: only domain-expert review or verification against a known fact base catches this.
Distributional mismatch
The opening example. Marginals are right, joints are wrong. Or the marginals are subtly wrong in a way that reading cannot detect: the real intent mix is 34 percent billing, the synthetic is 41 percent, and no human reading 30 rows will notice a seven-point shift.
Measuring distribution fidelity with real numbers
Categorical columns: total variation distance and chi-square
Start with the simplest useful statistic. Total variation distance is half the sum of absolute differences between two probability vectors:
| Intent | Real proportion | Synthetic proportion | Absolute difference |
|---|---|---|---|
| billing | 0.34 | 0.41 | 0.07 |
| technical | 0.28 | 0.33 | 0.05 |
| shipping | 0.21 | 0.16 | 0.05 |
| account | 0.11 | 0.07 | 0.04 |
| other | 0.06 | 0.03 | 0.03 |
| Sum of absolute differences | 0.24 |
TVD = 0.24 ÷ 2 = 0.12. Interpretation is direct: 12 percent of the synthetic probability mass sits in the wrong category. A useful working threshold is TVD below 0.05 for columns the model will condition on; 0.12 is a fail.
TVD tells you the size of the discrepancy. A chi-square test tells you whether it could be sampling noise. With 2,000 synthetic rows and expected counts taken from the real proportions:
category expected observed (O-E)^2 / Ebilling 680 820 19600/680 = 28.82technical 560 660 10000/560 = 17.86shipping 420 320 10000/420 = 23.81account 220 140 6400/220 = 29.09other 120 60 3600/120 = 30.00 ---------------- chi-square statistic = 129.58 df = 4, critical value at alpha 0.05 = 9.49129.58 against a critical value of 9.49 is not close. This is a real distributional difference, not noise. Note the useful detail buried in the per-category column: the largest contributions come from account (29.09) and other (30.00), the two smallest classes — the generator is starving the tail, which is the fingerprint of mode collapse.
Numeric columns: the Kolmogorov–Smirnov statistic
For a continuous column, compare the empirical cumulative distributions and take the largest vertical gap.
| Order value | Real ECDF | Synthetic ECDF | Gap |
|---|---|---|---|
| 25 | 0.19 | 0.27 | 0.08 |
| 50 | 0.42 | 0.55 | 0.13 |
| 75 | 0.61 | 0.77 | 0.16 |
| 120 | 0.78 | 0.86 | 0.08 |
| 300 | 0.96 | 0.99 | 0.03 |
The KS statistic is the maximum gap, D = 0.16 at a value of 75. With 2,000 rows in each sample, the critical value at the 5 percent level is 1.36 × √(2 ÷ 2000) = 1.36 × 0.0316 = 0.043. Since 0.16 is nearly four times that, the distributions differ. And again the shape of the discrepancy is informative: the synthetic ECDF is above the real one everywhere, meaning synthetic orders are systematically cheaper. The generator compressed the expensive tail.
Joint structure: correlation preservation
This is the check that would have saved the team six weeks. Compute the correlation matrix on real data and on synthetic data, and compare the off-diagonal entries.
REAL SYNTHETIC val ten ref val ten refval 1.00 -0.18 0.62 val 1.00 -0.05 0.11ten -0.18 1.00 -0.31 ten -0.05 1.00 -0.09ref 0.62 -0.31 1.00 ref 0.11 -0.09 1.00absolute differences on the upper triangle: val-ten: |-0.18 - (-0.05)| = 0.13 val-ref: | 0.62 - 0.11 | = 0.51 ten-ref: |-0.31 - (-0.09)| = 0.22 mean absolute difference = (0.13 + 0.51 + 0.22) / 3 = 0.287 Frobenius norm = sqrt(0.13^2 + 0.51^2 + 0.22^2) = sqrt(0.0169 + 0.2601 + 0.0484) = sqrt(0.3254) = 0.570Every synthetic correlation is closer to zero than its real counterpart. That is the signature of a generator sampling each column independently from its own marginal — a very common bug, and one that no single-column check will ever reveal. A workable threshold is mean absolute correlation difference below 0.10, with no single pair exceeding 0.15.
Check the joints, not just the marginals. Almost every column can be individually perfect while the dataset as a whole teaches the wrong thing.
The test that subsumes the others: train a discriminator
Label real rows 1 and synthetic rows 0, shuffle, and train a gradient-boosted classifier to tell them apart with proper cross-validation. This is a classifier two-sample test, and its ROC-AUC is a single number for overall fidelity.
| Discriminator AUC | Reading |
|---|---|
| 0.50–0.55 | Indistinguishable. As good as this gets. |
| 0.55–0.70 | Distinguishable but usable for augmentation alongside real data. |
| 0.70–0.85 | Clearly different. Investigate before training on it. |
| Above 0.85 | Trivially separable. Do not ship. |
The diagnostic value is in the feature importances. In the refund case the discriminator scored AUC 0.94, and the top feature by a wide margin was the interaction between order value and refund amount — which named the bug directly. After the generator was rebuilt to model the joint distribution, AUC dropped to 0.61 and the correlation gap on val-ref fell from 0.51 to 0.06.
Measuring diversity
Fidelity says the synthetic data looks like the real data on average. Diversity says it is not the same twenty rows repeated.
| Metric | How it is computed | Real corpus | Synthetic corpus |
|---|---|---|---|
| Distinct-2 | unique bigrams ÷ total bigrams | 9,030 ÷ 12,900 = 0.70 | 5,720 ÷ 14,300 = 0.40 |
| Self-BLEU | mean BLEU of each row against the rest; higher is worse | 0.21 | 0.58 |
| Embedding dispersion | mean pairwise cosine distance of sentence embeddings | 0.71 | 0.43 |
| Cluster coverage | fraction of real k-means clusters (k = 50) with at least one synthetic neighbour | 1.00 by definition | 34 of 50 = 0.68 |
The synthetic corpus has more bigrams in total (14,300 against 12,900) and fewer unique ones — it is longer and more repetitive at once. Cluster coverage is the most actionable of the four: 16 of the 50 regions of real behaviour have no synthetic representation at all, and inspecting those 16 clusters tells you exactly which prompt axes are missing.
How bias gets in, and how to catch it
Bias enters synthetic data through four doors, and they compound.
| Entry point | Mechanism | Example |
|---|---|---|
| Seed data | Whatever skew the real examples carry is inherited | Few-shot examples drawn from one region's tickets |
| Generator priors | Associations baked into the base model | "a nurse" produces mostly female names |
| Prompt wording | Unstated defaults | "a professional" yields corporate, urban, English-first personas |
| Filtering | Quality filters correlate with group membership | A fluency filter rejects non-standard dialects at a higher rate, stripping them from the final set |
The fourth is the one teams miss, because the filter feels neutral. If your LLM-as-judge scores "clear and professional" and rejects below 4 out of 5, and the rejection rate is 12 percent for standard-register text and 38 percent for regional dialect, you have built a filter that removes dialect from your training data. Always compute rejection rates by group, not just overall.
Representation and parity, computed
Representation is the easy check: compare group proportions against the real population and flag anything outside a tolerance band. If real support tickets are 52 percent female-coded names and synthetic is 31 percent, that is a 21-point gap and a fail.
Parity is the check that matters more, because it looks at outcomes rather than counts. Suppose your synthetic loan applications carry an approved label:
real data synthetic datagroup A approval rate 0.62 0.68group B approval rate 0.55 0.41demographic parity difference: real 0.62 - 0.55 = 0.07 synthetic 0.68 - 0.41 = 0.27disparate impact ratio (min / max): real 0.55 / 0.62 = 0.887 synthetic 0.41 / 0.68 = 0.603 <-- below the 0.80 four-fifths ruleThe real data had a 7-point gap. The synthetic data has a 27-point gap — the generator did not merely inherit the bias, it amplified it nearly fourfold. This happens because generators sample towards the mode: a weak real association becomes a strong synthetic one. Any model trained on this data will learn the amplified version.
Generators do not copy bias, they sharpen it. Always compare the synthetic disparity against the real disparity, not against zero.
Model collapse: when models train on their own output
There is a failure that no single-batch check will find, because it only appears across generations. Train a generator on real data, generate, train the next generator on that output, repeat. Each round is a resampling step, and resampling loses information that never comes back.
The clean version of the argument uses one parameter. Say the true quantity is normal with variance 1.0. At each generation you draw n = 1,000 samples and refit by maximum likelihood, whose variance estimator has expectation (n−1)/n times the current variance — a factor of 0.999 per round:
| Generation | Expected variance | Ratio to original |
|---|---|---|
| 0 | 1.000 | 100% |
| 100 | 0.905 | 90% |
| 500 | 0.606 | 61% |
| 1,000 | 0.368 | 37% |
| 5,000 | 0.007 | 0.7% |
That is the slow part. The fast part is the tails. An event with true probability 0.001 appears in a 1,000-sample draw with probability 1−0.9991000=0.632 — so 37 percent of the time it does not appear at all. Once it is absent from one generation's training data, the next generator assigns it zero probability and it is gone permanently. Rare intents, minority dialects, unusual clinical presentations, the fraud pattern with 23 examples: the tail disappears first, and it disappears in a handful of rounds rather than thousands.
Two things make this practical rather than theoretical. First, generation-to-generation contamination happens by accident — a public corpus scraped this year contains last year's generated text. Second, teams do it deliberately: filter a model's output with the same model as judge, then fine-tune on what survives. That loop is a collapse loop with extra steps, because the judge shares the generator's blind spots.
What actually helps: keep a fixed real-data anchor in every training mix rather than letting the synthetic fraction drift upward; track the synthetic fraction as a recorded field; and monitor tail metrics — minimum per-class frequency, entropy of the label distribution, cluster coverage — across generations rather than only in-batch.
Privacy is a measurement, not an assumption
Synthetic data is not automatically anonymous. A generator fitted to 10,000 real records carries information about those records in its parameters and can reproduce them. Two checks, both cheap:
Distance to closest record. For each synthetic row, find its nearest neighbour in the real training set, and separately in a real holdout set the generator never saw. If the generator learned the distribution, those two distances should be similar.
5th-percentile DCR vs TRAINING set = 0.0145th-percentile DCR vs HOLDOUT set = 0.041ratio = 0.34 (healthy is near 1.0)exact duplicates of training rows = 312 / 10,000 = 3.1%A ratio of 0.34 means the generator sits three times closer to the data it memorised than to unseen real data, and 312 rows are literal copies of real people's records republished under the word "synthetic".
Membership inference. An attacker holding one candidate record asks whether it was in the training set. In a medical or financial context that alone is disclosure. The attack exploits the same memorisation — training records sit in denser regions of the synthetic output — and is scored by AUC, where 0.50 is random guessing. An attack reaching 0.71 is a real leak. The mitigation with a formal guarantee is differential privacy during generator training, which bounds any single record's influence at a measurable cost in fidelity; the mitigation without a guarantee is deleting near-duplicates and capping generator capacity.
The validation pipeline and the approval gate
Automated checks catch structure and statistics. They cannot catch a plausible sentence that states a false policy, so the last gate is human — but a targeted one. Random sampling wastes expert time. Sample stratified by the axes you generated along, plus the rows your automated checks scored worst, plus the rows nearest the decision boundary of the real-versus-synthetic discriminator.
Two reviewers on the same 100 rows, with Cohen's kappa computed on their verdicts, is the calibration step people skip. If kappa is 0.34, your reviewers do not share a definition of "acceptable" and their verdicts on the other 39,900 rows mean nothing.
1GATES = [2 ("schema", lambda r, s: s.schema_failure_rate < 0.02),3 ("domain_rules", lambda r, s: s.rule_violation_rate < 0.01),4 ("duplicates", lambda r, s: s.near_dup_rate < 0.03),5 ("tvd_intent", lambda r, s: s.tvd("intent") < 0.05),6 ("ks_value", lambda r, s: s.ks("order_value") < 0.05),7 ("corr_gap", lambda r, s: s.mean_abs_corr_diff < 0.10),8 ("discriminator", lambda r, s: s.c2st_auc < 0.70),9 ("diversity", lambda r, s: s.distinct_2 > 0.60),10 ("coverage", lambda r, s: s.cluster_coverage > 0.85),11 ("parity_delta", lambda r, s: s.parity_gap - r.parity_gap < 0.05),12 ("dcr_ratio", lambda r, s: s.dcr_ratio > 0.80),13 ("exact_copies", lambda r, s: s.exact_copies == 0),14 ("human_accept", lambda r, s: s.human_accept_rate > 0.90),15 ("downstream", lambda r, s: s.real_test_f1_lift > 0.02),16]Every threshold in that list is a decision made before generation. Thresholds set afterwards are set in favour of whatever you produced, which is not a gate — it is a rubber stamp.
What this means when you build
Write the acceptance thresholds into the repository before you write the generator. They are a specification, and a specification agreed after the fact is worthless. Fifteen lines of thresholds is enough to start.
Make the discriminator test your first diagnostic, not your last. It takes twenty minutes to build, returns a single number, and its feature importances point at the broken column. Running it before any hand-rolled statistic saves hours of checking things that were fine.
Report three numbers together, always: fidelity on real held-out data, diversity, and privacy leakage. Any one of them can be made excellent by wrecking another. Copy the real data verbatim and fidelity is perfect and privacy is zero. Generate pure noise and privacy is perfect and fidelity is zero. A synthetic dataset is only credible when all three are reported side by side, with the thresholds they were judged against.