Synthetic Data Generation

Quality and Bias Considerations


A team generated 40,000 synthetic e-commerce orders to test a refund-risk model. They eyeballed 30 rows. Order values looked plausible, dates were in range, customer names looked like names. Sign-off took ten minutes.

Six weeks later the model was flagging 31 percent of orders as refund risks in production against an expected 4 percent. The post-mortem found it in one line: in the real data, order value and refund amount correlate at 0.62 — expensive orders get refunded more, and for more. In the synthetic data the correlation was 0.11. The generator had learned both marginal distributions beautifully and the relationship between them not at all. Every single column passed inspection. The dataset was still useless, because the thing the model needed to learn lived in the joint distribution, and nobody had looked at the joint distribution.

Eyeballing rows tells you whether values are plausible. It cannot tell you whether the dataset is faithful. Those are different properties and only one of them can be checked by reading.

Four failure modes 30 eyeballed rows cannot showevery rowplausible, all alikedistinct-n, embedding spreadimpossible field combinationsdomain rules and range checksconfident invented entitiesverification against a sourceeach column fine on its ownKS, TVD, correlation deltaWhat it looks likeWhat catches itMode collapseUnrealistic artefactsHallucinated factsDistributional mismatchA discriminator trained to tell real from synthetic subsumes all four tests.
Reading rows one at a time tests plausibility per row, and every one of these failures lives in the distribution across 40,000 of them.

Why synthetic data needs a QA process of its own

Real data has errors, but they are honest errors: a sensor drifts, a form field is mistyped, a join drops rows. They are usually visible and usually local. Synthetic data has a different pathology — the errors are systematic and confident. A generator does not fail on 3 percent of rows; it applies a wrong assumption to 100 percent of rows, consistently, in a format that passes every schema check you have.

There is also a circularity trap. If you evaluate a model trained on synthetic data using a synthetic test set, you are measuring how well the model reproduces the generator. That number can be 0.96 while the real-world number is 0.58. Every validation gate below assumes you have a real held-out set that the generation pipeline has never touched.

The only number that means anything is performance on real data the generator never saw. Everything else is diagnostics.

Four failure modes, and how each one looks

Mode collapse and low diversity

The generator finds a small number of high-probability outputs and returns them repeatedly with cosmetic variation. In text: 8,000 tickets, of which 1,470 open with "I am writing to". In images: 2,000 faces, all frontal, all with the same three-point studio lighting. In tabular data: an age column whose real distribution has a mode at 34 and a long tail to 91, reproduced with 78 percent of mass between 28 and 42.

Symptom: per-row quality is high and set-level utility is low. Models trained on it overfit fast and generalise badly.

Unrealistic artefacts

Things no real record would contain. Six-fingered hands and unreadable pseudo-text in images. In tabular data: an account created on 2024-03-14 with a first order dated 2023-11-02, or a customer whose total_spend is less than the sum of their orders. These are violations of constraints the generator was never told about, because nobody wrote them down.

Symptom: a rules check catches them instantly, if you write the rules. In practice a dozen or so domain assertions catch most tabular artefacts.

Hallucinated facts

Text that is fluent, well-formed and false. A synthetic support dialogue in which the agent quotes "our standard 45-day return window" when the company's window is 30 days. This is the most dangerous mode, because it is invisible to every automated check — the text is grammatical, on-topic, correctly labelled, and not a duplicate. Train a support assistant on 5,000 such dialogues and you have taught it a policy that does not exist.

Symptom: only domain-expert review or verification against a known fact base catches this.

Distributional mismatch

The opening example. Marginals are right, joints are wrong. Or the marginals are subtly wrong in a way that reading cannot detect: the real intent mix is 34 percent billing, the synthetic is 41 percent, and no human reading 30 rows will notice a seven-point shift.

Measuring distribution fidelity with real numbers

Categorical columns: total variation distance and chi-square

Start with the simplest useful statistic. Total variation distance is half the sum of absolute differences between two probability vectors:

IntentReal proportionSynthetic proportionAbsolute difference
billing0.340.410.07
technical0.280.330.05
shipping0.210.160.05
account0.110.070.04
other0.060.030.03
Sum of absolute differences0.24

TVD = 0.24 ÷ 2 = 0.12. Interpretation is direct: 12 percent of the synthetic probability mass sits in the wrong category. A useful working threshold is TVD below 0.05 for columns the model will condition on; 0.12 is a fail.

TVD tells you the size of the discrepancy. A chi-square test tells you whether it could be sampling noise. With 2,000 synthetic rows and expected counts taken from the real proportions:

Text
category    expected   observed   (O-E)^2 / Ebilling        680        820      19600/680  = 28.82technical      560        660      10000/560  = 17.86shipping       420        320      10000/420  = 23.81account        220        140       6400/220  = 29.09other          120         60       3600/120  = 30.00                                    ----------------                          chi-square statistic = 129.58  df = 4, critical value at alpha 0.05 = 9.49

129.58 against a critical value of 9.49 is not close. This is a real distributional difference, not noise. Note the useful detail buried in the per-category column: the largest contributions come from account (29.09) and other (30.00), the two smallest classes — the generator is starving the tail, which is the fingerprint of mode collapse.

Numeric columns: the Kolmogorov–Smirnov statistic

For a continuous column, compare the empirical cumulative distributions and take the largest vertical gap.

Order valueReal ECDFSynthetic ECDFGap
250.190.270.08
500.420.550.13
750.610.770.16
1200.780.860.08
3000.960.990.03

The KS statistic is the maximum gap, D = 0.16 at a value of 75. With 2,000 rows in each sample, the critical value at the 5 percent level is 1.36 × √(2 ÷ 2000) = 1.36 × 0.0316 = 0.043. Since 0.16 is nearly four times that, the distributions differ. And again the shape of the discrepancy is informative: the synthetic ECDF is above the real one everywhere, meaning synthetic orders are systematically cheaper. The generator compressed the expensive tail.

Joint structure: correlation preservation

This is the check that would have saved the team six weeks. Compute the correlation matrix on real data and on synthetic data, and compare the off-diagonal entries.

Text
REAL                       SYNTHETIC        val   ten   ref            val   ten   refval    1.00 -0.18  0.62     val   1.00 -0.05  0.11ten   -0.18  1.00 -0.31     ten  -0.05  1.00 -0.09ref    0.62 -0.31  1.00     ref   0.11 -0.09  1.00absolute differences on the upper triangle:  val-ten:  |-0.18 - (-0.05)| = 0.13  val-ref:  | 0.62 -   0.11 | = 0.51  ten-ref:  |-0.31 - (-0.09)| = 0.22  mean absolute difference = (0.13 + 0.51 + 0.22) / 3 = 0.287  Frobenius norm = sqrt(0.13^2 + 0.51^2 + 0.22^2)                 = sqrt(0.0169 + 0.2601 + 0.0484) = sqrt(0.3254) = 0.570

Every synthetic correlation is closer to zero than its real counterpart. That is the signature of a generator sampling each column independently from its own marginal — a very common bug, and one that no single-column check will ever reveal. A workable threshold is mean absolute correlation difference below 0.10, with no single pair exceeding 0.15.

Check the joints, not just the marginals. Almost every column can be individually perfect while the dataset as a whole teaches the wrong thing.

The test that subsumes the others: train a discriminator

Label real rows 1 and synthetic rows 0, shuffle, and train a gradient-boosted classifier to tell them apart with proper cross-validation. This is a classifier two-sample test, and its ROC-AUC is a single number for overall fidelity.

Discriminator AUCReading
0.50–0.55Indistinguishable. As good as this gets.
0.55–0.70Distinguishable but usable for augmentation alongside real data.
0.70–0.85Clearly different. Investigate before training on it.
Above 0.85Trivially separable. Do not ship.

The diagnostic value is in the feature importances. In the refund case the discriminator scored AUC 0.94, and the top feature by a wide margin was the interaction between order value and refund amount — which named the bug directly. After the generator was rebuilt to model the joint distribution, AUC dropped to 0.61 and the correlation gap on val-ref fell from 0.51 to 0.06.

Measuring diversity

Fidelity says the synthetic data looks like the real data on average. Diversity says it is not the same twenty rows repeated.

MetricHow it is computedReal corpusSynthetic corpus
Distinct-2unique bigrams ÷ total bigrams9,030 ÷ 12,900 = 0.705,720 ÷ 14,300 = 0.40
Self-BLEUmean BLEU of each row against the rest; higher is worse0.210.58
Embedding dispersionmean pairwise cosine distance of sentence embeddings0.710.43
Cluster coveragefraction of real k-means clusters (k = 50) with at least one synthetic neighbour1.00 by definition34 of 50 = 0.68

The synthetic corpus has more bigrams in total (14,300 against 12,900) and fewer unique ones — it is longer and more repetitive at once. Cluster coverage is the most actionable of the four: 16 of the 50 regions of real behaviour have no synthetic representation at all, and inspecting those 16 clusters tells you exactly which prompt axes are missing.

How bias gets in, and how to catch it

Bias enters synthetic data through four doors, and they compound.

Entry pointMechanismExample
Seed dataWhatever skew the real examples carry is inheritedFew-shot examples drawn from one region's tickets
Generator priorsAssociations baked into the base model"a nurse" produces mostly female names
Prompt wordingUnstated defaults"a professional" yields corporate, urban, English-first personas
FilteringQuality filters correlate with group membershipA fluency filter rejects non-standard dialects at a higher rate, stripping them from the final set

The fourth is the one teams miss, because the filter feels neutral. If your LLM-as-judge scores "clear and professional" and rejects below 4 out of 5, and the rejection rate is 12 percent for standard-register text and 38 percent for regional dialect, you have built a filter that removes dialect from your training data. Always compute rejection rates by group, not just overall.

Representation and parity, computed

Representation is the easy check: compare group proportions against the real population and flag anything outside a tolerance band. If real support tickets are 52 percent female-coded names and synthetic is 31 percent, that is a 21-point gap and a fail.

Parity is the check that matters more, because it looks at outcomes rather than counts. Suppose your synthetic loan applications carry an approved label:

Text
                       real data        synthetic datagroup A approval rate     0.62               0.68group B approval rate     0.55               0.41demographic parity difference:   real       0.62 - 0.55 = 0.07   synthetic  0.68 - 0.41 = 0.27disparate impact ratio (min / max):   real       0.55 / 0.62 = 0.887   synthetic  0.41 / 0.68 = 0.603      <-- below the 0.80 four-fifths rule

The real data had a 7-point gap. The synthetic data has a 27-point gap — the generator did not merely inherit the bias, it amplified it nearly fourfold. This happens because generators sample towards the mode: a weak real association becomes a strong synthetic one. Any model trained on this data will learn the amplified version.

Generators do not copy bias, they sharpen it. Always compare the synthetic disparity against the real disparity, not against zero.

Model collapse: when models train on their own output

There is a failure that no single-batch check will find, because it only appears across generations. Train a generator on real data, generate, train the next generator on that output, repeat. Each round is a resampling step, and resampling loses information that never comes back.

The clean version of the argument uses one parameter. Say the true quantity is normal with variance 1.0. At each generation you draw n = 1,000 samples and refit by maximum likelihood, whose variance estimator has expectation (n−1)/n times the current variance — a factor of 0.999 per round:

GenerationExpected varianceRatio to original
01.000100%
1000.90590%
5000.60661%
1,0000.36837%
5,0000.0070.7%

That is the slow part. The fast part is the tails. An event with true probability 0.001 appears in a 1,000-sample draw with probability 1−0.9991000=0.6321 - 0.999^{1000} = 0.632 — so 37 percent of the time it does not appear at all. Once it is absent from one generation's training data, the next generator assigns it zero probability and it is gone permanently. Rare intents, minority dialects, unusual clinical presentations, the fraud pattern with 23 examples: the tail disappears first, and it disappears in a handful of rounds rather than thousands.

Two things make this practical rather than theoretical. First, generation-to-generation contamination happens by accident — a public corpus scraped this year contains last year's generated text. Second, teams do it deliberately: filter a model's output with the same model as judge, then fine-tune on what survives. That loop is a collapse loop with extra steps, because the judge shares the generator's blind spots.

What actually helps: keep a fixed real-data anchor in every training mix rather than letting the synthetic fraction drift upward; track the synthetic fraction as a recorded field; and monitor tail metrics — minimum per-class frequency, entropy of the label distribution, cluster coverage — across generations rather than only in-batch.

Privacy is a measurement, not an assumption

Synthetic data is not automatically anonymous. A generator fitted to 10,000 real records carries information about those records in its parameters and can reproduce them. Two checks, both cheap:

Distance to closest record. For each synthetic row, find its nearest neighbour in the real training set, and separately in a real holdout set the generator never saw. If the generator learned the distribution, those two distances should be similar.

Text
5th-percentile DCR vs TRAINING set = 0.0145th-percentile DCR vs HOLDOUT  set = 0.041ratio = 0.34          (healthy is near 1.0)exact duplicates of training rows = 312 / 10,000 = 3.1%

A ratio of 0.34 means the generator sits three times closer to the data it memorised than to unseen real data, and 312 rows are literal copies of real people's records republished under the word "synthetic".

Membership inference. An attacker holding one candidate record asks whether it was in the training set. In a medical or financial context that alone is disclosure. The attack exploits the same memorisation — training records sit in denser regions of the synthetic output — and is scored by AUC, where 0.50 is random guessing. An attack reaching 0.71 is a real leak. The mitigation with a formal guarantee is differential privacy during generator training, which bounds any single record's influence at a measurable cost in fidelity; the mitigation without a guarantee is deleting near-duplicates and capping generator capacity.

The validation pipeline and the approval gate

Automated checks catch structure and statistics. They cannot catch a plausible sentence that states a false policy, so the last gate is human — but a targeted one. Random sampling wastes expert time. Sample stratified by the axes you generated along, plus the rows your automated checks scored worst, plus the rows nearest the decision boundary of the real-versus-synthetic discriminator.

Two reviewers on the same 100 rows, with Cohen's kappa computed on their verdicts, is the calibration step people skip. If kappa is 0.34, your reviewers do not share a definition of "acceptable" and their verdicts on the other 39,900 rows mean nothing.

Python
GATES = [    ("schema",        lambda r, s: s.schema_failure_rate < 0.02),    ("domain_rules",  lambda r, s: s.rule_violation_rate < 0.01),    ("duplicates",    lambda r, s: s.near_dup_rate < 0.03),    ("tvd_intent",    lambda r, s: s.tvd("intent") < 0.05),    ("ks_value",      lambda r, s: s.ks("order_value") < 0.05),    ("corr_gap",      lambda r, s: s.mean_abs_corr_diff < 0.10),    ("discriminator", lambda r, s: s.c2st_auc < 0.70),    ("diversity",     lambda r, s: s.distinct_2 > 0.60),    ("coverage",      lambda r, s: s.cluster_coverage > 0.85),    ("parity_delta",  lambda r, s: s.parity_gap - r.parity_gap < 0.05),    ("dcr_ratio",     lambda r, s: s.dcr_ratio > 0.80),    ("exact_copies",  lambda r, s: s.exact_copies == 0),    ("human_accept",  lambda r, s: s.human_accept_rate > 0.90),    ("downstream",    lambda r, s: s.real_test_f1_lift > 0.02),]

Every threshold in that list is a decision made before generation. Thresholds set afterwards are set in favour of whatever you produced, which is not a gate — it is a rubber stamp.

What this means when you build

Write the acceptance thresholds into the repository before you write the generator. They are a specification, and a specification agreed after the fact is worthless. Fifteen lines of thresholds is enough to start.

Make the discriminator test your first diagnostic, not your last. It takes twenty minutes to build, returns a single number, and its feature importances point at the broken column. Running it before any hand-rolled statistic saves hours of checking things that were fine.

Report three numbers together, always: fidelity on real held-out data, diversity, and privacy leakage. Any one of them can be made excellent by wrecking another. Copy the real data verbatim and fidelity is perfect and privacy is zero. Generate pure noise and privacy is perfect and fidelity is zero. A synthetic dataset is only credible when all three are reported side by side, with the thresholds they were judged against.