Course Content
Synthetic Data Generation
3 sections · 7 lessons
Mini Project: Synthetic Customer Support Dialogue Generation
Two thousand synthetic support dialogues, generated over a weekend. An intent classifier trained on them scored 0.91 macro-F1 on a held-out slice of that same synthetic data. The team shipped it. On the first 300 real tickets it scored 0.44.
The post-mortem took ten minutes. Of the 2,000 dialogues, 1,317 opened with a customer message beginning "Hi, I'm having an issue with my recent order" — the same nine words, in 66 percent of the corpus. The distinct-2 score (unique word pairs divided by total word pairs) was 0.11, meaning nearly nine in ten bigrams were repeats of one already seen. The classifier had not learned to recognise ten intents. It had learned one template, and it recognised that template beautifully.
That is the whole problem with synthetic data in one dataset. Generating text is easy and getting cheaper every quarter. Generating text that is different from the other text you just generated, and then proving that it makes a downstream model better on real inputs, is the actual engineering.
This project builds that pipeline end to end: a schema, a generator, a diversity regime, a filter, augmentation, multi-turn extension, an evaluation suite, and — the part that is usually skipped — a three-arm downstream test against real, hand-labelled tickets. Every model call is against a hosted API, and the whole thing costs about 52 dollars of tokens to run once.
What you are building, and how you know it works
These are the acceptance criteria. A submission that produces a beautiful dataset failing row 9 has failed the project, because row 9 is the only row that says the dataset is worth anything.
| # | Criterion | Target | Measured by |
|---|---|---|---|
| 1 | Final dataset size | ≥ 2,500 dialogues after filtering | Line count of final_dataset.jsonl |
| 2 | Intent balance | Every intent between 5% and 12% share | Value counts over 12 intents |
| 3 | Lexical diversity | distinct-2 ≥ 0.30 | distinct_n(texts, 2) |
| 4 | Opening concentration | Most common opening 8-gram < 5% of dialogues | Counter over first 8 tokens |
| 5 | Near-duplicates | < 1.0% of pairs at Jaccard ≥ 0.85 | MinHash + LSH |
| 6 | Topic spread | Largest KMeans cluster (k=12) ≤ 25% | Sentence embeddings + KMeans |
| 7 | Turn structure | Mean 6–9 turns, standard deviation ≥ 2.5 | Turn counts per transcript |
| 8 | Multi-turn consistency | ≤ 5% of dialogues flagged for contradiction | Entity and timeline checker |
| 9 | Downstream lift | Macro-F1 on 300 real tickets improves by ≥ 0.05 over real-data-only | Three-arm training run |
| 10 | Safety | Zero unreviewed bias or toxicity flags | Regex sweep + moderation pass + human sample |
Architecture
schema.py 12 intents x 8 personas x 5 tones x 5 resolutions | = 2,400 label combinations (you cannot cover them all) v seeds.py deterministic quota walk over intent x persona | + a random concrete brief drawn from 36,000 options | -> 2,000 scenario dicts, written to disk BEFORE any API call v generate.py one call per scenario, rotated few-shot pool, | forbid-list of common openings | -> raw_dialogues.jsonl (2,000) v filter.py structural rules -> character-break -> PII -> near-dup | -> accepted.jsonl (1,653), rejected.jsonl (347) v augment.py paraphrase customer turns + persona/tone swap | source_id preserved on every derived row | -> augmented.jsonl (2,975) v extend.py context-conditioned continuation with a state block | -> final_dataset.jsonl (2,914 after consistency drops) v evaluate.py diversity, coverage, clustering, bias -> report.json v downstream.py 3 arms x 4 real-data sizes, tested on 300 REAL tickets -> the only number that decides anythingDesign the schema before you design the prompt
The schema is the axis list along which the dataset must vary. Write it first, because once you have generated 2,000 dialogues it is far too late to discover that none of them involve a customer who was already told the wrong thing by a previous agent.
1# schema.py2from dataclasses import dataclass, field34@dataclass(frozen=True)5class SupportSchema:6 intents = (7 "billing_dispute", "refund_request", "return_exchange",8 "technical_issue", "account_access", "subscription_cancellation",9 "shipping_delay", "product_defect", "feature_question",10 "general_inquiry", "warranty_claim", "order_modification",11 )12 personas = (13 "frustrated_first_time_customer", "calm_long_time_customer",14 "confused_non_technical_user", "impatient_business_customer",15 "polite_but_persistent_customer", "anxious_urgent_customer",16 "customer_already_given_wrong_answer", "terse_power_user",17 )18 tones = ("empathetic", "professional", "apologetic", "efficient", "friendly")19 resolutions = (20 "fully_resolved", "escalated_to_specialist", "partial_resolution",21 "resolved_with_compensation", "pending_followup",22 )23 min_turns, max_turns = 3, 102425def combination_count(s=SupportSchema()):26 return len(s.intents) * len(s.personas) * len(s.tones) * len(s.resolutions)27 # 12 * 8 * 5 * 5 = 2,400Two thousand dialogues cannot cover 2,400 combinations, and chasing that is the wrong goal anyway. Decide which axes must be balanced and which may simply be sampled. Intent must be balanced, because the downstream classifier predicts it. Intent crossed with persona should be balanced, because a classifier that has only ever seen billing_dispute from calm customers will fail on an angry one. Tone and resolution can be sampled freely.
Balance the two that matter with a deterministic quota walk rather than random.choice. Twelve intents times eight personas is 96 cells; 2,000 dialogues divided by 96 gives 20.83 per cell, so a round-robin assigns exactly 20 or 21 to each. Uniform random sampling would give a Poisson spread with standard deviation √20.83 = 4.6, putting real cells anywhere between about 8 and 34. The quota costs nothing, is reproducible, and removes a source of noise you would otherwise have to explain away later.
The mistake the schema does not fix
Putting persona=impatient_business_customer in the prompt does not make the dialogue different. It makes the label different. If the model writes essentially the same conversation regardless, you have a perfectly balanced set of labels sitting on top of one story, and every coverage table you produce will look excellent while the dataset is worthless.
Intended coverage is what your schema asked for. Realised coverage is what the generated text actually contains. Only the second one exists, and only the second one is worth measuring.
Generation: put the entropy in the input, not the sampler
The single most effective anti-collapse move is to give every request a concrete brief that differs from every other brief. Labels are abstract and the model rounds them off; concrete nouns are not, and it cannot.
1# seeds.py2import itertools, random34PRODUCTS = [...] # 40 fictional SKUs with names and price points5FAILURES = [...] # 15 things that go wrong: "arrived cracked", "charged twice", ...6TENURES = ["first order", "6 months", "2 years", "5 years", "lapsed"]7PRIOR = [0, 1, 2, 3] # previous contacts about this issue8TRAJECTORY = ["cools_down", "escalates", "stays_flat"]9# 40 * 15 * 5 * 4 * 3 = 36,000 distinct briefs, for 2,000 dialogues1011def build_seeds(n=2000, seed=7):12 rng = random.Random(seed)13 s = SupportSchema()14 cells = itertools.cycle(itertools.product(s.intents, s.personas)) # quota walk15 out = []16 for i in range(n):17 intent, persona = next(cells)18 out.append({19 "id": f"synth_{i:05d}",20 "intent": intent, "persona": persona,21 "tone": rng.choice(s.tones),22 "resolution": rng.choice(s.resolutions),23 "target_turns": rng.randint(s.min_turns, s.max_turns),24 "product": rng.choice(PRODUCTS),25 "failure": rng.choice(FAILURES),26 "tenure": rng.choice(TENURES),27 "prior_contacts": rng.choice(PRIOR),28 "trajectory": rng.choice(TRAJECTORY),29 "order_id": f"{rng.choice('ABCDEFGH')}-{rng.randint(10000, 99999)}",30 "amount": round(rng.uniform(9.99, 480.00), 2),31 "order_date": rng.choice(DATES),32 "exemplar_pool": i % 8, # rotate few-shot examples33 })34 return outWrite the seeds to disk before you spend a single token. A coverage gap is a two-minute fix in a table of 2,000 dicts and a two-day forensic exercise inside 2,000 transcripts.
1# generate.py2PROMPT = """Write one realistic customer support conversation.34BRIEF (the customer knows all of this; the agent starts knowing none of it)5 intent: {intent} persona: {persona}6 agent tone: {tone} target customer turns: {target_turns}7 ends in: {resolution}8 product: {product} what went wrong: {failure}9 order {order_id} for {amount} placed {order_date}10 customer tenure: {tenure}, has contacted us {prior_contacts} time(s) before11 emotional arc: {trajectory}1213RULES14- Show the persona through word choice and message length. Never name it.15- Use the order id, amount and date above. Invent no other identifiers.16- The customer reveals detail gradually. The agent asks before assuming.17- Do NOT begin with any of these openings: {forbidden}18- Strict alternation, each line prefixed CUSTOMER: or AGENT:.1920EXAMPLES OF THE REGISTER WANTED (do not copy their content)21{exemplars}2223Write the conversation:"""2425def generate(seed, forbidden_openings, exemplar_pools, temperature=1.0):26 prompt = PROMPT.format(27 **seed,28 forbidden="; ".join(forbidden_openings[:20]),29 exemplars=render(exemplar_pools[seed["exemplar_pool"]]))30 return call_model(prompt, max_tokens=1200, temperature=temperature)The forbidden line is populated from a running Counter of the first eight tokens of every dialogue accepted so far. It is a feedback loop: the more the generator converges, the harder the prompt pushes against it.
What one full run costs
| Stage | Calls | Mean input tokens | Mean output tokens | Cost (USD) |
|---|---|---|---|---|
| Seed construction | 0 | — | — | 0.00 |
| Dialogue generation | 2,000 | 450 | 600 | 20.70 |
| Paraphrase augmentation | 800 | 700 | 600 | 8.88 |
| Persona/tone swap | 800 | 700 | 620 | 9.12 |
| Multi-turn extension | 600 | 900 | 300 | 4.32 |
| State extraction for consistency | 3,600 | 400 | 80 | 8.64 |
| Total | 7,800 | 51.66 |
Priced at example rates of USD 3.00 per million input tokens and USD 15.00 per million output — check your provider's current prices before budgeting. Check one row: 2,000 × 450 = 900,000 input at USD 3.00 per million is USD 2.70, plus 2,000 × 600 = 1,200,000 output at USD 15.00 per million is USD 18.00, giving USD 20.70. Move the state-extraction row to a cheap model first: USD 8.64 is 17 percent of the budget for work a small model does perfectly well.
Mode collapse: measuring it, then beating it
Mode collapse is when a generator, asked for variety, keeps producing minor variations on one output. It is the default behaviour, not an edge case, because the model is sampling from a distribution whose peak is the most typical support conversation in its training data — and it will return to that peak every time unless something in the prompt pushes it away.
Four cheap metrics catch it. Run all four; each sees a failure the others miss.
1def distinct_n(texts, n=2):2 grams = [tuple(t.lower().split()[i:i+n])3 for t in texts4 for i in range(max(0, len(t.split()) - n + 1))]5 return len(set(grams)) / len(grams) if grams else 0.067def opening_concentration(texts, k=8):8 from collections import Counter9 heads = Counter(" ".join(t.lower().split()[:k]) for t in texts)10 return heads.most_common(1)[0][1] / len(texts)1112def cluster_spread(texts, k=12):13 emb = embed(texts) # any sentence-embedding model14 labels = KMeans(k, n_init=10, random_state=0).fit_predict(emb)15 sizes = Counter(labels)16 return {"largest_share": max(sizes.values()) / len(texts),17 "silhouette": silhouette_score(emb, labels)}1819def near_duplicate_rate(texts, threshold=0.85):20 # MinHash + LSH: linear-ish, not quadratic. See note below.21 lsh = MinHashLSH(threshold=threshold, num_perm=128)22 sigs = [minhash(shingles(t)) for t in texts]23 for i, s in enumerate(sigs):24 lsh.insert(i, s)25 dup = sum(1 for i, s in enumerate(sigs) if len(lsh.query(s)) > 1)26 return dup / len(texts)Use MinHash, not an all-pairs string comparison. Comparing 1,653 dialogues pairwise is 1,653 × 1,652 / 2 = 1,365,378 comparisons, which at roughly 0.4 ms each is about nine minutes — tolerable. At 20,000 dialogues it is 199,990,000 comparisons and 22 hours, which is not.
The intervention ladder
Measured on the same 2,000 seeds, changing one thing at a time.
| Configuration | distinct-2 | Largest cluster | Top opening 8-gram | Near-dupes |
|---|---|---|---|---|
| Schema labels only, temperature 0.7 | 0.11 | 61% | 66% | 8.7% |
| + temperature raised to 1.0 | 0.14 | 55% | 51% | 6.2% |
| + concrete seed brief per sample | 0.23 | 38% | 12% | 2.4% |
| + rotated few-shot pools (8 pools of 3) | 0.29 | 27% | 6% | 1.1% |
| + forbid-list of the 20 commonest openings | 0.34 | 21% | 2% | 0.6% |
Read the second row carefully, because raising the temperature is the intervention everybody tries first and it is the one that helps least: distinct-2 moved 0.11 to 0.14 while the largest cluster barely budged. Temperature changes which synonym gets picked. It does not change which story gets told. The seed brief moved distinct-2 by three times as much, because it changed the story.
Sampling randomness varies the wording. Input randomness varies the world. Only the second one produces a dataset that teaches a model something new.
Temperature is also not free: past about 1.0 it raises the malformed-generation rate, so the filter deletes what it bought.
Filtering, and what a healthy reject rate looks like
Generate then filter. Cheap deterministic rules first, embeddings second, model calls last — in that order, because the first stage is free and removes most of the work for the later ones.
1def screen(text, seed):2 bad = []3 lines = [l for l in text.strip().split("\n") if l.strip()]4 cust = [l for l in lines if l.upper().startswith("CUSTOMER:")]5 agent = [l for l in lines if l.upper().startswith("AGENT:")]67 if len(cust) < 3: bad.append("too_few_turns")8 if abs(len(cust) - len(agent)) > 1: bad.append("unbalanced")9 if re.search(r"(?i)as an ai|language model|i cannot", text):10 bad.append("broke_character")11 if re.search(r"\b\d{3}-\d{2}-\d{4}\b|\b\d{16}\b", text):12 bad.append("pii_pattern")13 if len(text.split()) < 40: bad.append("too_short")14 if seed["order_id"] not in text: bad.append("ignored_brief")15 if re.search(BIAS_MARKERS, text) or re.search(TOXICITY_MARKERS, text):16 bad.append("safety_review")17 return bad| Reject reason | Count of 2,000 | Share | What it usually means |
|---|---|---|---|
| Near-duplicate of an accepted dialogue | 143 | 7.2% | Collapse the ladder above did not fully remove |
| Fewer than 3 customer turns | 88 | 4.4% | Model resolved the issue in one reply |
| Unbalanced alternation | 51 | 2.6% | Format drift; tighten the format rule |
| Broke character | 34 | 1.7% | Refusal-adjacent scenario, often billing disputes |
| Real-looking PII pattern | 19 | 1.0% | Model invented a plausible card or SSN |
| Under 40 words | 12 | 0.6% | Truncation or an empty generation |
| Total rejected | 347 | 17.4% | Accepted: 1,653 |
Both extremes of the pass rate are bugs. A pass rate above about 97 percent means the filter is not catching anything and should be tested by hand-feeding it a deliberately broken dialogue. A pass rate below about 50 percent means the prompt is asking for something the model will not reliably produce — the fix belongs in the prompt, never in a looser filter. Around 80 to 90 percent is where a well-specified generator lands.
Rejects are data. Group them by intent before deleting them: if broke_character concentrates in billing_dispute, that intent is now quietly under-represented, and criterion 2 will catch it later at greater expense.
The regex safety sweep is a first pass, not a safety programme. It catches slurs and a short list of stereotype phrases and nothing else. Follow it with a moderation-API pass over the accepted set, and read a stratified sample of 100 dialogues yourself — one per intent-persona group. Skew is the bias failure that no regex sees: a set where every technical_issue customer is confused_non_technical_user teaches a model an association nobody intended.
Augmentation: what it buys and what it cannot
Two strategies, applied to 40 percent of the accepted set (661 dialogues, two variants each, giving 2,975 rows). Paraphrase rewrites only the customer lines, leaving agent lines and structure untouched, which keeps the intent label valid by construction. Persona and tone swap regenerates the same scenario as a different customer, holding intent and resolution fixed.
Every derived row carries source_id pointing at the dialogue it came from. This is not bookkeeping — it is the thing that stops the whole evaluation being a lie. A paraphrase shares roughly three quarters of its bigrams with its source. If one lands in your training split and the other in your test split, your test score is measuring memorisation, and it will look wonderful. Split on source_id, always.
Measure what augmentation actually did. In this run, size rose 1.80× while distinct-2 moved from 0.34 to 0.33 and intent coverage was unchanged by construction.
Augmentation buys surface robustness, not coverage. If the model's failure is "has never seen a warranty claim", 1,322 rewrites of the dialogues it has seen will not help, and the metrics will not tell you that.
Multi-turn extension without contradictions
Extending a dialogue by conditioning only on the last message produces turn 7 announcing that order A-4471 has not shipped, three turns after turn 4 gave its tracking number. Condition on the full transcript and on an explicit state block, so the constraints arrive as facts rather than buried in prose.
1state = {"facts_asserted": [], "commitments": [], "open_questions": [],2 "resolution_status": "open"}34def extend(sample, state, extra_turns=2):5 out = call_model(CONTINUE_PROMPT.format(6 transcript=sample["transcript"],7 state=render(state), # "agent said: shipped 14 May"8 goal=sample["resolution"], persona=sample["persona"],9 n=extra_turns), max_tokens=512)10 return out1112def contradiction_flags(before, after):13 pat = r"#?\b[A-H]-\d{5}\b|\b\d{1,3}\.\d{2}\b|\b\d{1,2} (Jan|Feb|Mar|Apr|May)\b"14 known = set(re.findall(pat, before))15 fresh = set(re.findall(pat, after)) - known16 return ["new_identifier:" + i for i in fresh] if known else []Of 600 extended dialogues, 61 were flagged and dropped, giving the final 2,914. A cheaper and better alternative to dropping is repair: truncate to the last clean turn, append the violated constraint as an explicit instruction, and regenerate forward. A contradiction at turn 7 does not invalidate turns 1 to 6, and repair typically recovers three quarters of the flagged set for one extra call.
Proving the dataset helps
Every metric so far is a diagnostic. None of them is evidence. distinct-2 of 0.34 tells you the corpus is not one template repeated; it says nothing about whether a model trained on it does better on real tickets. The only experiment that answers that needs a test set your generator never touched: 300 real, hand-labelled tickets, stratified across the 12 intents, held out and never used for tuning.
Then vary two things — how much real data you have, and which synthetic set you add — and train the same classifier each time.
| Real training examples | Real only | + 2,000 collapsed synthetic | + 2,000 diverse synthetic |
|---|---|---|---|
| 0 | — | 0.38 | 0.57 |
| 50 | 0.31 | 0.42 | 0.55 |
| 200 | 0.61 | 0.63 | 0.76 |
| 800 | 0.79 | 0.74 | 0.82 |
| 3,000 | 0.86 | 0.78 | 0.85 |
Three findings, none of them visible from inside the dataset. First, the collapsed set adds 0.02 at 200 real examples — and with 300 test items the standard error on a score near 0.6 is √(0.6 × 0.4 / 300) = 0.028, so a 95 percent interval is roughly ±5.5 points. A gain of 0.02 is indistinguishable from noise, and reporting it as a win is the most common way these projects mislead their own authors. Second, the diverse set adds 0.15 at the same point, which is comfortably outside the noise band. Third, both synthetic sets stop helping as real data grows, and the collapsed one actively hurts from 800 examples onward, dragging 0.86 down to 0.78.
That last row is the most useful thing in the table. Synthetic data is a substitute for real data you do not yet have. Its value decays as you collect the real thing, so the synthetic-to-real ratio is a hyperparameter to tune, not a setting to maximise.
Synthetic data is judged exactly one way: does a model trained with it do better on real inputs than a model trained without it? Everything else is a diagnostic for why, never evidence that it worked.
What goes wrong, and what to do about it
| Symptom | Cause | Fix |
|---|---|---|
| 0.91 on held-out synthetic, 0.44 on real tickets | The test split is synthetic; you measured template recognition | Hold out real, hand-labelled data and never generate against it |
| distinct-2 under 0.15, openings near-identical | Schema labels are in the prompt but not driving the story | Inject a concrete per-sample brief drawn from thousands of options |
| Raising temperature barely moved diversity | Temperature varies wording, not narrative | Vary the input; keep temperature near 1.0 and stop there |
| Accuracy jumped after augmentation | Paraphrase pairs split across train and test | Group-split on source_id |
| Adding more synthetic data lowers real-test score | Synthetic distribution diverges from real; volume amplifies the gap | Tune the synthetic-to-real ratio; re-check as real data grows |
| Filter pass rate 99 percent | Filter is too weak to reject anything | Feed it a deliberately broken dialogue and confirm it fails |
| Filter pass rate under 50 percent | Prompt asks for output the model will not produce | Fix the prompt; loosening the filter hides the problem |
| Every intent present, one intent still fails | Realised coverage differs from intended coverage | Cluster the generated text and inspect that intent's cluster |
| Turn 7 contradicts turn 4 | Continuation conditioned on recent text, not on state | Extract facts after each turn and inject them as constraints |
| Bias scan returns zero flags | The regex list has four entries | Add a moderation pass plus a stratified human read of 100 rows |
What this means when you build
Write downstream.py first, before generate.py. It is a hundred lines and it forces the two decisions that determine whether the project succeeds: what the real evaluation set is, and what "better" means numerically. Teams that build it last discover in week three that they have no real labelled data to test against, and by then the only honest answer is that the dataset's value is unknown.
Keep the seeds in version control, separate from the generator. Is billing under-represented? Do angry customers only appear on refunds? Those questions take seconds against a table of seeds and days against a pile of transcripts.
Store the state object and the source lineage alongside every transcript. A dialogue carrying its asserted facts, commitments and resolution status is not only an intent-classification row; it is a labelled example for state tracking and escalation prediction, free.
And keep the noise band in front of you whenever you report a number. With a 300-item test set, differences under about 5 points mean nothing — which is exactly the size of gain that gets celebrated in most write-ups of synthetic data.