Synthetic Data Generation

Mini Project: Synthetic Customer Support Dialogue Generation


Two thousand synthetic support dialogues, generated over a weekend. An intent classifier trained on them scored 0.91 macro-F1 on a held-out slice of that same synthetic data. The team shipped it. On the first 300 real tickets it scored 0.44.

The post-mortem took ten minutes. Of the 2,000 dialogues, 1,317 opened with a customer message beginning "Hi, I'm having an issue with my recent order" — the same nine words, in 66 percent of the corpus. The distinct-2 score (unique word pairs divided by total word pairs) was 0.11, meaning nearly nine in ten bigrams were repeats of one already seen. The classifier had not learned to recognise ten intents. It had learned one template, and it recognised that template beautifully.

That is the whole problem with synthetic data in one dataset. Generating text is easy and getting cheaper every quarter. Generating text that is different from the other text you just generated, and then proving that it makes a downstream model better on real inputs, is the actual engineering.

This project builds that pipeline end to end: a schema, a generator, a diversity regime, a filter, augmentation, multi-turn extension, an evaluation suite, and — the part that is usually skipped — a three-arm downstream test against real, hand-labelled tickets. Every model call is against a hosted API, and the whole thing costs about 52 dollars of tokens to run once.

The pipeline that would have caught 0.44Schema and non-goals fixed before the promptEntropy in the input, not in the samplerFilter: rules, exact and near-duplicatesMulti-turn extension with a state scratchpadScore on held-out REAL tickets only
0.91 macro-F1 on a synthetic hold-out measures how self-similar the generator is; only real tickets can measure whether the dataset taught the classifier anything.

What you are building, and how you know it works

These are the acceptance criteria. A submission that produces a beautiful dataset failing row 9 has failed the project, because row 9 is the only row that says the dataset is worth anything.

#CriterionTargetMeasured by
1Final dataset size≥ 2,500 dialogues after filteringLine count of final_dataset.jsonl
2Intent balanceEvery intent between 5% and 12% shareValue counts over 12 intents
3Lexical diversitydistinct-2 ≥ 0.30distinct_n(texts, 2)
4Opening concentrationMost common opening 8-gram < 5% of dialoguesCounter over first 8 tokens
5Near-duplicates< 1.0% of pairs at Jaccard ≥ 0.85MinHash + LSH
6Topic spreadLargest KMeans cluster (k=12) ≤ 25%Sentence embeddings + KMeans
7Turn structureMean 6–9 turns, standard deviation ≥ 2.5Turn counts per transcript
8Multi-turn consistency≤ 5% of dialogues flagged for contradictionEntity and timeline checker
9Downstream liftMacro-F1 on 300 real tickets improves by ≥ 0.05 over real-data-onlyThree-arm training run
10SafetyZero unreviewed bias or toxicity flagsRegex sweep + moderation pass + human sample

Architecture

Text
  schema.py            12 intents x 8 personas x 5 tones x 5 resolutions       |                = 2,400 label combinations (you cannot cover them all)       v  seeds.py             deterministic quota walk over intent x persona       |               + a random concrete brief drawn from 36,000 options       |               -> 2,000 scenario dicts, written to disk BEFORE any API call       v  generate.py          one call per scenario, rotated few-shot pool,       |               forbid-list of common openings       |               -> raw_dialogues.jsonl (2,000)       v  filter.py            structural rules -> character-break -> PII -> near-dup       |               -> accepted.jsonl (1,653), rejected.jsonl (347)       v  augment.py           paraphrase customer turns + persona/tone swap       |               source_id preserved on every derived row       |               -> augmented.jsonl (2,975)       v  extend.py            context-conditioned continuation with a state block       |               -> final_dataset.jsonl (2,914 after consistency drops)       v  evaluate.py          diversity, coverage, clustering, bias  -> report.json       v  downstream.py        3 arms x 4 real-data sizes, tested on 300 REAL tickets                       -> the only number that decides anything

Design the schema before you design the prompt

The schema is the axis list along which the dataset must vary. Write it first, because once you have generated 2,000 dialogues it is far too late to discover that none of them involve a customer who was already told the wrong thing by a previous agent.

Python
# schema.pyfrom dataclasses import dataclass, field@dataclass(frozen=True)class SupportSchema:    intents = (        "billing_dispute", "refund_request", "return_exchange",        "technical_issue", "account_access", "subscription_cancellation",        "shipping_delay", "product_defect", "feature_question",        "general_inquiry", "warranty_claim", "order_modification",    )    personas = (        "frustrated_first_time_customer", "calm_long_time_customer",        "confused_non_technical_user", "impatient_business_customer",        "polite_but_persistent_customer", "anxious_urgent_customer",        "customer_already_given_wrong_answer", "terse_power_user",    )    tones = ("empathetic", "professional", "apologetic", "efficient", "friendly")    resolutions = (        "fully_resolved", "escalated_to_specialist", "partial_resolution",        "resolved_with_compensation", "pending_followup",    )    min_turns, max_turns = 3, 10def combination_count(s=SupportSchema()):    return len(s.intents) * len(s.personas) * len(s.tones) * len(s.resolutions)    # 12 * 8 * 5 * 5 = 2,400

Two thousand dialogues cannot cover 2,400 combinations, and chasing that is the wrong goal anyway. Decide which axes must be balanced and which may simply be sampled. Intent must be balanced, because the downstream classifier predicts it. Intent crossed with persona should be balanced, because a classifier that has only ever seen billing_dispute from calm customers will fail on an angry one. Tone and resolution can be sampled freely.

Balance the two that matter with a deterministic quota walk rather than random.choice. Twelve intents times eight personas is 96 cells; 2,000 dialogues divided by 96 gives 20.83 per cell, so a round-robin assigns exactly 20 or 21 to each. Uniform random sampling would give a Poisson spread with standard deviation √20.83 = 4.6, putting real cells anywhere between about 8 and 34. The quota costs nothing, is reproducible, and removes a source of noise you would otherwise have to explain away later.

The mistake the schema does not fix

Putting persona=impatient_business_customer in the prompt does not make the dialogue different. It makes the label different. If the model writes essentially the same conversation regardless, you have a perfectly balanced set of labels sitting on top of one story, and every coverage table you produce will look excellent while the dataset is worthless.

Intended coverage is what your schema asked for. Realised coverage is what the generated text actually contains. Only the second one exists, and only the second one is worth measuring.

Generation: put the entropy in the input, not the sampler

The single most effective anti-collapse move is to give every request a concrete brief that differs from every other brief. Labels are abstract and the model rounds them off; concrete nouns are not, and it cannot.

Python
# seeds.pyimport itertools, randomPRODUCTS   = [...]   # 40 fictional SKUs with names and price pointsFAILURES   = [...]   # 15 things that go wrong: "arrived cracked", "charged twice", ...TENURES    = ["first order", "6 months", "2 years", "5 years", "lapsed"]PRIOR      = [0, 1, 2, 3]              # previous contacts about this issueTRAJECTORY = ["cools_down", "escalates", "stays_flat"]# 40 * 15 * 5 * 4 * 3 = 36,000 distinct briefs, for 2,000 dialoguesdef build_seeds(n=2000, seed=7):    rng = random.Random(seed)    s = SupportSchema()    cells = itertools.cycle(itertools.product(s.intents, s.personas))  # quota walk    out = []    for i in range(n):        intent, persona = next(cells)        out.append({            "id": f"synth_{i:05d}",            "intent": intent, "persona": persona,            "tone": rng.choice(s.tones),            "resolution": rng.choice(s.resolutions),            "target_turns": rng.randint(s.min_turns, s.max_turns),            "product": rng.choice(PRODUCTS),            "failure": rng.choice(FAILURES),            "tenure": rng.choice(TENURES),            "prior_contacts": rng.choice(PRIOR),            "trajectory": rng.choice(TRAJECTORY),            "order_id": f"{rng.choice('ABCDEFGH')}-{rng.randint(10000, 99999)}",            "amount": round(rng.uniform(9.99, 480.00), 2),            "order_date": rng.choice(DATES),            "exemplar_pool": i % 8,        # rotate few-shot examples        })    return out

Write the seeds to disk before you spend a single token. A coverage gap is a two-minute fix in a table of 2,000 dicts and a two-day forensic exercise inside 2,000 transcripts.

Python
# generate.pyPROMPT = """Write one realistic customer support conversation.BRIEF (the customer knows all of this; the agent starts knowing none of it)  intent: {intent}          persona: {persona}  agent tone: {tone}        target customer turns: {target_turns}  ends in: {resolution}  product: {product}        what went wrong: {failure}  order {order_id} for {amount} placed {order_date}  customer tenure: {tenure}, has contacted us {prior_contacts} time(s) before  emotional arc: {trajectory}RULES- Show the persona through word choice and message length. Never name it.- Use the order id, amount and date above. Invent no other identifiers.- The customer reveals detail gradually. The agent asks before assuming.- Do NOT begin with any of these openings: {forbidden}- Strict alternation, each line prefixed CUSTOMER: or AGENT:.EXAMPLES OF THE REGISTER WANTED (do not copy their content){exemplars}Write the conversation:"""def generate(seed, forbidden_openings, exemplar_pools, temperature=1.0):    prompt = PROMPT.format(        **seed,        forbidden="; ".join(forbidden_openings[:20]),        exemplars=render(exemplar_pools[seed["exemplar_pool"]]))    return call_model(prompt, max_tokens=1200, temperature=temperature)

The forbidden line is populated from a running Counter of the first eight tokens of every dialogue accepted so far. It is a feedback loop: the more the generator converges, the harder the prompt pushes against it.

What one full run costs

StageCallsMean input tokensMean output tokensCost (USD)
Seed construction0——0.00
Dialogue generation2,00045060020.70
Paraphrase augmentation8007006008.88
Persona/tone swap8007006209.12
Multi-turn extension6009003004.32
State extraction for consistency3,600400808.64
Total7,80051.66

Priced at example rates of USD 3.00 per million input tokens and USD 15.00 per million output — check your provider's current prices before budgeting. Check one row: 2,000 × 450 = 900,000 input at USD 3.00 per million is USD 2.70, plus 2,000 × 600 = 1,200,000 output at USD 15.00 per million is USD 18.00, giving USD 20.70. Move the state-extraction row to a cheap model first: USD 8.64 is 17 percent of the budget for work a small model does perfectly well.

Mode collapse: measuring it, then beating it

Mode collapse is when a generator, asked for variety, keeps producing minor variations on one output. It is the default behaviour, not an edge case, because the model is sampling from a distribution whose peak is the most typical support conversation in its training data — and it will return to that peak every time unless something in the prompt pushes it away.

Four cheap metrics catch it. Run all four; each sees a failure the others miss.

Python
def distinct_n(texts, n=2):    grams = [tuple(t.lower().split()[i:i+n])             for t in texts             for i in range(max(0, len(t.split()) - n + 1))]    return len(set(grams)) / len(grams) if grams else 0.0def opening_concentration(texts, k=8):    from collections import Counter    heads = Counter(" ".join(t.lower().split()[:k]) for t in texts)    return heads.most_common(1)[0][1] / len(texts)def cluster_spread(texts, k=12):    emb = embed(texts)                      # any sentence-embedding model    labels = KMeans(k, n_init=10, random_state=0).fit_predict(emb)    sizes = Counter(labels)    return {"largest_share": max(sizes.values()) / len(texts),            "silhouette": silhouette_score(emb, labels)}def near_duplicate_rate(texts, threshold=0.85):    # MinHash + LSH: linear-ish, not quadratic. See note below.    lsh = MinHashLSH(threshold=threshold, num_perm=128)    sigs = [minhash(shingles(t)) for t in texts]    for i, s in enumerate(sigs):        lsh.insert(i, s)    dup = sum(1 for i, s in enumerate(sigs) if len(lsh.query(s)) > 1)    return dup / len(texts)

Use MinHash, not an all-pairs string comparison. Comparing 1,653 dialogues pairwise is 1,653 × 1,652 / 2 = 1,365,378 comparisons, which at roughly 0.4 ms each is about nine minutes — tolerable. At 20,000 dialogues it is 199,990,000 comparisons and 22 hours, which is not.

The intervention ladder

Measured on the same 2,000 seeds, changing one thing at a time.

Configurationdistinct-2Largest clusterTop opening 8-gramNear-dupes
Schema labels only, temperature 0.70.1161%66%8.7%
+ temperature raised to 1.00.1455%51%6.2%
+ concrete seed brief per sample0.2338%12%2.4%
+ rotated few-shot pools (8 pools of 3)0.2927%6%1.1%
+ forbid-list of the 20 commonest openings0.3421%2%0.6%

Read the second row carefully, because raising the temperature is the intervention everybody tries first and it is the one that helps least: distinct-2 moved 0.11 to 0.14 while the largest cluster barely budged. Temperature changes which synonym gets picked. It does not change which story gets told. The seed brief moved distinct-2 by three times as much, because it changed the story.

Sampling randomness varies the wording. Input randomness varies the world. Only the second one produces a dataset that teaches a model something new.

Temperature is also not free: past about 1.0 it raises the malformed-generation rate, so the filter deletes what it bought.

Filtering, and what a healthy reject rate looks like

Generate then filter. Cheap deterministic rules first, embeddings second, model calls last — in that order, because the first stage is free and removes most of the work for the later ones.

Python
def screen(text, seed):    bad = []    lines = [l for l in text.strip().split("\n") if l.strip()]    cust  = [l for l in lines if l.upper().startswith("CUSTOMER:")]    agent = [l for l in lines if l.upper().startswith("AGENT:")]    if len(cust) < 3:                          bad.append("too_few_turns")    if abs(len(cust) - len(agent)) > 1:        bad.append("unbalanced")    if re.search(r"(?i)as an ai|language model|i cannot", text):                                               bad.append("broke_character")    if re.search(r"\b\d{3}-\d{2}-\d{4}\b|\b\d{16}\b", text):                                               bad.append("pii_pattern")    if len(text.split()) < 40:                 bad.append("too_short")    if seed["order_id"] not in text:           bad.append("ignored_brief")    if re.search(BIAS_MARKERS, text) or re.search(TOXICITY_MARKERS, text):                                               bad.append("safety_review")    return bad
Reject reasonCount of 2,000ShareWhat it usually means
Near-duplicate of an accepted dialogue1437.2%Collapse the ladder above did not fully remove
Fewer than 3 customer turns884.4%Model resolved the issue in one reply
Unbalanced alternation512.6%Format drift; tighten the format rule
Broke character341.7%Refusal-adjacent scenario, often billing disputes
Real-looking PII pattern191.0%Model invented a plausible card or SSN
Under 40 words120.6%Truncation or an empty generation
Total rejected34717.4%Accepted: 1,653

Both extremes of the pass rate are bugs. A pass rate above about 97 percent means the filter is not catching anything and should be tested by hand-feeding it a deliberately broken dialogue. A pass rate below about 50 percent means the prompt is asking for something the model will not reliably produce — the fix belongs in the prompt, never in a looser filter. Around 80 to 90 percent is where a well-specified generator lands.

Rejects are data. Group them by intent before deleting them: if broke_character concentrates in billing_dispute, that intent is now quietly under-represented, and criterion 2 will catch it later at greater expense.

The regex safety sweep is a first pass, not a safety programme. It catches slurs and a short list of stereotype phrases and nothing else. Follow it with a moderation-API pass over the accepted set, and read a stratified sample of 100 dialogues yourself — one per intent-persona group. Skew is the bias failure that no regex sees: a set where every technical_issue customer is confused_non_technical_user teaches a model an association nobody intended.

Augmentation: what it buys and what it cannot

Two strategies, applied to 40 percent of the accepted set (661 dialogues, two variants each, giving 2,975 rows). Paraphrase rewrites only the customer lines, leaving agent lines and structure untouched, which keeps the intent label valid by construction. Persona and tone swap regenerates the same scenario as a different customer, holding intent and resolution fixed.

Every derived row carries source_id pointing at the dialogue it came from. This is not bookkeeping — it is the thing that stops the whole evaluation being a lie. A paraphrase shares roughly three quarters of its bigrams with its source. If one lands in your training split and the other in your test split, your test score is measuring memorisation, and it will look wonderful. Split on source_id, always.

Measure what augmentation actually did. In this run, size rose 1.80× while distinct-2 moved from 0.34 to 0.33 and intent coverage was unchanged by construction.

Augmentation buys surface robustness, not coverage. If the model's failure is "has never seen a warranty claim", 1,322 rewrites of the dialogues it has seen will not help, and the metrics will not tell you that.

Multi-turn extension without contradictions

Extending a dialogue by conditioning only on the last message produces turn 7 announcing that order A-4471 has not shipped, three turns after turn 4 gave its tracking number. Condition on the full transcript and on an explicit state block, so the constraints arrive as facts rather than buried in prose.

Python
state = {"facts_asserted": [], "commitments": [], "open_questions": [],         "resolution_status": "open"}def extend(sample, state, extra_turns=2):    out = call_model(CONTINUE_PROMPT.format(        transcript=sample["transcript"],        state=render(state),                 # "agent said: shipped 14 May"        goal=sample["resolution"], persona=sample["persona"],        n=extra_turns), max_tokens=512)    return outdef contradiction_flags(before, after):    pat = r"#?\b[A-H]-\d{5}\b|\b\d{1,3}\.\d{2}\b|\b\d{1,2} (Jan|Feb|Mar|Apr|May)\b"    known = set(re.findall(pat, before))    fresh = set(re.findall(pat, after)) - known    return ["new_identifier:" + i for i in fresh] if known else []

Of 600 extended dialogues, 61 were flagged and dropped, giving the final 2,914. A cheaper and better alternative to dropping is repair: truncate to the last clean turn, append the violated constraint as an explicit instruction, and regenerate forward. A contradiction at turn 7 does not invalidate turns 1 to 6, and repair typically recovers three quarters of the flagged set for one extra call.

Proving the dataset helps

Every metric so far is a diagnostic. None of them is evidence. distinct-2 of 0.34 tells you the corpus is not one template repeated; it says nothing about whether a model trained on it does better on real tickets. The only experiment that answers that needs a test set your generator never touched: 300 real, hand-labelled tickets, stratified across the 12 intents, held out and never used for tuning.

Then vary two things — how much real data you have, and which synthetic set you add — and train the same classifier each time.

Real training examplesReal only+ 2,000 collapsed synthetic+ 2,000 diverse synthetic
0—0.380.57
500.310.420.55
2000.610.630.76
8000.790.740.82
3,0000.860.780.85

Three findings, none of them visible from inside the dataset. First, the collapsed set adds 0.02 at 200 real examples — and with 300 test items the standard error on a score near 0.6 is √(0.6 × 0.4 / 300) = 0.028, so a 95 percent interval is roughly ±5.5 points. A gain of 0.02 is indistinguishable from noise, and reporting it as a win is the most common way these projects mislead their own authors. Second, the diverse set adds 0.15 at the same point, which is comfortably outside the noise band. Third, both synthetic sets stop helping as real data grows, and the collapsed one actively hurts from 800 examples onward, dragging 0.86 down to 0.78.

That last row is the most useful thing in the table. Synthetic data is a substitute for real data you do not yet have. Its value decays as you collect the real thing, so the synthetic-to-real ratio is a hyperparameter to tune, not a setting to maximise.

Synthetic data is judged exactly one way: does a model trained with it do better on real inputs than a model trained without it? Everything else is a diagnostic for why, never evidence that it worked.

What goes wrong, and what to do about it

SymptomCauseFix
0.91 on held-out synthetic, 0.44 on real ticketsThe test split is synthetic; you measured template recognitionHold out real, hand-labelled data and never generate against it
distinct-2 under 0.15, openings near-identicalSchema labels are in the prompt but not driving the storyInject a concrete per-sample brief drawn from thousands of options
Raising temperature barely moved diversityTemperature varies wording, not narrativeVary the input; keep temperature near 1.0 and stop there
Accuracy jumped after augmentationParaphrase pairs split across train and testGroup-split on source_id
Adding more synthetic data lowers real-test scoreSynthetic distribution diverges from real; volume amplifies the gapTune the synthetic-to-real ratio; re-check as real data grows
Filter pass rate 99 percentFilter is too weak to reject anythingFeed it a deliberately broken dialogue and confirm it fails
Filter pass rate under 50 percentPrompt asks for output the model will not produceFix the prompt; loosening the filter hides the problem
Every intent present, one intent still failsRealised coverage differs from intended coverageCluster the generated text and inspect that intent's cluster
Turn 7 contradicts turn 4Continuation conditioned on recent text, not on stateExtract facts after each turn and inject them as constraints
Bias scan returns zero flagsThe regex list has four entriesAdd a moderation pass plus a stratified human read of 100 rows

What this means when you build

Write downstream.py first, before generate.py. It is a hundred lines and it forces the two decisions that determine whether the project succeeds: what the real evaluation set is, and what "better" means numerically. Teams that build it last discover in week three that they have no real labelled data to test against, and by then the only honest answer is that the dataset's value is unknown.

Keep the seeds in version control, separate from the generator. Is billing under-represented? Do angry customers only appear on refunds? Those questions take seconds against a table of seeds and days against a pile of transcripts.

Store the state object and the source lineage alongside every transcript. A dialogue carrying its asserted facts, commitments and resolution status is not only an intent-classification row; it is a labelled example for state tracking and escalation prediction, free.

And keep the noise band in front of you whenever you report a number. With a 300-item test set, differences under about 5 points mean nothing — which is exactly the size of gain that gets celebrated in most write-ups of synthetic data.