Course Content
Synthetic Data Generation
3 sections · 7 lessons
Prompt-Based Generation and Filtering
Marco generated 12,000 synthetic support tickets overnight. The run cost USD 78 in API charges. His schema validator passed 11,304 of them, which felt like a 94 percent success rate, so he appended them to his 5,000 real tickets and retrained the intent classifier.
Macro-F1 on the real held-out set went from 0.742 to 0.719. The synthetic data made the model measurably worse.
The autopsy took a day and found three separate problems. 2,138 of the 11,304 rows were near-duplicates of another row in the same file — same complaint, different name. 1,604 carried a label that contradicted their own text, mostly billing complaints tagged as account. And 410 dialogues confidently referenced a 45-day returns window that the company does not offer; every one of those was fluent, on-topic, correctly labelled and completely false.
Schema validation had caught none of it, because none of it is a schema problem. Generation is the easy half. The filtering pipeline is where a synthetic dataset is actually made, and the number that matters is not how many rows you produced but how many survived.
The generation prompt as a specification
A prompt for a chat assistant asks for one good answer. A prompt for data generation defines a population: rows that code will parse, filters will score, and a model will treat as ground truth. Write it like a spec.
Role, constraints, and — the part people omit — non-goals
Positive instructions tell the model what to produce. Non-goals tell it what to stop doing, and a model's helpful defaults are the largest source of unusable rows. Left unconstrained, a model asked for a customer complaint will resolve the complaint, sign off with a name, invent an order number in a format you do not use, and end with an offer to help further.
ROLEYou are building a labelled dataset of inbound customer support messagesfor an intent classifier. You are not a support agent.MUST- Write only the customer's message. Never the agent's reply.- Length 25-180 words unless the length hint says otherwise.- Use exactly one intent from: billing, technical, shipping, account, other.- Reference only these products: Orbit Router, Orbit Router Pro, Orbit Mesh.MUST NOT- Do not resolve, apologise for, or explain the issue.- Do not invent policies, refund windows, warranty terms or prices.- Do not include signatures, greetings addressed to a named agent, or "Thank you for your time".- Do not mention any competitor or any real company.OUTPUTReturn only a JSON object matching the schema. No prose, no code fences.The "do not invent policies" line is doing the heaviest lifting. It is what stops the 45-day returns window — the failure that no automated check downstream can catch.
Where the schema instruction goes
Put the schema at the end, immediately before generation begins, not buried in the middle of a long system prompt. Models attend most reliably to instructions nearest the generation point, and format compliance measurably improves when the format rule is last. If your provider supports constrained decoding against a JSON schema, use it as well — but keep the textual instruction, because it also shapes the content, not only the syntax.
Exemplars: rotate them, and make some of them hard
Two rules, both routinely broken. Rotate. If the same three examples appear in all 12,000 calls, every row is a variation on three seeds. Draw three exemplars per call at random from a pool of thirty and the anchoring is spread. Include the ugly ones. If all exemplars are clean and unambiguous, the generator never produces the ticket that mixes a billing question into a delivery complaint, misspells the product name, and trails off mid-sentence — which is precisely the ticket your classifier fails on in production. One deliberately messy exemplar in three shifts the difficulty distribution of the entire output.
Diversity is designed, not sampled
Sampling the same prompt 12,000 times draws 12,000 times from one conditional distribution. The number of distinct outputs is bounded by how many modes that distribution has, not by how many draws you take. Raising temperature widens the modes slightly; it does not create new ones.
Rotate explicit axes
1import itertools, random23AXES = {4 "persona": ["terse power user", "confused first-time buyer",5 "small-business owner in a hurry", "polite persistent retiree",6 "typing on a phone with typos"],7 "topic": ["duplicate charge", "package marked delivered but missing",8 "promo code rejected", "cannot cancel subscription",9 "device drops wifi nightly", "password reset loop",10 "wrong item shipped", "refund not received"],11 "channel": ["email", "live chat", "phone transcript"],12 "length": ["under 30 words", "60-90 words", "200 words, rambling"],13}1415KEYS = list(AXES)16CELLS = list(itertools.product(*(AXES[k] for k in KEYS))) # 5*8*3*3 = 360 cells17random.shuffle(CELLS) # shuffle so a partial run stays balanced1819def cell_prompt(cell, pool, template):20 spec = dict(zip(KEYS, cell))21 shots = random.sample(pool, 3)22 return template.format(**spec, examples="\n\n".join(shots))Shuffling matters more than it looks. Iterating the product in order means a run that dies at 40 percent has generated every row for the first two personas and none for the last three.
Temperature per field, not per call
Free-text and categorical labels want opposite settings. Text at temperature 0.3 is stilted and repetitive; a label at temperature 1.2 is a coin flip weighted by nothing useful. If your provider lets you generate fields in separate calls, use temperature 0.9, top_p 0.95 for the message body and temperature 0.1 for the label. If you must produce both in one structured response, keep temperature at 0.7–0.9 and accept that you will need a label-verification gate later — you will need one anyway. (All of this assumes your model exposes sampling parameters; some current reasoning models ignore or reject them, which leaves the prompt as your only lever.)
Rotate the template itself
Even with 360 axis cells, one template imposes one rhetorical shape. Keep three or four structurally different templates — one that gives the persona as a first-person brief, one that describes the situation in third person, one that supplies a partial transcript to continue — and sample among them. Measured on 2,000 rows, adding template rotation on top of axis rotation lifted distinct-3 from 0.61 to 0.68 and cut the near-duplicate rate from 4.4 percent to 2.1 percent.
Measure it, or you are guessing
| Strategy (2,000 rows each) | Distinct-3 | Near-dup rate | Cluster coverage (k=50) |
|---|---|---|---|
| Single prompt, T = 0.7 | 0.31 | 18.4% | 0.42 |
| Single prompt, T = 1.3 | 0.39 | 11.7% | 0.48 |
| 360 axis cells, T = 0.8 | 0.61 | 4.4% | 0.79 |
| 360 cells × 4 templates, T = 0.8 | 0.68 | 2.1% | 0.86 |
Raising temperature by 0.6 bought 0.08 of distinct-3. Designing the variation axes bought 0.37, at lower temperature and better coherence.
Structured output: shape without truth
Define the row schema once and reuse the same object for generation, validation, and consumption. The classic bug is drift — the prompt says sentiment, the validator checks tone, and a gate silently passes 100 percent of rows because it is reading a field that does not exist.
1from pydantic import BaseModel, Field2from typing import Literal34class Ticket(BaseModel):5 text: str = Field(min_length=40, max_length=1200)6 intent: Literal["billing", "technical", "shipping", "account", "other"]7 sentiment: Literal["angry", "frustrated", "neutral", "polite"]8 channel: Literal["email", "chat", "phone_transcript"]910# MODEL comes from config. Both SDKs turn the Pydantic class into a JSON schema11# for constrained decoding, then validate the reply back into a Ticket.1213# Provider A: OpenAI Responses API14resp = openai_client.responses.parse(15 model=MODEL, input=prompt, text_format=Ticket)16ticket = resp.output_parsed1718# Provider B: Anthropic Messages API, native structured outputs19resp = anthropic_client.messages.parse(20 model=MODEL, max_tokens=1500,21 messages=[{"role": "user", "content": prompt}],22 output_format=Ticket)23ticket = resp.parsed_outputConstrained decoding masks the token distribution at each step so that only tokens keeping the output schema-valid can be sampled. Invalid JSON becomes impossible rather than unlikely — parse failure rates drop from roughly 6–8 percent to effectively zero. Constraints the decoder does not enforce, such as min_length, are still checked when the SDK validates the parsed object. You will also see the older trick of forcing a tool call to get JSON back; it still works on many models, but some current models reject a forced tool_choice, so prefer the native structured-output parameter.
And it changes nothing about quality. Every one of these is schema-perfect and worthless: a billing complaint labelled shipping; the 47th near-copy of the same complaint; a fluent reference to a returns policy that does not exist. Structured output is the cheapest gate, so it goes first — but a gate that checks only shape catches only shape.
Gate 1: schema and domain rules
Schema validation is free and deterministic. Domain rules cost you fifteen lines and catch a surprising share of what schema validation misses:
1RULES = [2 ("no_agent_voice", lambda t: "I'm sorry to hear" not in t.text),3 ("no_signature", lambda t: not SIGNOFF_RE.search(t.text)),4 ("real_product", lambda t: all(p in PRODUCTS for p in extract_products(t.text))),5 ("no_policy_claim", lambda t: not POLICY_RE.search(t.text)),6 ("word_count", lambda t: 20 <= len(t.text.split()) <= 260),7 ("label_in_taxonomy", lambda t: t.intent in TAXONOMY),8]On Marco's 12,000 rows: schema alone passed 11,304 (94.2 percent). Adding the six domain rules removed a further 388, including 341 of the 410 fabricated-policy rows, because a regex for \d+[- ]day (return|refund|warranty) catches most invented policies. That regex is thirty seconds of work and removed 83 percent of the most dangerous failure class.
Gate 2: exact and near-duplicate removal
Exact duplicates are easy — normalise whitespace and case, hash, drop repeats. That took 611 rows out of Marco's file. Near-duplicates are the real problem: same complaint, different name and product, cosine similarity 0.94, indistinguishable to a hash.
Choosing the threshold with evidence
Hand-label 400 known duplicate pairs and 4,000 known distinct pairs, then sweep:
| Cosine threshold | Duplicates caught (of 400) | Recall | False positives |
|---|---|---|---|
| 0.98 | 168 | 0.42 | 3 |
| 0.95 | 210 | 0.53 | 12 |
| 0.92 | 318 | 0.80 | 96 |
| 0.88 | 372 | 0.93 | 480 |
0.92 is the knee: dropping from 0.92 to 0.88 buys 54 more true duplicates at a cost of 384 more false positives — a rate of seven good rows destroyed per duplicate removed. Above 0.92 you are leaving most duplicates in. This sweep takes an afternoon and is the difference between a principled threshold and a number someone remembered from a blog post.
Scaling it: MinHash and LSH
All-pairs comparison of 17,000 rows is 144 million comparisons. MinHash with locality-sensitive hashing avoids that. Represent each document by the set of its 3-word shingles and estimate Jaccard similarity:
|A intersect B| = 42 |A union B| = 97Jaccard J = 42 / 97 = 0.433With 128 hash permutations, the MinHash estimate hasstandard error = sqrt(J(1-J)/128) = sqrt(0.2455/128) = 0.044LSH banding: 128 signature rows split into b = 32 bands of r = 4.P(pair becomes a candidate) = 1 - (1 - J^r)^b J = 0.80 -> 1 - (1 - 0.4096)^32 = 1.000 J = 0.42 -> approximately 0.64 (the S-curve threshold, (1/b)^(1/r)) J = 0.30 -> 1 - (1 - 0.0081)^32 = 0.229Choose b and r so the S-curve's midpoint sits just below your target similarity. LSH turns 144 million comparisons into a few million candidate pairs, which you then score exactly.
Gate 3: LLM-as-judge, and whether your judge is any good
A judge model reads a row and scores it. This catches label errors and incoherence that no rule can express. It also introduces a new component whose accuracy you now have to measure — and almost nobody does.
Marco's judge, on 200 rows he had also labelled by hand:
human accept human rejectjudge accept 148 22 (170)judge reject 14 16 ( 30) (162) (38) (200)raw agreement = (148 + 16) / 200 = 0.820chance agreement = (170/200)(162/200) + (30/200)(38/200) = 0.6885 + 0.0285 = 0.717Cohen's kappa = (0.820 - 0.717) / (1 - 0.717) = 0.103 / 0.283 = 0.364Eighty-two percent agreement sounds respectable. Kappa of 0.364 says it is barely better than a judge that accepts everything, because 81 percent of rows are acceptable anyway — agreeing by default is cheap. This is the trap: report kappa, never raw agreement, whenever the base rate is skewed.
Three changes moved Marco's kappa from 0.364 to 0.68:
- Score dimensions separately. Not "is this good, 1–5", but three binary judgements: does the label match the text, does it state any policy or fact not in the provided context, would a real customer plausibly write this. Concrete questions get consistent answers.
- Give the rubric examples. Two accepted and two rejected rows with one-line reasons, inside the judge prompt.
- Use a different model family from the generator. Judges show self-preference: they tend to rate output from their own model family higher than equally good output from another. A judge grading its own output is a filter with the generator's blind spots built in.
An unmeasured judge is not a quality gate. It is a second generator whose output you have decided to trust.
Run the judge at temperature 0 where the model allows it, ask for the verdict field before any explanation only if you do not need the reasoning, and log every score — the distribution of judge scores across axis cells is one of the best diagnostics you will get for free.
Gate 4: rejection sampling to a balanced target
After three gates you have a pile of accepted rows whose composition reflects where generation happened to succeed, not what you need. Judges reject long rambling messages more often than short ones, so the surviving set is systematically shorter than intended.
Fix it by sampling to a quota. Target 8,000 rows across 40 reporting cells is 200 per cell. Take min(200, available) from each cell, note the shortfalls, and regenerate only the deficient cells:
cells at or above quota: 28 -> take 200 each = 5,600cells below quota: 12 -> take all = 1,280 total = 6,880shortfall = 8,000 - 6,880 = 1,120 rows across 12 cellsat an end-to-end yield of 0.463, regenerating those 12 cells needs1,120 / 0.463 = 2,419 additional raw generationsTargeted regeneration is far cheaper than generating uniformly more and hoping. And when a cell keeps under-delivering after two attempts, that is information: something in the prompt for "200 words, rambling" plus "phone transcript" is producing rows your judge rejects, and the fix is in the prompt, not in the sample size.
The funnel, priced
| Stage | Rows in | Removed | Rows out | Stage yield |
|---|---|---|---|---|
| Raw generation | — | — | 12,000 | — |
| Schema + domain rules | 12,000 | 1,084 | 10,916 | 91.0% |
| Exact duplicates | 10,916 | 611 | 10,305 | 94.4% |
| Near-duplicates (0.92) | 10,305 | 1,344 | 8,961 | 87.0% |
| LLM judge | 8,961 | 1,910 | 7,051 | 78.7% |
| Quota sampling | 7,051 | 1,496 | 5,555 | 78.8% |
| End-to-end yield | 46.3% |
Plan from the end. To ship 8,000 rows you must generate 8,000 ÷ 0.463 = 17,279. At roughly USD 0.0065 per raw row that is about USD 112 in generation. The judge runs on the 12,900 rows that reach it, at about USD 0.0009 each — USD 12. Embeddings for deduplication, at roughly 120 tokens per row, cost under a dollar. Filtering is about 10 percent of the cost of generation and removes 54 percent of the rows. Anyone arguing that a gate is too expensive has not done this arithmetic.
Yield is also a metric, not just a planning input. A yield that drops from 46 percent to 31 percent between runs means the generator, the model version, or the prompt changed. Log it per run and per cell.
What this means when you build
Build the gates before the generator. It feels backwards and it is not: gates written first are honest thresholds, and gates written after inspecting output are thresholds chosen to pass the output you already have. Fifteen rules and a judge rubric are an afternoon's work.
Keep every rejected row, with the gate that rejected it and the reason. The rejects are the most informative artefact in the pipeline. When the judge is throwing out 41 percent of rows from one axis cell, the rejected rows tell you which instruction the generator is misreading, and one prompt edit recovers thousands of rows.
Measure the judge against humans on 200 rows before trusting it on 200,000, and report kappa rather than agreement. Then re-measure after any prompt or model change. A pipeline whose most consequential filter has never been validated is a pipeline that will produce Marco's result: a clean, well-formatted, schema-valid dataset that makes the model worse.