Course Content
Synthetic Data Generation
3 sections · 7 lessons
LLMs and Diffusion Models as Generators
A team needed 500 customer complaints to train an intent classifier. They wrote one prompt — "Write a realistic customer complaint about a late delivery" — set temperature to 0.7, and ran it 500 times. The generation took eleven minutes and cost under two dollars. It looked like a win.
Then someone grepped the output. The phrase "absolutely furious" appeared in 214 of the 500 complaints. Ninety-one of them opened with the exact words "I am writing to express". Counting distinct three-word openings across the whole file gave 61. The classifier trained on this data hit 0.96 on a synthetic validation split and 0.58 on real tickets, because it had learned that "absolutely furious" means late delivery — a rule that holds in the synthetic world and nowhere else.
Nothing was broken. The model did exactly what sampling from a fixed conditional distribution does: it returned high-probability text, 500 times, from the same distribution. Understanding why that produces near-duplicates, and what to change, is the difference between a generator and a photocopier.
Why a language model is a data generator at all
A language model computes, for a given prefix, a probability distribution over the next token. Generation means repeatedly sampling from that distribution and appending the result. So a prompt does not request a document; it selects a conditional distribution, and every sample you take is a draw from that same distribution.
This reframing explains the failure above immediately. If you draw 500 times from one distribution, the number of distinct outcomes you see is bounded by how many high-probability modes that distribution has — not by how many times you sample. If the prompt "Write a realistic customer complaint about a late delivery" has roughly 40 well-worn ways of starting, you will see roughly those 40 openings whether you sample 500 times or 5,000.
You do not get diversity by sampling more from one prompt. You get it by sampling from more prompts.
Anatomy of a data-generation prompt
A prompt written for a chat assistant and a prompt written to manufacture training rows are different objects. The assistant prompt optimises for one good answer. The generation prompt optimises for a population of rows that will be parsed by code, filtered automatically, and used as supervision.
| Component | Purpose | What goes wrong without it |
|---|---|---|
| Role | Fixes register and vocabulary ("You are a support ticket triage dataset builder") | Output drifts into assistant voice: helpful, polite, uniformly well-punctuated |
| Positive constraints | Length range, required fields, allowed label set | Rows vary from 8 to 800 words; labels appear that are not in your taxonomy |
| Non-goals | "Do not resolve the issue. Do not include a signature. Do not mention brand names." | The model helpfully adds things you must strip later, inconsistently |
| Schema instruction | Exact output shape, field names, types | Prose wrapped around JSON; markdown fences; trailing commentary |
| Few-shot exemplars | Demonstrates tone, length and difficulty far better than description | The model's default style dominates; hard cases never appear |
| Variation axis | The one slot that changes per call: persona, product, channel, mood | 500 near-copies |
Few-shot exemplars: two rules people break
First, rotate them. If the same three examples appear in every call, the model anchors on them and your 500 rows are variations on three seeds. Sample three exemplars per call from a pool of thirty and the anchoring is spread across the pool.
Second, include the hard cases. If all your exemplars are clean, unambiguous complaints, the model will not invent the ticket that mixes a billing question with a delivery complaint and contains a typo in the order number — which is exactly the ticket your classifier fails on in production. One deliberately messy exemplar in three changes the difficulty distribution of the whole output.
Structured output that actually parses
"Return JSON" in the prompt gets you valid JSON maybe 92 percent of the time. The other 8 percent is markdown fences, a preamble, a trailing "Let me know if you'd like more examples", or a hallucinated extra field. At 80,000 rows, 8 percent is 6,400 parse failures.
Two mechanisms fix this at the API level rather than the prompt level. Constrained decoding against a JSON schema masks the token distribution at each step so only tokens that keep the output schema-valid can be sampled — invalid output becomes impossible, not unlikely. A tool marked strict makes the model emit arguments that match a declared input schema, which is the same idea wearing a different name.
1from pydantic import BaseModel, Field2from typing import Literal34class SupportTicket(BaseModel):5 text: str = Field(min_length=40, max_length=900)6 intent: Literal["billing", "technical", "shipping", "account", "other"]7 sentiment: Literal["angry", "frustrated", "neutral", "polite"]8 channel: Literal["email", "chat", "phone_transcript"]9 contains_order_id: bool1011# One schema object, reused three ways:12# 1. constrained decoding at generation time13# 2. validation gate after generation14# 3. the contract your training pipeline reads15SCHEMA = SupportTicket.model_json_schema()Define the schema once and reuse the same object for generation, validation and consumption. The common bug is a schema that drifts: the prompt says sentiment, the validator checks tone, and 100 percent of rows silently fail a check nobody reads.
Structured output is not quality control
This is where people get it wrong. Constrained decoding guarantees shape, not truth and not usefulness. Every one of the following is schema-valid and worthless:
- Text describing a billing problem, labelled
intent: "shipping" contains_order_id: trueon text with no order ID in it- The same complaint, near-verbatim, for the 47th time
- A complaint mentioning a refund policy the company does not have
Schema validation is the cheapest filter in the pipeline and it must be the first, not the only one.
The diversity problem, and the knobs that address it
Temperature, computed
Temperature rescales the logits before the softmax: pi=exp(zi/T)/∑jexp(zj/T). Take three candidate tokens with logits 3.0, 2.0 and 1.0 and work it through.
| Temperature | P(token A) | P(token B) | P(token C) | Effect |
|---|---|---|---|---|
| 0.5 | 0.867 | 0.117 | 0.016 | Near-deterministic; C is essentially unreachable |
| 1.0 | 0.665 | 0.245 | 0.090 | The model's own distribution |
| 1.5 | 0.563 | 0.289 | 0.148 | Flatter; C now appears once in seven |
Notice the size of the effect. Halving temperature from 1.0 to 0.5 cut the third token's probability from 0.090 to 0.016 — a factor of 5.6. Raising it to 1.5 only lifted that token to 0.148, a factor of 1.6. Temperature is far more powerful at suppressing variety than at creating it, which is why cranking it to 1.4 to fix duplicate output mostly produces the same 40 modes plus occasional garbled tokens.
Top-p, computed
Nucleus sampling keeps the smallest set of tokens whose probabilities sum to at least p, then renormalises. With a sorted tail of 0.42, 0.21, 0.13, 0.08, 0.06, 0.04, 0.03, 0.02, 0.01:
cumulative: 0.42 0.63 0.76 0.84 0.90 0.94 0.97 0.99 1.00top_p = 0.90 keeps the first 5 tokens, discards the restrenormalised (divide by 0.90): 0.467 0.233 0.144 0.089 0.067Four tokens carrying 10 percent of the mass between them are removed entirely. That is the point — those are usually where incoherence comes from — but it is also why the outputs converge. Typical settings: temperature 0.9, top_p 0.95 for open-ended text; temperature 0.3, top_p 1.0 for the label field, if you can generate fields separately. Some current models, reasoning models in particular, ignore or reject temperature and top_p; on those, variety has to come from the prompt — which, as the next section shows, is where it works best anyway.
The knob that actually works: varying the prompt
The measured comparison, on 500 generated tickets each way:
| Strategy | Distinct-3 ratio | Exact/near duplicates |
|---|---|---|
| One prompt, T = 0.7 | 0.31 | 18.4% |
| One prompt, T = 1.3 | 0.39 | 11.7% |
| 25 personas × 20 topics, T = 0.8 | 0.68 | 2.1% |
Distinct-3 is the count of unique three-word sequences divided by the total count of three-word sequences; higher means less repetition. Raising temperature by 0.6 bought 0.08. Building a 500-cell grid of prompt variations bought 0.37, at lower temperature — so the text is also more coherent.
1import random, itertools23PERSONAS = ["a terse power user", "a first-time buyer who is confused",4 "a small-business owner in a hurry", "a polite but persistent retiree",5 "someone typing on a phone with typos"]6TOPICS = ["a duplicate charge", "a package marked delivered but missing",7 "a promo code that was rejected", "a subscription they cannot cancel"]8CHANNELS = ["email", "live chat", "phone transcript"]9LENGTHS = ["under 30 words", "60-90 words", "over 200 words, rambling"]1011CELLS = list(itertools.product(PERSONAS, TOPICS, CHANNELS, LENGTHS))12random.shuffle(CELLS) # 5 * 4 * 3 * 3 = 180 distinct prompt cells1314def build_prompt(cell, exemplar_pool):15 persona, topic, channel, length = cell16 shots = random.sample(exemplar_pool, 3) # rotate the few-shot set too17 return PROMPT_TEMPLATE.format(18 persona=persona, topic=topic, channel=channel,19 length=length, examples="\n\n".join(shots))Walk the cells in shuffled order rather than looping over the outer axis, so that a run interrupted at 40 percent still has balanced coverage rather than every persona being "a terse power user".
Diffusion models as image generators
A diffusion model is trained on a destructive process run backwards. Take a real image, add Gaussian noise in small steps until it is pure static, and train a network to predict the noise that was added at each step. To generate, start from pure static and repeatedly subtract the predicted noise. After 30 or so steps you have an image that was never photographed.
The parameters that matter
| Parameter | What it controls | Practical range | Failure at the extreme |
|---|---|---|---|
| Steps | How many denoising iterations | 20–50 with a modern scheduler | Below 10: soft, mushy. Above 60: cost with no visible gain |
| Guidance scale (CFG) | How hard the prompt pulls the image | 5–9 | Below 3: prompt ignored. Above 15: oversaturated, burnt edges, warped anatomy |
| Seed | The initial noise tensor | Record it, always | Unrecorded seed means the dataset is unreproducible |
| Negative prompt | A second conditioning the sampler pushes away from | "blurry, watermark, text, extra fingers" | Overloading it flattens the whole output distribution |
| Scheduler | The step-size schedule for denoising | DPM++ 2M, Euler a, DDIM | Mismatched scheduler and step count gives visible banding |
These ranges suit classic Stable Diffusion-style models (SD 1.5, SDXL). Newer flow-matching models such as FLUX.1 and Stable Diffusion 3.5 use their own schedulers and much lower guidance — FLUX.1-dev defaults to a guidance scale of 3.5, and it ignores a negative prompt unless you enable true classifier-free guidance — so start from the settings on the model card.
Guidance is worth understanding rather than tuning blindly. At each step the model predicts noise twice — once conditioned on your prompt, once unconditioned — and the sampler uses ϵ^=ϵuncond+w(ϵcond−ϵuncond). With w = 1 the conditional term is used as-is. With w = 7.5 you are extrapolating 7.5 times past the conditional prediction, away from the unconditional one. That extrapolation is what makes images obey prompts, and it is also why high guidance produces over-contrasted, high-saturation pictures: you are pushing past the region the model was trained on.
Conditioning beyond the text prompt
For dataset work, text prompts are the weakest form of control. Four stronger ones:
- img2img with a strength parameter. Start the denoising from a noised version of a real image rather than pure static. Strength 0.25 changes lighting and texture while keeping composition and — critically — keeping bounding boxes valid. Strength 0.8 keeps only the vague layout.
- Inpainting. Regenerate only a masked region. This is how you place a rare defect onto 400 photographs of a real product: the background stays real, the defect varies, and you know the mask so you get the segmentation label for free.
- Structural conditioning (ControlNet and similar). Condition on a depth map, edge map, segmentation mask or human pose. Because you supply the structure, you already hold the ground-truth annotation. Generate 5,000 images from 5,000 pose skeletons and you have 5,000 labelled keypoint examples with zero annotation cost.
- Subject adaptation (LoRA, DreamBooth). Fine-tune a small adapter on 15–30 photographs of one specific object so the generator can place that object into new scenes. This is also the biggest memorisation risk in image synthesis — a LoRA trained on 20 photographs of a person will reproduce that person.
Generating a systematic grid, not a pile
Random prompts give random coverage. A dataset needs balanced coverage, which means enumerating the conditions you care about and generating a fixed count per cell.
1import itertools2import torch34AXES = {5 "defect": ["hairline crack", "surface scratch", "discolouration", "none"],6 "lighting": ["harsh overhead", "soft diffuse", "low warehouse light"],7 "angle": ["top-down", "45 degrees", "eye level"],8 "material": ["matte plastic", "brushed metal"],9}1011keys = list(AXES)12cells = list(itertools.product(*(AXES[k] for k in keys))) # 4*3*3*2 = 72 cells13PER_CELL = 40 # 2,880 images total1415for cell_idx, cell in enumerate(cells):16 spec = dict(zip(keys, cell))17 prompt = ("product photo of a {material} housing, {defect}, "18 "{lighting}, shot {angle}").format(**spec)19 for k in range(PER_CELL):20 seed = cell_idx * 10_000 + k # deterministic and reconstructable21 image = pipe(prompt, negative_prompt=NEG, num_inference_steps=30,22 guidance_scale=7.0,23 generator=torch.Generator("cuda").manual_seed(seed)).images[0]24 save(image, meta={**spec, "seed": seed, "model": MODEL_ID})The meta dictionary is not optional bookkeeping. When the defect detector later shows 0.91 recall on cracks and 0.44 on discolouration, the per-cell metadata is what lets you find out that discolouration only rendered convincingly under soft diffuse lighting.
Choosing between them
| Dimension | LLM (text) | Diffusion (images) |
|---|---|---|
| Label comes from | The same generation call — ask for text and label together | The conditioning input — mask, pose, or the grid cell |
| Output validity | Enforceable by constrained decoding | Not enforceable; must be filtered visually or by a classifier |
| Typical unit cost | USD 0.003–0.01 per row via API | USD 0.002 self-hosted, USD 0.04 hosted |
| Dominant failure | Repetition and hallucinated facts | Anatomy, text-in-image, physically impossible geometry |
| Controllability | Very high — you can state constraints in words | Low from text alone, very high with structural conditioning |
| Human review speed | Slow: minutes per row to read carefully | Fast: a reviewer clears 400 thumbnails in an hour |
That last row is quietly important for planning. Image review is roughly ten times faster per item than text review, so image pipelines can afford a much higher human-review sampling rate for the same budget.
Cost, throughput and licensing
Cost, worked
Text, 80,000 rows at 850 input and 260 output tokens, at the same example rates: 68.0 million input tokens at USD 3 per million is USD 204, and 20.8 million output tokens at USD 15 per million is USD 312 — about USD 516 in total. Routing generation to a cheaper model at, say, USD 0.80 and USD 4.00 per million drops that to USD 54 and USD 83, about USD 137. A batch endpoint at a 50 percent discount halves it again.
Images, 10,000 at 1024×1024, 30 steps, 3.8 seconds each on a GPU costing USD 2.20 per hour: 3.8 ÷ 3600 × 2.20 = USD 0.00232 per image, so USD 23.20 in compute — but 10,000 × 3.8 s = 10.6 hours on one GPU, or 1.3 hours across eight. The same 10,000 images through a hosted API at USD 0.04 each cost USD 400. Self-hosting wins on unit cost by 17× and loses on setup time; below about 5,000 images the API is almost always the right call.
Throughput and rate limits
Unit cost rarely decides the schedule; rate limits do. Suppose your account allows 400,000 output tokens per minute and each row uses 260. The ceiling is 400,000 ÷ 260 = 1,538 rows per minute, so 80,645 rows needs at least 52 minutes of pure generation. Real pipelines hit perhaps 40 percent of the ceiling once retries, backoff and long-tail latency are included, so budget around 2.2 hours. Practical measures:
- Ask for k rows per call rather than one. Ten rows per response amortises the 850-token prompt across ten outputs, cutting input cost by about 90 percent — but rows within one response are more similar to each other, so keep k at 5–10 and never at 50.
- Use asynchronous concurrency with a bounded semaphore, and back off on 429 rather than retrying immediately.
- Checkpoint to disk every few hundred rows. A pipeline that loses two hours of generation to an unhandled exception is a pipeline that will lose it again.
Licensing and provenance
Three questions have to be answered before generated data enters a training set, and they are legal questions with engineering consequences.
| Question | Why it bites | What to record |
|---|---|---|
| Do the provider's terms permit using outputs to train a model? | Several commercial APIs restrict using outputs to build a competing model. Open-weight models under permissive licences generally do not. | Provider, model ID, terms version, date of generation |
| Can the generator reproduce copyrighted or personal material? | Image models memorise; a LoRA on 20 photographs of a person reproduces that person | Which real assets were used for adaptation, and a duplicate check against them |
| Can you prove which rows are synthetic later? | Auditors, and your future self, will ask | A per-row provenance field, plus C2PA content credentials for images where supported |
If you cannot answer "which model produced this row, with which prompt and seed, on which date" for every row in your dataset, you do not have a dataset — you have a pile.
What this means when you build
Design the variation axes before you write the prompt. The single highest-leverage decision in a generation pipeline is not the model, the temperature or the prompt wording — it is the list of dimensions along which rows are allowed to differ. Write that list down as data, count the cells, and decide how many rows per cell you need. If you cannot name at least three axes with four or more values each, you will produce near-duplicates no matter what else you do.
Generate the label in the same call as the content, but validate it separately. Asking a model to write a complaint and assign its intent in one structured response is cheap and mostly correct. Mostly is the operative word: expect 3–8 percent label noise, and plan a second pass where a different model, or a classifier trained on real data, re-labels a sample and you measure agreement.
Record everything at generation time, because you cannot reconstruct it afterwards. Model ID, prompt template version, axis values, sampling parameters, seed, timestamp. It costs a few hundred bytes per row. When the classifier underperforms on one slice three months later, that metadata is the difference between a two-hour diagnosis and a rebuild from scratch.