Synthetic Data Generation

LLMs and Diffusion Models as Generators


A team needed 500 customer complaints to train an intent classifier. They wrote one prompt — "Write a realistic customer complaint about a late delivery" — set temperature to 0.7, and ran it 500 times. The generation took eleven minutes and cost under two dollars. It looked like a win.

Then someone grepped the output. The phrase "absolutely furious" appeared in 214 of the 500 complaints. Ninety-one of them opened with the exact words "I am writing to express". Counting distinct three-word openings across the whole file gave 61. The classifier trained on this data hit 0.96 on a synthetic validation split and 0.58 on real tickets, because it had learned that "absolutely furious" means late delivery — a rule that holds in the synthetic world and nowhere else.

Nothing was broken. The model did exactly what sampling from a fixed conditional distribution does: it returned high-probability text, 500 times, from the same distribution. Understanding why that produces near-duplicates, and what to change, is the difference between a generator and a photocopier.

Where the diversity actually comes fromTurning up temperature• Reweights one conditional distribution• Varies wording inside a single scenario• Past about 1.0, coherence degrades first• 500 runs of one prompt, one complaintVarying the prompt• Changes which distribution you sample• Rotates product, fault, tone, region• Quality stays where you tuned it• 25 personas × 20 topics, one row each
Sampling knobs move entropy within a scenario; only the input can add scenarios, which is why one prompt run 500 times yields 500 near-copies.

Why a language model is a data generator at all

A language model computes, for a given prefix, a probability distribution over the next token. Generation means repeatedly sampling from that distribution and appending the result. So a prompt does not request a document; it selects a conditional distribution, and every sample you take is a draw from that same distribution.

This reframing explains the failure above immediately. If you draw 500 times from one distribution, the number of distinct outcomes you see is bounded by how many high-probability modes that distribution has — not by how many times you sample. If the prompt "Write a realistic customer complaint about a late delivery" has roughly 40 well-worn ways of starting, you will see roughly those 40 openings whether you sample 500 times or 5,000.

You do not get diversity by sampling more from one prompt. You get it by sampling from more prompts.

Anatomy of a data-generation prompt

A prompt written for a chat assistant and a prompt written to manufacture training rows are different objects. The assistant prompt optimises for one good answer. The generation prompt optimises for a population of rows that will be parsed by code, filtered automatically, and used as supervision.

ComponentPurposeWhat goes wrong without it
RoleFixes register and vocabulary ("You are a support ticket triage dataset builder")Output drifts into assistant voice: helpful, polite, uniformly well-punctuated
Positive constraintsLength range, required fields, allowed label setRows vary from 8 to 800 words; labels appear that are not in your taxonomy
Non-goals"Do not resolve the issue. Do not include a signature. Do not mention brand names."The model helpfully adds things you must strip later, inconsistently
Schema instructionExact output shape, field names, typesProse wrapped around JSON; markdown fences; trailing commentary
Few-shot exemplarsDemonstrates tone, length and difficulty far better than descriptionThe model's default style dominates; hard cases never appear
Variation axisThe one slot that changes per call: persona, product, channel, mood500 near-copies

Few-shot exemplars: two rules people break

First, rotate them. If the same three examples appear in every call, the model anchors on them and your 500 rows are variations on three seeds. Sample three exemplars per call from a pool of thirty and the anchoring is spread across the pool.

Second, include the hard cases. If all your exemplars are clean, unambiguous complaints, the model will not invent the ticket that mixes a billing question with a delivery complaint and contains a typo in the order number — which is exactly the ticket your classifier fails on in production. One deliberately messy exemplar in three changes the difficulty distribution of the whole output.

Structured output that actually parses

"Return JSON" in the prompt gets you valid JSON maybe 92 percent of the time. The other 8 percent is markdown fences, a preamble, a trailing "Let me know if you'd like more examples", or a hallucinated extra field. At 80,000 rows, 8 percent is 6,400 parse failures.

Two mechanisms fix this at the API level rather than the prompt level. Constrained decoding against a JSON schema masks the token distribution at each step so only tokens that keep the output schema-valid can be sampled — invalid output becomes impossible, not unlikely. A tool marked strict makes the model emit arguments that match a declared input schema, which is the same idea wearing a different name.

Python
from pydantic import BaseModel, Fieldfrom typing import Literalclass SupportTicket(BaseModel):    text: str = Field(min_length=40, max_length=900)    intent: Literal["billing", "technical", "shipping", "account", "other"]    sentiment: Literal["angry", "frustrated", "neutral", "polite"]    channel: Literal["email", "chat", "phone_transcript"]    contains_order_id: bool# One schema object, reused three ways:#   1. constrained decoding at generation time#   2. validation gate after generation#   3. the contract your training pipeline readsSCHEMA = SupportTicket.model_json_schema()

Define the schema once and reuse the same object for generation, validation and consumption. The common bug is a schema that drifts: the prompt says sentiment, the validator checks tone, and 100 percent of rows silently fail a check nobody reads.

Structured output is not quality control

This is where people get it wrong. Constrained decoding guarantees shape, not truth and not usefulness. Every one of the following is schema-valid and worthless:

  • Text describing a billing problem, labelled intent: "shipping"
  • contains_order_id: true on text with no order ID in it
  • The same complaint, near-verbatim, for the 47th time
  • A complaint mentioning a refund policy the company does not have

Schema validation is the cheapest filter in the pipeline and it must be the first, not the only one.

The diversity problem, and the knobs that address it

Temperature, computed

Temperature rescales the logits before the softmax: pi=exp⁡(zi/T)/∑jexp⁡(zj/T)p_i = \exp(z_i/T) \big/ \sum_j \exp(z_j/T). Take three candidate tokens with logits 3.0, 2.0 and 1.0 and work it through.

TemperatureP(token A)P(token B)P(token C)Effect
0.50.8670.1170.016Near-deterministic; C is essentially unreachable
1.00.6650.2450.090The model's own distribution
1.50.5630.2890.148Flatter; C now appears once in seven

Notice the size of the effect. Halving temperature from 1.0 to 0.5 cut the third token's probability from 0.090 to 0.016 — a factor of 5.6. Raising it to 1.5 only lifted that token to 0.148, a factor of 1.6. Temperature is far more powerful at suppressing variety than at creating it, which is why cranking it to 1.4 to fix duplicate output mostly produces the same 40 modes plus occasional garbled tokens.

Top-p, computed

Nucleus sampling keeps the smallest set of tokens whose probabilities sum to at least p, then renormalises. With a sorted tail of 0.42, 0.21, 0.13, 0.08, 0.06, 0.04, 0.03, 0.02, 0.01:

Text
cumulative:  0.42  0.63  0.76  0.84  0.90  0.94  0.97  0.99  1.00top_p = 0.90 keeps the first 5 tokens, discards the restrenormalised (divide by 0.90):  0.467   0.233   0.144   0.089   0.067

Four tokens carrying 10 percent of the mass between them are removed entirely. That is the point — those are usually where incoherence comes from — but it is also why the outputs converge. Typical settings: temperature 0.9, top_p 0.95 for open-ended text; temperature 0.3, top_p 1.0 for the label field, if you can generate fields separately. Some current models, reasoning models in particular, ignore or reject temperature and top_p; on those, variety has to come from the prompt — which, as the next section shows, is where it works best anyway.

The knob that actually works: varying the prompt

The measured comparison, on 500 generated tickets each way:

StrategyDistinct-3 ratioExact/near duplicates
One prompt, T = 0.70.3118.4%
One prompt, T = 1.30.3911.7%
25 personas × 20 topics, T = 0.80.682.1%

Distinct-3 is the count of unique three-word sequences divided by the total count of three-word sequences; higher means less repetition. Raising temperature by 0.6 bought 0.08. Building a 500-cell grid of prompt variations bought 0.37, at lower temperature — so the text is also more coherent.

Python
import random, itertoolsPERSONAS = ["a terse power user", "a first-time buyer who is confused",            "a small-business owner in a hurry", "a polite but persistent retiree",            "someone typing on a phone with typos"]TOPICS   = ["a duplicate charge", "a package marked delivered but missing",            "a promo code that was rejected", "a subscription they cannot cancel"]CHANNELS = ["email", "live chat", "phone transcript"]LENGTHS  = ["under 30 words", "60-90 words", "over 200 words, rambling"]CELLS = list(itertools.product(PERSONAS, TOPICS, CHANNELS, LENGTHS))random.shuffle(CELLS)   # 5 * 4 * 3 * 3 = 180 distinct prompt cellsdef build_prompt(cell, exemplar_pool):    persona, topic, channel, length = cell    shots = random.sample(exemplar_pool, 3)     # rotate the few-shot set too    return PROMPT_TEMPLATE.format(        persona=persona, topic=topic, channel=channel,        length=length, examples="\n\n".join(shots))

Walk the cells in shuffled order rather than looping over the outer axis, so that a run interrupted at 40 percent still has balanced coverage rather than every persona being "a terse power user".

Diffusion models as image generators

A diffusion model is trained on a destructive process run backwards. Take a real image, add Gaussian noise in small steps until it is pure static, and train a network to predict the noise that was added at each step. To generate, start from pure static and repeatedly subtract the predicted noise. After 30 or so steps you have an image that was never photographed.

The parameters that matter

ParameterWhat it controlsPractical rangeFailure at the extreme
StepsHow many denoising iterations20–50 with a modern schedulerBelow 10: soft, mushy. Above 60: cost with no visible gain
Guidance scale (CFG)How hard the prompt pulls the image5–9Below 3: prompt ignored. Above 15: oversaturated, burnt edges, warped anatomy
SeedThe initial noise tensorRecord it, alwaysUnrecorded seed means the dataset is unreproducible
Negative promptA second conditioning the sampler pushes away from"blurry, watermark, text, extra fingers"Overloading it flattens the whole output distribution
SchedulerThe step-size schedule for denoisingDPM++ 2M, Euler a, DDIMMismatched scheduler and step count gives visible banding

These ranges suit classic Stable Diffusion-style models (SD 1.5, SDXL). Newer flow-matching models such as FLUX.1 and Stable Diffusion 3.5 use their own schedulers and much lower guidance — FLUX.1-dev defaults to a guidance scale of 3.5, and it ignores a negative prompt unless you enable true classifier-free guidance — so start from the settings on the model card.

Guidance is worth understanding rather than tuning blindly. At each step the model predicts noise twice — once conditioned on your prompt, once unconditioned — and the sampler uses ϵ^=ϵuncond+w (ϵcond−ϵuncond)\hat{\epsilon} = \epsilon_{uncond} + w\,(\epsilon_{cond} - \epsilon_{uncond}). With w = 1 the conditional term is used as-is. With w = 7.5 you are extrapolating 7.5 times past the conditional prediction, away from the unconditional one. That extrapolation is what makes images obey prompts, and it is also why high guidance produces over-contrasted, high-saturation pictures: you are pushing past the region the model was trained on.

Conditioning beyond the text prompt

For dataset work, text prompts are the weakest form of control. Four stronger ones:

  • img2img with a strength parameter. Start the denoising from a noised version of a real image rather than pure static. Strength 0.25 changes lighting and texture while keeping composition and — critically — keeping bounding boxes valid. Strength 0.8 keeps only the vague layout.
  • Inpainting. Regenerate only a masked region. This is how you place a rare defect onto 400 photographs of a real product: the background stays real, the defect varies, and you know the mask so you get the segmentation label for free.
  • Structural conditioning (ControlNet and similar). Condition on a depth map, edge map, segmentation mask or human pose. Because you supply the structure, you already hold the ground-truth annotation. Generate 5,000 images from 5,000 pose skeletons and you have 5,000 labelled keypoint examples with zero annotation cost.
  • Subject adaptation (LoRA, DreamBooth). Fine-tune a small adapter on 15–30 photographs of one specific object so the generator can place that object into new scenes. This is also the biggest memorisation risk in image synthesis — a LoRA trained on 20 photographs of a person will reproduce that person.

Generating a systematic grid, not a pile

Random prompts give random coverage. A dataset needs balanced coverage, which means enumerating the conditions you care about and generating a fixed count per cell.

Python
import itertoolsimport torchAXES = {    "defect":   ["hairline crack", "surface scratch", "discolouration", "none"],    "lighting": ["harsh overhead", "soft diffuse", "low warehouse light"],    "angle":    ["top-down", "45 degrees", "eye level"],    "material": ["matte plastic", "brushed metal"],}keys = list(AXES)cells = list(itertools.product(*(AXES[k] for k in keys)))   # 4*3*3*2 = 72 cellsPER_CELL = 40                                               # 2,880 images totalfor cell_idx, cell in enumerate(cells):    spec = dict(zip(keys, cell))    prompt = ("product photo of a {material} housing, {defect}, "              "{lighting}, shot {angle}").format(**spec)    for k in range(PER_CELL):        seed = cell_idx * 10_000 + k        # deterministic and reconstructable        image = pipe(prompt, negative_prompt=NEG, num_inference_steps=30,                     guidance_scale=7.0,                     generator=torch.Generator("cuda").manual_seed(seed)).images[0]        save(image, meta={**spec, "seed": seed, "model": MODEL_ID})

The meta dictionary is not optional bookkeeping. When the defect detector later shows 0.91 recall on cracks and 0.44 on discolouration, the per-cell metadata is what lets you find out that discolouration only rendered convincingly under soft diffuse lighting.

Choosing between them

DimensionLLM (text)Diffusion (images)
Label comes fromThe same generation call — ask for text and label togetherThe conditioning input — mask, pose, or the grid cell
Output validityEnforceable by constrained decodingNot enforceable; must be filtered visually or by a classifier
Typical unit costUSD 0.003–0.01 per row via APIUSD 0.002 self-hosted, USD 0.04 hosted
Dominant failureRepetition and hallucinated factsAnatomy, text-in-image, physically impossible geometry
ControllabilityVery high — you can state constraints in wordsLow from text alone, very high with structural conditioning
Human review speedSlow: minutes per row to read carefullyFast: a reviewer clears 400 thumbnails in an hour

That last row is quietly important for planning. Image review is roughly ten times faster per item than text review, so image pipelines can afford a much higher human-review sampling rate for the same budget.

Cost, throughput and licensing

Cost, worked

Text, 80,000 rows at 850 input and 260 output tokens, at the same example rates: 68.0 million input tokens at USD 3 per million is USD 204, and 20.8 million output tokens at USD 15 per million is USD 312 — about USD 516 in total. Routing generation to a cheaper model at, say, USD 0.80 and USD 4.00 per million drops that to USD 54 and USD 83, about USD 137. A batch endpoint at a 50 percent discount halves it again.

Images, 10,000 at 1024×1024, 30 steps, 3.8 seconds each on a GPU costing USD 2.20 per hour: 3.8 ÷ 3600 × 2.20 = USD 0.00232 per image, so USD 23.20 in compute — but 10,000 × 3.8 s = 10.6 hours on one GPU, or 1.3 hours across eight. The same 10,000 images through a hosted API at USD 0.04 each cost USD 400. Self-hosting wins on unit cost by 17× and loses on setup time; below about 5,000 images the API is almost always the right call.

Throughput and rate limits

Unit cost rarely decides the schedule; rate limits do. Suppose your account allows 400,000 output tokens per minute and each row uses 260. The ceiling is 400,000 ÷ 260 = 1,538 rows per minute, so 80,645 rows needs at least 52 minutes of pure generation. Real pipelines hit perhaps 40 percent of the ceiling once retries, backoff and long-tail latency are included, so budget around 2.2 hours. Practical measures:

  • Ask for k rows per call rather than one. Ten rows per response amortises the 850-token prompt across ten outputs, cutting input cost by about 90 percent — but rows within one response are more similar to each other, so keep k at 5–10 and never at 50.
  • Use asynchronous concurrency with a bounded semaphore, and back off on 429 rather than retrying immediately.
  • Checkpoint to disk every few hundred rows. A pipeline that loses two hours of generation to an unhandled exception is a pipeline that will lose it again.

Licensing and provenance

Three questions have to be answered before generated data enters a training set, and they are legal questions with engineering consequences.

QuestionWhy it bitesWhat to record
Do the provider's terms permit using outputs to train a model?Several commercial APIs restrict using outputs to build a competing model. Open-weight models under permissive licences generally do not.Provider, model ID, terms version, date of generation
Can the generator reproduce copyrighted or personal material?Image models memorise; a LoRA on 20 photographs of a person reproduces that personWhich real assets were used for adaptation, and a duplicate check against them
Can you prove which rows are synthetic later?Auditors, and your future self, will askA per-row provenance field, plus C2PA content credentials for images where supported

If you cannot answer "which model produced this row, with which prompt and seed, on which date" for every row in your dataset, you do not have a dataset — you have a pile.

What this means when you build

Design the variation axes before you write the prompt. The single highest-leverage decision in a generation pipeline is not the model, the temperature or the prompt wording — it is the list of dimensions along which rows are allowed to differ. Write that list down as data, count the cells, and decide how many rows per cell you need. If you cannot name at least three axes with four or more values each, you will produce near-duplicates no matter what else you do.

Generate the label in the same call as the content, but validate it separately. Asking a model to write a complaint and assign its intent in one structured response is cheap and mostly correct. Mostly is the operative word: expect 3–8 percent label noise, and plan a second pass where a different model, or a classifier trained on real data, re-labels a sample and you measure agreement.

Record everything at generation time, because you cannot reconstruct it afterwards. Model ID, prompt template version, axis values, sampling parameters, seed, timestamp. It costs a few hundred bytes per row. When the classifier underperforms on one slice three months later, that metadata is the difference between a two-hour diagnosis and a rebuild from scratch.