Synthetic Data Generation

Multi-Turn and Context-Aware Synthesis


Six hundred synthetic support dialogues, generated overnight, read beautifully one turn at a time. Then a reviewer read one end to end. At turn 3 the agent says order A-4471 shipped on Tuesday and gives a tracking number. At turn 7 the same agent says the order has not shipped yet and offers to expedite it. Neither turn is wrong on its own. Together they are nonsense.

An automated check across all 600 dialogues found that 247 of them — 41.2 percent — contained at least one contradiction. And the dialogues that survived had a different problem: they averaged 5.2 turns against 8.4 for real conversations, because the synthetic customer accepted the first answer every time. Not one of the 600 contained a customer who misunderstood the agent, which happens in roughly a third of real support conversations.

Both problems come from the same source. Generating a conversation is not generating a longer document. It is simulating two parties who know different things, want different things, and are not both trying to be helpful.

One dialogue, read end to endturn 1turn 2turn 3turn 4turn 5turn 6turn 7nullshippedTuesdaynot yetshippedAt 97 percent consistency per turn, an 8-turn dialogue is fully coherent about 78 percent of the time.
Errors compound multiplicatively along the chain, and every turn reads perfectly on its own — only the pair contradicts.

Why one long generation is not a conversation

Errors compound multiplicatively

Every turn is a fresh opportunity to contradict something already established. If the per-turn probability of a contradiction is p, the probability that a dialogue of n turns is clean is (1−p)n(1-p)^n.

Per-turn error rate4 turns8 turns12 turns20 turns
0.027.8%14.9%21.5%33.2%
0.0415.1%27.9%38.7%55.8%
0.0828.4%48.7%63.2%81.1%

Those are probabilities that the dialogue contains at least one contradiction. Work through the 0.04 column at 8 turns: 0.968=0.72140.96^8 = 0.7214, so 27.9 percent of dialogues are broken. A per-turn error rate of 4 percent sounds like a good model. At conversation length it is a 28 percent scrap rate — and at 12 turns with a slightly weaker prompt, 63 percent.

This is why single-turn quality metrics mislead so badly here. You cannot fix a multi-turn pipeline by improving average turn quality from 96 to 97 percent; you fix it by making the per-turn error rate structurally small, which means giving the generator something to check against.

The single-narrator problem

Ask one model to write a whole dialogue and it writes a scene, not an interaction. It knows both sides' intentions, so the customer asks exactly the question the agent is ready to answer, and the agent's answer is exactly what the customer needed. Real conversations are full of friction precisely because each party lacks the other's information.

The measurable signature is in the turn-count distribution:

PropertyReal dialoguesSingle-narrator syntheticTwo-agent synthetic
Mean turns8.45.28.1
Standard deviation4.11.33.8
Over 12 turns22%0%19%
Contains a misunderstanding34%0%29%
KS statistic on turn count—0.380.07

A standard deviation of 1.3 against a real 4.1 is the tell. The single-narrator generator has one idea of what a support conversation is and produces it every time.

If one model can see both sides' knowledge, it will never generate the misunderstanding — and misunderstandings are most of what a support model needs to learn to handle.

Context cost grows quadratically

Each turn's prompt contains all previous turns. With a 400-token system prompt, 120 tokens per turn, and an 80-token state block, the input tokens for a 12-turn dialogue are:

Text
per call k:   400 (system) + 80 (state) + 120 * (k - 1) (history)sum over k = 1..12:   system+state:  12 * 480                 =  5,760   history:       120 * (0+1+...+11)                = 120 * 66                 =  7,920                                    total  = 13,680 input tokens   output:        12 * 120                 =  1,440 output tokens

At example rates of USD 3 per million input and USD 15 per million output tokens, that is USD 0.041 + USD 0.022 = USD 0.063 per dialogue, against about USD 0.023 for a single-shot generation of the same content — roughly 2.7× the cost. Acceptable. But the growth is quadratic in turns: at 40 turns the same arithmetic gives 112,800 input tokens — 8.2 times the total for only 3.3 times the turns. Two mitigations, both worth having by default:

  • Cap the verbatim window. Keep the last four turns in full and replace everything older with a running 60-token summary. Input per call becomes roughly constant at about 1,020 tokens instead of growing.
  • Cache the static prefix. System prompt, persona and few-shot examples are identical across every call in a dialogue. Prompt caching typically bills cached input at around a tenth of the normal rate, which removes most of the fixed cost.

The two-agent self-play pattern

Run two separately-prompted models, each with private information the other cannot see, and let them alternate. The customer knows the problem, the emotional state, and what would satisfy them. The agent knows the policies, the tools, and the account record. Neither sees the other's brief.

Python
CUSTOMER_SYS = """You are a customer contacting support. You are NOT an assistant.YOUR SITUATION (the agent does not know this):  problem: {problem}  order: {order_id}, placed {order_date}  what would satisfy you: {resolution_wanted}  mood: {mood}  technical skill: {skill}RULES- Reveal information gradually. Do not dump every detail in turn one.- If the agent's answer is vague, push back. Do not thank them for a non-answer.- Never speak as the agent. Never solve your own problem.- Write one message, 10-70 words, in your own register. No signature.- If you are satisfied, say so plainly and stop."""AGENT_SYS = """You are a support agent for Orbit Networks.YOU MAY USE ONLY THESE FACTS: {policy_context}ACCOUNT RECORD VISIBLE TO YOU: {account_record}RULES- Never invent a policy, refund window, price or delivery date.- If a fact is not in your context, say you need to check and ask a question.- Everything you have already told the customer is listed in STATE. Never  contradict it.- One message, 15-90 words."""

The line "You are NOT an assistant" earns its place. Without it, the customer model drifts into helpfulness within three turns and starts suggesting solutions to its own problem — the most common and most obvious tell in synthetic dialogue data.

Keeping state: the scratchpad

The contradiction problem is a memory problem. The model has the transcript, but salient facts are buried in prose and get lost in attention. The fix is to maintain an explicit structured state, update it after every turn, and inject it into every prompt as facts rather than as narrative.

Python
state = {    "turn": 0,    "facts_asserted_by_agent": [],   # "order A-4471 shipped Tuesday 14 May"    "facts_asserted_by_customer": [],    "commitments": [],               # "refund of 42.00 to be issued in 3 days"    "open_questions": [],            # asked but not yet answered    "resolution_status": "open",     # open | partial | resolved | escalated    "customer_mood": "frustrated",   # updated as the dialogue moves}def run_dialogue(scenario, max_turns=14):    transcript, state = [], init_state(scenario)    for t in range(max_turns):        speaker = "customer" if t % 2 == 0 else "agent"        sys = CUSTOMER_SYS if speaker == "customer" else AGENT_SYS        msg = generate(            system=sys.format(**scenario),            history=window(transcript, keep=4, summarise_older=True),            state_block=render_state(state),            directive=dynamics_for_turn(t, scenario))        transcript.append({"role": speaker, "text": msg})        state = update_state(state, speaker, msg)   # a cheap extraction call        if state["resolution_status"] in ("resolved", "escalated"):            break    return transcript, state

The state update is a separate, small model call that extracts new commitments and asserted facts from the turn just produced. It costs perhaps 150 input and 60 output tokens — around 15 percent added to the dialogue cost. The measured return:

ConfigurationPer-turn contradiction rateDialogues clean at 8 turns
Full transcript only0.04072.1%
Transcript + state block0.00993.0%

0.9918=0.9300.991^8 = 0.930, so the scrap rate falls from 27.9 percent to 7.0 percent. Fifteen percent more cost to recover a fifth of your dataset is not a close decision.

Give the generator a short list of facts it must not contradict, and it will mostly not contradict them. Give it a long transcript and hope, and it will.

Injecting realistic dynamics

Clean, cooperative dialogues are the ones a model handles already. The value in synthetic conversation data is in the awkward ones, and awkwardness has to be scheduled deliberately because neither agent will produce it spontaneously.

DynamicWhat it teaches the downstream modelHow to inject itReal-corpus rate
Gradual disclosureAsk clarifying questions instead of guessingWithhold two scenario fields until turn 3+~always
MisunderstandingDetect and recover from being misreadDirective: "misread the agent's last message as being about billing"34%
CorrectionUpdate state when the user retracts somethingDirective: "you gave the wrong order number earlier; correct it now"18%
Topic shiftHandle a second issue without dropping the firstDirective: "raise an unrelated billing question mid-thread"21%
Escalation arcRecognise rising frustration and change registerStep mood: neutral → frustrated → angry on a schedule27%
InterruptionCope with a turn that ignores the question askedDirective: "ignore the agent's question, repeat your demand"14%
AbandonmentRecognise a conversation that never resolvesTerminate at a random turn with no resolution9%

Sample dynamics from these measured rates rather than applying them uniformly, or your synthetic corpus will be more chaotic than reality — a different distributional error with the same cause.

Python
DYNAMICS = [("misunderstanding", 0.34), ("correction", 0.18),            ("topic_shift", 0.21), ("escalation", 0.27),            ("interruption", 0.14), ("abandonment", 0.09)]def plan_dynamics(n_turns, rng):    """Decide up front which turns carry which dynamic."""    plan = {}    for name, rate in DYNAMICS:        if rng.random() < rate:            # never on turn 0 or the final turn            plan[rng.randrange(1, max(2, n_turns - 1))] = name    return plan

Planning the schedule before the dialogue starts, rather than deciding turn by turn, keeps the arc coherent — an escalation that begins at turn 2 and resolves at turn 9 needs to be known at turn 2.

One more knob: sample the target turn count for each dialogue from a distribution fitted to the real corpus before generation begins, and put it in the customer's brief ("you will be satisfied at around turn 9"). This alone moved the KS statistic on turn-count from 0.38 to 0.07 in the comparison table above.

Validating multi-turn coherence

Row-level filters do not see the failure. You need checks that read the whole dialogue.

Python
def check_dialogue(transcript, state, scenario):    issues = []    # 1. Entity consistency: every ID, date and amount mentioned more than    #    once must be mentioned identically.    for kind, values in extract_entities(transcript).items():        if len(set(values)) > 1:            issues.append(("entity", kind, sorted(set(values))))    # 2. Temporal consistency: no event dated before the order date, no    #    "already shipped" followed by "not yet shipped".    issues += check_timeline(transcript, scenario["order_date"])    # 3. Resolution consistency: if the last turn claims resolution, the    #    customer's stated goal must actually have been addressed.    if state["resolution_status"] == "resolved":        if not goal_addressed(transcript, scenario["resolution_wanted"]):            issues.append(("premature_resolution", None, None))    # 4. Persona drift: the customer must not start solving the problem,    #    and register must not converge on the agent's.    issues += check_persona(transcript, scenario)    # 5. Structural: alternating speakers, no empty turns, length in range.    issues += check_structure(transcript)    return issues

Run over the original 600 dialogues:

Issue categoryDialogues affectedShare of the 600
Entity contradiction (ID, date, amount)11819.7%
Temporal impossibility6110.2%
Premature resolution447.3%
Persona drift244.0%
Total flagged24741.2%

Repair beats discard

A contradiction at turn 7 does not invalidate turns 1 to 6. Truncate to the last good turn, append the violated constraint as an explicit instruction, and regenerate forward. Of the 247 flagged dialogues, 189 (76.5 percent) passed on the second attempt and 58 were discarded after two failures. Discarding all 247 would have thrown away 41 percent of a run that cost real money; repair recovered three quarters of it.

Cap repairs at two attempts. A dialogue that fails three times is telling you the scenario itself is incoherent — usually a policy context that cannot satisfy the customer's stated goal — and the fix belongs in the scenario generator.

The check no automation performs

Everything above verifies internal consistency. None of it verifies that the agent's statements are true of your business. A dialogue in which the agent smoothly explains a 45-day returns window is perfectly coherent and, if your window is 30 days, perfectly poisonous — and a model fine-tuned on 5,000 such dialogues will confidently tell real customers the wrong policy. The mitigations are a hard fact base in the agent's context, a regex sweep for policy-shaped claims, and a domain expert reading a stratified sample.

What this means when you build

Separate the scenario generator from the dialogue generator. The scenario — problem, account record, policy context, target length, planned dynamics, desired resolution — is cheap, structured, and enumerable across the axes you care about. Generate 5,000 scenarios first, check their coverage as a table, and only then run dialogues. Debugging a coverage gap in a table of scenarios takes minutes; finding it inside 5,000 transcripts takes days.

Store the state object alongside the transcript in your final dataset. It is a free, machine-readable annotation: which facts were asserted, what was committed, whether it resolved. That turns a pile of conversations into a supervised dataset for state tracking, commitment extraction and escalation prediction, at no extra generation cost.

Validate at the dialogue level from the first day, not after the first big run. The 41 percent contradiction rate was discoverable on a sample of 20 dialogues in the first hour. Every multi-turn pipeline needs its coherence checker written before its first full run, because a per-turn error rate small enough to be invisible in a spot check is large enough to break most of your conversations.