Advanced Prompting and Reasoning

Cognitive Patterns - Debate & Reflection


A team running a document-analysis pipeline scored 82 out of 100 on their evaluation set. So they added what looks free — a second pass appended to the same conversation:

Text
Now review your answer above. Identify any mistakes and providea corrected answer.

Accuracy fell to 80.

The per-question diffs showed the shape of it. Of the 82 correct answers, 6 became wrong. Of the 18 wrong ones, 4 became right. Net: minus two, plus a 60% increase in cost and latency.

The rates tell the opposite story: the critic repaired 4 of 18 wrong answers (a 22% repair rate) and damaged 6 of 82 correct ones (a 7.3% damage rate). It was three times more likely to help than to harm on any given answer, and still lost, because there were four and a half times more correct answers available to damage.

The break-even follows directly. With repair rate ff, damage rate dd and accuracy aa, the critic gains ground only when f(1−a)>d af(1-a) > d\,a, which rearranges to

a  <  ff+d  =  0.220.22+0.073  =  0.75a \;<\; \frac{f}{f + d} \;=\; \frac{0.22}{0.22 + 0.073} \;=\; 0.75

Below about 75% baseline accuracy this critic helps; above it, it hurts. The team was at 82%.

That threshold is not universal — ff and dd depend on task and critic prompt. What is universal: every self-critique step has both rates, nobody measures the damage rate, and base rates decide the outcome. The rest of this lesson is about making ff large and dd small.

Why one critique pass is nearly free and nearly uselessCritique in the same thread• Conditioned on the answer it is judging• Adds no information the context lacked• Usually returns 'looks correct'• Scored 82, still scores 82Debate in fresh contexts• One side argues for, one against• Neither has seen the other's reasoning• A judge reads both and decides• Costs three calls, finds real defects
Self-critique fails because the critic already believes the answer; independence, not more passes, is what surfaces the error.

Why naive self-critique fails

Two mechanical facts, and neither is about the model being "too agreeable".

The critique is conditioned on the answer

When the review happens in the same conversation, the original answer sits in the context. The model is predicting what follows a confident, fluent answer plus the phrase "review your answer". In its training distribution, that is overwhelmingly agreement or minor elaboration. The prior pulls towards "this is correct".

Instruct it the other way — "identify the mistakes" — and the prior flips just as hard. The most probable continuation of "Error 1:" is a description of an error, present or not. The model is not checking; it is completing a pattern whose shape you specified.

An unconditional instruction to find errors manufactures errors, and an unconditional instruction to review manufactures approval. In both cases the instruction determines the output more strongly than the content does.

Those 6 damaged answers are exactly this: told to find mistakes in a correct answer, the model found something — usually a defensible reinterpretation of the question — and "corrected" right into wrong.

The critique adds no information

The critique is sampled from the same weights, on the same prompt, with the same knowledge. If the model did not know the tax threshold when it answered, it does not know it when it reviews. Self-critique redistributes probability mass already there; it cannot import a fact.

This is the hard ceiling, and it explains every reliable result here: self-correction works to the extent that something outside the model's own generation enters the loop. A test suite, a compiler, a database, a retrieved document, a checklist constraining what counts as an answer — or, more subtly, a change in what the critique is conditioned on. Four levers:

LeverWhat it changesTypical effect
External verifier — tests, interpreter, database, searchInjects ground truth the model could not generateLarge and reliable
Fresh context — critique with the original reasoning hiddenRemoves the self-conditioning that produces agreementModerate
Role asymmetry — instruct a position, not a reviewConditions generation on advocacy rather than on the answerModerate
Decomposed checks — one narrow question at a timeSpends a whole forward pass per claim instead of one global impressionModerate, cheap to add

The last one works like this. "Is this answer correct?" gives one holistic judgement from a single pass over a long text. "Does the figure 4,320 in paragraph 3 follow from the two inputs in paragraph 2?" gives a whole pass to that one claim, in immediate context. Ten narrow checks recover errors one broad check reliably misses, and each is cheap.

The debate pattern

Debate applies role asymmetry deliberately. Instead of one model giving a balanced view, you run separate calls each committed to a position, then a third that adjudicates.

Asked for a balanced analysis, the model generates the modal discussion of the topic — the most typical considerations, usually the most obvious ones. Instructed to argue a position, generation is conditioned on advocacy, and the strongest available case becomes likely rather than unlikely. You steer into two regions of the distribution and compare what you find.

Worked example: deprecating an API

The question: should we deprecate the v1 API? The facts: v1 carries 12% of request volume, three enterprise customers are contractually on it until March, and maintaining both versions costs two engineer-months per quarter.

A single neutral prompt returns: "There are trade-offs. Deprecating reduces maintenance burden but risks customer disruption. Consider a phased migration with clear communication." True, generic, decides nothing.

Text
ADVOCATE A — argue FOR deprecating v1Two engineer-months per quarter against 12% of traffic is eightengineer-months a year — a third of an engineer, permanently,for a shrinking minority. Worse, the cost is not just time: everyv1-compatible feature has to be designed twice, so the tax landson new work, not on maintenance. Three enterprise contracts runto March, which is a hard date rather than an obstacle: announcenow, migrate them individually with engineering support, and endof life in April. Waiting does not shrink the 12% — nothingmigrates until there is a deadline.
Text
ADVOCATE B — argue AGAINST deprecating v1The 12% figure hides the distribution. If those three enterpriseaccounts are a large share of revenue, the exposure is measured incontract value, not request count — and that number is not onthe table. "Contractually until March" is a floor, not a ceiling:renewal conversations will start in January, and arriving with aforced migration is the wrong opening. Two engineer-months aquarter is also the cost of the status quo, not of deprecation.Deprecation itself costs migration tooling, support, documentationand at least one incident. Compare the real alternatives: freezev1 for new signups, stop building compatible features, and letattrition take the number down without a fight.

Neither is what a neutral pass produces. Advocate B surfaced two things that change the decision: request share is the wrong denominator when revenue is concentrated, and "deprecate" and "keep" are not the only options. That second point is debate's classic value — adversarial framing exposes the false dichotomy in the question.

Structuring it

Text
BADGive me the arguments for and against deprecating the v1 API,then tell me what you'd do.

One generation, so the "against" section is conditioned on the "for" section already written — a rebuttal shaped by the first half, not an independent case. The verdict then sits at the end of a long text, where the pull is towards a compromise splitting the difference. This is why single-pass "on one hand, on the other hand" analysis reliably lands on "consider a phased approach".

Text
GOOD  — three separate callsCALL 1 (advocate A, no knowledge of B)  You are arguing FOR deprecating the v1 API. Make the strongest  honest case. Use only the facts below; if a number you need is  missing, say which one and why it matters. Do not hedge, do not  present the other side. Maximum 200 words.  FACTS: [...]CALL 2 (advocate B, no knowledge of A) — mirror image, AGAINSTCALL 3 (judge, sees both cases, not the original question framing)  Two analysts have argued opposite positions on the same  decision. Evaluate them.  For each of these criteria, state which case is stronger and  quote the specific sentence that decides it:    1. Correct use of the quantitative evidence    2. Assumptions asserted without support    3. Options not considered by either side    4. Risks that are reversible vs irreversible  Then give a verdict, and state the single piece of missing  information that would most change it.  You may not introduce arguments neither analyst made.

Four constraints in the judge prompt do specific work. Per-criterion findings before the verdict conditions it on four completed analyses, not a general impression. Quote the deciding sentence anchors the judgement to text that exists — the cheapest guard against invented reasoning. Name the missing information converts a forced binary into something actionable; here, "what share of revenue do those three accounts represent" beats any verdict. No new arguments keeps the judge judging rather than becoming a third advocate.

Judge biases, and what to do about them

BiasMechanismMitigation
Position biasAttention is not uniform over position; which slot is favoured varies by model and prompt, but is rarely neutralRun the judge twice with the cases swapped. If the verdict flips, the debate was a tie; report it as one
Length biasMore text means more specific-sounding detail, and detail reads as evidenceImpose an identical word cap on both advocates and enforce it before judging
Confidence biasAssertive phrasing is more probable in text that turned out to be right, so it correlates with quality in training data but not in advocacyTell the judge that both sides were instructed not to hedge, so confidence carries no signal
Self-preferenceA model scores text closer to its own generation distribution more highlyUnavoidable when one model plays all roles; treat verdicts as evidence, not as findings

Debate plus repeated sampling

Combine the two: run the debate several times, alternating which advocate goes first, and count verdicts. Six runs — three with A first, three with B first — then read the pattern, not the total.

ResultInterpretationAction
6–0 for one sideGenuinely one-sided on the evidence givenAct on it
4–2, consistent across orderingsReal but contestedAct, and mitigate the losing side's strongest point
3–3, and the verdict always matches whichever side went lastYou have measured position bias, not meritDo not act on the verdict; get the missing information the judges named

That third row earns the extra calls. Without swapping the order you would have seen 3–3 and read it as a close call. It is no signal at all, and the difference is worth six calls on any decision that matters.

Reflection that actually finds things

Reflection means having the model examine its own output before it is used. What separates reflection that works from reflection that damages 7% of correct answers is how specific the questions are.

A concrete case

Python
def median(values):    values.sort()    n = len(values)    return values[n // 2]

Three real defects. It mutates the caller's list as a side effect. It is wrong for even-length input — median([1, 2, 3, 4]) returns 3, not 2.5. It raises IndexError on an empty list.

Text
BADReview this code and identify any issues.

Open-ended, so the output is whatever "code review comments" typically look like: a missing docstring, type hints, naming. All conditioned on the genre rather than on this function, and all three actual bugs missed. Worse, run it on a correct function and it still produces comments, because "no issues found" is a very improbable continuation of "identify any issues".

Text
GOODCheck this function against each item below. For each, answerPASS or FAIL. On FAIL, give a concrete input and the wrong outputit produces. If you cannot construct such an input, answer PASS.1. Does it modify any argument in place?2. Trace it by hand on [1, 2, 3, 4]. What exactly is returned?   Is that the median?3. Trace it on []. What happens?4. Trace it on a single-element list [5].5. Does it assume the input is already sorted, numeric, or non-null   anywhere without checking?

Every difference is mechanical. Each item is a narrow question, so the model spends a full pass on it with the relevant code in immediate context. "Trace by hand on [1, 2, 3, 4]" forces concrete execution — the model writes n = 4, n // 2 = 2, values[2] = 3, and the error appears in the tokens instead of needing to be spotted in one leap. Requiring a failing input to justify a FAIL is the anti-hallucination guard: inventing a plausible witness for a bug that does not exist is far harder than inventing the complaint. And "if you cannot construct such an input, answer PASS" licenses PASS explicitly, so it stops being an improbable continuation.

The single highest-value change you can make to any critique prompt: require the critic to produce a concrete witness for every defect it claims, and explicitly permit "no defect found".

A reusable self-questioning frame

When you cannot write task-specific checks, these five generalise, in descending order of yield:

  1. Which claims here are not supported by the input I was given? Quote each verbatim. (Catches fabrication.)
  2. Recompute every number independently. Show the arithmetic. (Catches the most common silent error.)
  3. What did the question ask that this answer does not address? (Catches partial answers, which read as complete.)
  4. What assumption, if false, would change the conclusion? Is it stated anywhere? (Catches hidden premises.)
  5. Is anything here inconsistent with anything else here? Quote both. (Catches drift in long outputs.)

Iterative refinement, and how it goes wrong

The loop is draft → critique → revise → repeat. Genuinely useful, and it has a near-universal failure mode.

An unbounded refinement loop on a product description, rubric score out of 10 each round:

RoundWordsScoreWhat the critic said
11806.0"Does not mention the integration story"
22407.5"Could be more specific about the target user"
33107.5"Would benefit from a concrete example"
44207.0"The opening is now buried"

Round 2 was a real improvement. Rounds 3 and 4 were not, and round 4 was a regression — the fix in round 3 caused the complaint in round 4. Meanwhile the text grew by 133%.

The cause is structural. A critic asked for feedback produces feedback, because "no changes needed" is an improbable continuation of a request for critique. A reviser given feedback acts on it, because ignoring instructions is improbable too. Neither side has a stopping incentive, so the loop random-walks over defensible edits — and since additions are easier to justify than deletions, it drifts towards longer.

Bounding it

Python
def refine(task, draft, rubric, max_rounds=3, target=8.0):    history = [(0, draft, None)]    for rnd in range(1, max_rounds + 1):        critique = call(f"""Score this draft against the rubric. Output in this exact order:FINDINGS: up to 3 items. Each must name a specific rubric criterion  and quote the exact text at fault. If the draft meets every  criterion, write "none".SCORE: 0-10VERDICT: one of {{revise, accept}}. Answer accept if the score is  {target} or above, or if your findings are stylistic preferences  rather than rubric failures.RUBRIC:{rubric}DRAFT:{draft}""")        score, verdict, findings = parse(critique)        if verdict == "accept" or score >= target:            return draft, score, rnd, "target met"        if history and score <= history[-1][0]:            return history[-1][1], history[-1][0], rnd, "score regressed; kept previous draft"        draft = call(f"""Revise the draft to fix ONLY the findings listed. Constraints:  - do not exceed {int(len(draft.split()) * 1.1)} words  - do not change anything the findings did not mention  - if a finding cannot be fixed without breaking another rubric    criterion, leave it and say so in one line at the endFINDINGS:{findings}DRAFT:{draft}""")        history.append((score, draft, findings))    return draft, score, max_rounds, "round limit reached"

Five guards, each aimed at a named failure:

  • "If the draft meets every criterion, write none" — makes the empty critique a licensed output.
  • An explicit accept verdict — gives the critic a way to stop the loop, which it otherwise lacks.
  • A word ceiling of 110% — caps growth drift. Without it, expect 30–40% growth per round.
  • "Fix only the findings" — a reviser given free rein rewrites things that were fine, which is where regressions come from.
  • Regression detection — if the score drops, return the previous draft. Refinement must be monotonic in the metric or it is not refinement.

Self-correction against something real

Everything above softens the damage rate. Raising the repair rate needs an external verifier.

Take code generation on 100 tasks, first-attempt pass rate 61%.

LoopRound 1Round 2Why
"Review your code and fix any bugs" (no execution)60%59%The critique comes from the same weights that wrote the bug; damage rate slightly exceeds repair rate
Run the tests, feed failures back verbatim74%78%A failing assertion is information the model did not have and could not have produced

Same model, same task, same number of calls. The only difference is whether a real interpreter ran. Note too that the verified loop's second round adds just 4 points — what survives two rounds of real tests is mostly misunderstanding of the requirement, which more attempts cannot fix.

Python
def solve_with_tests(spec, tests, max_attempts=3):    code = call(f"Write a function for this specification.\n\n{spec}")    for attempt in range(max_attempts):        result = run_tests(code, tests)          # real execution, in a sandbox        if result.passed:            return code, attempt + 1, "passed"        code = call(f"""This implementation fails its tests.CODE:{code}TEST OUTPUT (verbatim — this is ground truth, not an opinion):{result.output}Fix the code so these tests pass. Before writing the fix, state inone sentence what the failing assertion proves about the currentbehaviour. Do not change the function signature.""")    return code, max_attempts, "still failing"

"This is ground truth, not an opinion" is not decoration. Without it, models argue with test output, explaining why the test is wrong, because disagreement is a plausible continuation of a criticism. Marking it authoritative shifts that continuation towards acceptance. And "state what the failing assertion proves before writing the fix" forces diagnosis before the first token of the patch.

Choosing the right check for the error

Different error types need different machinery, and using a model where code would do is the commonest waste here.

Error typeExampleBest checkerWhy not a model critique
Arithmetic"1,240 × 0.15 = 201"Evaluate the expression in codeThe critique needs the same multiplication the answer got wrong
Constraint violationPlan exceeds the stated budgetAssertion in codeDeterministic and free; a model call is strictly worse
Unsupported claimCites a policy clause that does not existRetrieval, then string match against the sourceThe model invented it once and will happily confirm it
Internal contradictionParagraph 2 says 30 days, paragraph 6 says 45Model, given a narrow "quote both statements" instructionNeeds language understanding — but only with a specific instruction
Scope driftAnswered a related question, not the one askedModel, comparing answer against requirements one by oneWorks, but only as an itemised checklist, not "did I answer the question?"
Stale assumptionUses last year's tax bandRetrieval against a dated sourceThe model has no way to know what today's value is

What this means when you build something

Treat every critique step as a component with two measurable properties, and refuse to ship one you have not measured.

Build a small set of cases where you know the right answer. Run the pipeline with the critic and without it. Record four numbers: correct answers damaged, wrong answers repaired, added cost, added latency. That gives dd, ff, and the price of both. Compute f/(f+d)f / (f + d) and compare it to your accuracy. If accuracy is already above that ratio, the critic is a liability no matter how sensible its comments read.

Then apply the routing rule:

  • An external verifier exists — tests, a schema, a database, a calculator, a retrievable source. Use it, feeding its raw output back verbatim and marked as authoritative. This is the only configuration with reliably large gains, and worth engineering a verifier where none exists.
  • No verifier, but a decision with real consequences. Use debate: separate calls, swapped ordering, symmetric length caps, a judge that must quote its evidence. Expect the value to come less from the verdict than from the missing information it names.
  • No verifier, routine output. Use decomposed checks with concrete witnesses required, or nothing. Never "review your answer and fix any mistakes" — that is the configuration measured above, and it cost the team two points of accuracy and 60% of their budget.

One last discipline: keep critique and revision in separate calls, and keep the critic's view of the original reasoning as narrow as the task allows. Every token of the original answer the critic sees pulls its output towards agreement. That is not a flaw you can prompt your way out of — it is what conditioning on context means — so control the context.