Course Content
Advanced Prompting and Reasoning
3 sections · 7 lessons
Cognitive Patterns - Debate & Reflection
A team running a document-analysis pipeline scored 82 out of 100 on their evaluation set. So they added what looks free — a second pass appended to the same conversation:
Now review your answer above. Identify any mistakes and providea corrected answer.Accuracy fell to 80.
The per-question diffs showed the shape of it. Of the 82 correct answers, 6 became wrong. Of the 18 wrong ones, 4 became right. Net: minus two, plus a 60% increase in cost and latency.
The rates tell the opposite story: the critic repaired 4 of 18 wrong answers (a 22% repair rate) and damaged 6 of 82 correct ones (a 7.3% damage rate). It was three times more likely to help than to harm on any given answer, and still lost, because there were four and a half times more correct answers available to damage.
The break-even follows directly. With repair rate f, damage rate d and accuracy a, the critic gains ground only when f(1−a)>da, which rearranges to
Below about 75% baseline accuracy this critic helps; above it, it hurts. The team was at 82%.
That threshold is not universal — f and d depend on task and critic prompt. What is universal: every self-critique step has both rates, nobody measures the damage rate, and base rates decide the outcome. The rest of this lesson is about making f large and d small.
Why naive self-critique fails
Two mechanical facts, and neither is about the model being "too agreeable".
The critique is conditioned on the answer
When the review happens in the same conversation, the original answer sits in the context. The model is predicting what follows a confident, fluent answer plus the phrase "review your answer". In its training distribution, that is overwhelmingly agreement or minor elaboration. The prior pulls towards "this is correct".
Instruct it the other way — "identify the mistakes" — and the prior flips just as hard. The most probable continuation of "Error 1:" is a description of an error, present or not. The model is not checking; it is completing a pattern whose shape you specified.
An unconditional instruction to find errors manufactures errors, and an unconditional instruction to review manufactures approval. In both cases the instruction determines the output more strongly than the content does.
Those 6 damaged answers are exactly this: told to find mistakes in a correct answer, the model found something — usually a defensible reinterpretation of the question — and "corrected" right into wrong.
The critique adds no information
The critique is sampled from the same weights, on the same prompt, with the same knowledge. If the model did not know the tax threshold when it answered, it does not know it when it reviews. Self-critique redistributes probability mass already there; it cannot import a fact.
This is the hard ceiling, and it explains every reliable result here: self-correction works to the extent that something outside the model's own generation enters the loop. A test suite, a compiler, a database, a retrieved document, a checklist constraining what counts as an answer — or, more subtly, a change in what the critique is conditioned on. Four levers:
| Lever | What it changes | Typical effect |
|---|---|---|
| External verifier — tests, interpreter, database, search | Injects ground truth the model could not generate | Large and reliable |
| Fresh context — critique with the original reasoning hidden | Removes the self-conditioning that produces agreement | Moderate |
| Role asymmetry — instruct a position, not a review | Conditions generation on advocacy rather than on the answer | Moderate |
| Decomposed checks — one narrow question at a time | Spends a whole forward pass per claim instead of one global impression | Moderate, cheap to add |
The last one works like this. "Is this answer correct?" gives one holistic judgement from a single pass over a long text. "Does the figure 4,320 in paragraph 3 follow from the two inputs in paragraph 2?" gives a whole pass to that one claim, in immediate context. Ten narrow checks recover errors one broad check reliably misses, and each is cheap.
The debate pattern
Debate applies role asymmetry deliberately. Instead of one model giving a balanced view, you run separate calls each committed to a position, then a third that adjudicates.
Asked for a balanced analysis, the model generates the modal discussion of the topic — the most typical considerations, usually the most obvious ones. Instructed to argue a position, generation is conditioned on advocacy, and the strongest available case becomes likely rather than unlikely. You steer into two regions of the distribution and compare what you find.
Worked example: deprecating an API
The question: should we deprecate the v1 API? The facts: v1 carries 12% of request volume, three enterprise customers are contractually on it until March, and maintaining both versions costs two engineer-months per quarter.
A single neutral prompt returns: "There are trade-offs. Deprecating reduces maintenance burden but risks customer disruption. Consider a phased migration with clear communication." True, generic, decides nothing.
ADVOCATE A — argue FOR deprecating v1Two engineer-months per quarter against 12% of traffic is eightengineer-months a year — a third of an engineer, permanently,for a shrinking minority. Worse, the cost is not just time: everyv1-compatible feature has to be designed twice, so the tax landson new work, not on maintenance. Three enterprise contracts runto March, which is a hard date rather than an obstacle: announcenow, migrate them individually with engineering support, and endof life in April. Waiting does not shrink the 12% — nothingmigrates until there is a deadline.ADVOCATE B — argue AGAINST deprecating v1The 12% figure hides the distribution. If those three enterpriseaccounts are a large share of revenue, the exposure is measured incontract value, not request count — and that number is not onthe table. "Contractually until March" is a floor, not a ceiling:renewal conversations will start in January, and arriving with aforced migration is the wrong opening. Two engineer-months aquarter is also the cost of the status quo, not of deprecation.Deprecation itself costs migration tooling, support, documentationand at least one incident. Compare the real alternatives: freezev1 for new signups, stop building compatible features, and letattrition take the number down without a fight.Neither is what a neutral pass produces. Advocate B surfaced two things that change the decision: request share is the wrong denominator when revenue is concentrated, and "deprecate" and "keep" are not the only options. That second point is debate's classic value — adversarial framing exposes the false dichotomy in the question.
Structuring it
BADGive me the arguments for and against deprecating the v1 API,then tell me what you'd do.One generation, so the "against" section is conditioned on the "for" section already written — a rebuttal shaped by the first half, not an independent case. The verdict then sits at the end of a long text, where the pull is towards a compromise splitting the difference. This is why single-pass "on one hand, on the other hand" analysis reliably lands on "consider a phased approach".
GOOD — three separate callsCALL 1 (advocate A, no knowledge of B) You are arguing FOR deprecating the v1 API. Make the strongest honest case. Use only the facts below; if a number you need is missing, say which one and why it matters. Do not hedge, do not present the other side. Maximum 200 words. FACTS: [...]CALL 2 (advocate B, no knowledge of A) — mirror image, AGAINSTCALL 3 (judge, sees both cases, not the original question framing) Two analysts have argued opposite positions on the same decision. Evaluate them. For each of these criteria, state which case is stronger and quote the specific sentence that decides it: 1. Correct use of the quantitative evidence 2. Assumptions asserted without support 3. Options not considered by either side 4. Risks that are reversible vs irreversible Then give a verdict, and state the single piece of missing information that would most change it. You may not introduce arguments neither analyst made.Four constraints in the judge prompt do specific work. Per-criterion findings before the verdict conditions it on four completed analyses, not a general impression. Quote the deciding sentence anchors the judgement to text that exists — the cheapest guard against invented reasoning. Name the missing information converts a forced binary into something actionable; here, "what share of revenue do those three accounts represent" beats any verdict. No new arguments keeps the judge judging rather than becoming a third advocate.
Judge biases, and what to do about them
| Bias | Mechanism | Mitigation |
|---|---|---|
| Position bias | Attention is not uniform over position; which slot is favoured varies by model and prompt, but is rarely neutral | Run the judge twice with the cases swapped. If the verdict flips, the debate was a tie; report it as one |
| Length bias | More text means more specific-sounding detail, and detail reads as evidence | Impose an identical word cap on both advocates and enforce it before judging |
| Confidence bias | Assertive phrasing is more probable in text that turned out to be right, so it correlates with quality in training data but not in advocacy | Tell the judge that both sides were instructed not to hedge, so confidence carries no signal |
| Self-preference | A model scores text closer to its own generation distribution more highly | Unavoidable when one model plays all roles; treat verdicts as evidence, not as findings |
Debate plus repeated sampling
Combine the two: run the debate several times, alternating which advocate goes first, and count verdicts. Six runs — three with A first, three with B first — then read the pattern, not the total.
| Result | Interpretation | Action |
|---|---|---|
| 6–0 for one side | Genuinely one-sided on the evidence given | Act on it |
| 4–2, consistent across orderings | Real but contested | Act, and mitigate the losing side's strongest point |
| 3–3, and the verdict always matches whichever side went last | You have measured position bias, not merit | Do not act on the verdict; get the missing information the judges named |
That third row earns the extra calls. Without swapping the order you would have seen 3–3 and read it as a close call. It is no signal at all, and the difference is worth six calls on any decision that matters.
Reflection that actually finds things
Reflection means having the model examine its own output before it is used. What separates reflection that works from reflection that damages 7% of correct answers is how specific the questions are.
A concrete case
1def median(values):2 values.sort()3 n = len(values)4 return values[n // 2]Three real defects. It mutates the caller's list as a side effect. It is wrong for even-length input — median([1, 2, 3, 4]) returns 3, not 2.5. It raises IndexError on an empty list.
BADReview this code and identify any issues.Open-ended, so the output is whatever "code review comments" typically look like: a missing docstring, type hints, naming. All conditioned on the genre rather than on this function, and all three actual bugs missed. Worse, run it on a correct function and it still produces comments, because "no issues found" is a very improbable continuation of "identify any issues".
GOODCheck this function against each item below. For each, answerPASS or FAIL. On FAIL, give a concrete input and the wrong outputit produces. If you cannot construct such an input, answer PASS.1. Does it modify any argument in place?2. Trace it by hand on [1, 2, 3, 4]. What exactly is returned? Is that the median?3. Trace it on []. What happens?4. Trace it on a single-element list [5].5. Does it assume the input is already sorted, numeric, or non-null anywhere without checking?Every difference is mechanical. Each item is a narrow question, so the model spends a full pass on it with the relevant code in immediate context. "Trace by hand on [1, 2, 3, 4]" forces concrete execution — the model writes n = 4, n // 2 = 2, values[2] = 3, and the error appears in the tokens instead of needing to be spotted in one leap. Requiring a failing input to justify a FAIL is the anti-hallucination guard: inventing a plausible witness for a bug that does not exist is far harder than inventing the complaint. And "if you cannot construct such an input, answer PASS" licenses PASS explicitly, so it stops being an improbable continuation.
The single highest-value change you can make to any critique prompt: require the critic to produce a concrete witness for every defect it claims, and explicitly permit "no defect found".
A reusable self-questioning frame
When you cannot write task-specific checks, these five generalise, in descending order of yield:
- Which claims here are not supported by the input I was given? Quote each verbatim. (Catches fabrication.)
- Recompute every number independently. Show the arithmetic. (Catches the most common silent error.)
- What did the question ask that this answer does not address? (Catches partial answers, which read as complete.)
- What assumption, if false, would change the conclusion? Is it stated anywhere? (Catches hidden premises.)
- Is anything here inconsistent with anything else here? Quote both. (Catches drift in long outputs.)
Iterative refinement, and how it goes wrong
The loop is draft → critique → revise → repeat. Genuinely useful, and it has a near-universal failure mode.
An unbounded refinement loop on a product description, rubric score out of 10 each round:
| Round | Words | Score | What the critic said |
|---|---|---|---|
| 1 | 180 | 6.0 | "Does not mention the integration story" |
| 2 | 240 | 7.5 | "Could be more specific about the target user" |
| 3 | 310 | 7.5 | "Would benefit from a concrete example" |
| 4 | 420 | 7.0 | "The opening is now buried" |
Round 2 was a real improvement. Rounds 3 and 4 were not, and round 4 was a regression — the fix in round 3 caused the complaint in round 4. Meanwhile the text grew by 133%.
The cause is structural. A critic asked for feedback produces feedback, because "no changes needed" is an improbable continuation of a request for critique. A reviser given feedback acts on it, because ignoring instructions is improbable too. Neither side has a stopping incentive, so the loop random-walks over defensible edits — and since additions are easier to justify than deletions, it drifts towards longer.
Bounding it
1def refine(task, draft, rubric, max_rounds=3, target=8.0):2 history = [(0, draft, None)]3 for rnd in range(1, max_rounds + 1):4 critique = call(f"""5Score this draft against the rubric. Output in this exact order:67FINDINGS: up to 3 items. Each must name a specific rubric criterion8 and quote the exact text at fault. If the draft meets every9 criterion, write "none".10SCORE: 0-1011VERDICT: one of {{revise, accept}}. Answer accept if the score is12 {target} or above, or if your findings are stylistic preferences13 rather than rubric failures.1415RUBRIC:16{rubric}1718DRAFT:19{draft}""")2021 score, verdict, findings = parse(critique)22 if verdict == "accept" or score >= target:23 return draft, score, rnd, "target met"24 if history and score <= history[-1][0]:25 return history[-1][1], history[-1][0], rnd, "score regressed; kept previous draft"2627 draft = call(f"""28Revise the draft to fix ONLY the findings listed. Constraints:29 - do not exceed {int(len(draft.split()) * 1.1)} words30 - do not change anything the findings did not mention31 - if a finding cannot be fixed without breaking another rubric32 criterion, leave it and say so in one line at the end3334FINDINGS:35{findings}3637DRAFT:38{draft}""")39 history.append((score, draft, findings))4041 return draft, score, max_rounds, "round limit reached"Five guards, each aimed at a named failure:
- "If the draft meets every criterion, write none" — makes the empty critique a licensed output.
- An explicit
acceptverdict — gives the critic a way to stop the loop, which it otherwise lacks. - A word ceiling of 110% — caps growth drift. Without it, expect 30–40% growth per round.
- "Fix only the findings" — a reviser given free rein rewrites things that were fine, which is where regressions come from.
- Regression detection — if the score drops, return the previous draft. Refinement must be monotonic in the metric or it is not refinement.
Self-correction against something real
Everything above softens the damage rate. Raising the repair rate needs an external verifier.
Take code generation on 100 tasks, first-attempt pass rate 61%.
| Loop | Round 1 | Round 2 | Why |
|---|---|---|---|
| "Review your code and fix any bugs" (no execution) | 60% | 59% | The critique comes from the same weights that wrote the bug; damage rate slightly exceeds repair rate |
| Run the tests, feed failures back verbatim | 74% | 78% | A failing assertion is information the model did not have and could not have produced |
Same model, same task, same number of calls. The only difference is whether a real interpreter ran. Note too that the verified loop's second round adds just 4 points — what survives two rounds of real tests is mostly misunderstanding of the requirement, which more attempts cannot fix.
1def solve_with_tests(spec, tests, max_attempts=3):2 code = call(f"Write a function for this specification.\n\n{spec}")3 for attempt in range(max_attempts):4 result = run_tests(code, tests) # real execution, in a sandbox5 if result.passed:6 return code, attempt + 1, "passed"7 code = call(f"""8This implementation fails its tests.910CODE:11{code}1213TEST OUTPUT (verbatim — this is ground truth, not an opinion):14{result.output}1516Fix the code so these tests pass. Before writing the fix, state in17one sentence what the failing assertion proves about the current18behaviour. Do not change the function signature.""")19 return code, max_attempts, "still failing""This is ground truth, not an opinion" is not decoration. Without it, models argue with test output, explaining why the test is wrong, because disagreement is a plausible continuation of a criticism. Marking it authoritative shifts that continuation towards acceptance. And "state what the failing assertion proves before writing the fix" forces diagnosis before the first token of the patch.
Choosing the right check for the error
Different error types need different machinery, and using a model where code would do is the commonest waste here.
| Error type | Example | Best checker | Why not a model critique |
|---|---|---|---|
| Arithmetic | "1,240 × 0.15 = 201" | Evaluate the expression in code | The critique needs the same multiplication the answer got wrong |
| Constraint violation | Plan exceeds the stated budget | Assertion in code | Deterministic and free; a model call is strictly worse |
| Unsupported claim | Cites a policy clause that does not exist | Retrieval, then string match against the source | The model invented it once and will happily confirm it |
| Internal contradiction | Paragraph 2 says 30 days, paragraph 6 says 45 | Model, given a narrow "quote both statements" instruction | Needs language understanding — but only with a specific instruction |
| Scope drift | Answered a related question, not the one asked | Model, comparing answer against requirements one by one | Works, but only as an itemised checklist, not "did I answer the question?" |
| Stale assumption | Uses last year's tax band | Retrieval against a dated source | The model has no way to know what today's value is |
What this means when you build something
Treat every critique step as a component with two measurable properties, and refuse to ship one you have not measured.
Build a small set of cases where you know the right answer. Run the pipeline with the critic and without it. Record four numbers: correct answers damaged, wrong answers repaired, added cost, added latency. That gives d, f, and the price of both. Compute f/(f+d) and compare it to your accuracy. If accuracy is already above that ratio, the critic is a liability no matter how sensible its comments read.
Then apply the routing rule:
- An external verifier exists — tests, a schema, a database, a calculator, a retrievable source. Use it, feeding its raw output back verbatim and marked as authoritative. This is the only configuration with reliably large gains, and worth engineering a verifier where none exists.
- No verifier, but a decision with real consequences. Use debate: separate calls, swapped ordering, symmetric length caps, a judge that must quote its evidence. Expect the value to come less from the verdict than from the missing information it names.
- No verifier, routine output. Use decomposed checks with concrete witnesses required, or nothing. Never "review your answer and fix any mistakes" — that is the configuration measured above, and it cost the team two points of accuracy and 60% of their budget.
One last discipline: keep critique and revision in separate calls, and keep the critic's view of the original reasoning as narrow as the task allows. Every token of the original answer the critic sees pulls its output towards agreement. That is not a flaw you can prompt your way out of — it is what conditioning on context means — so control the context.