Course Content
Advanced Prompting and Reasoning
3 sections · 7 lessons
Evaluating Reasoning Quality
An analytics assistant was asked for the total discount on two orders: 15% off 240, and 20% off 180. It replied:
15% of 240 = 3220% of 180 = 40Total discount = 32 + 40 = 72The harness checked the final number against the expected 72 and marked it correct. It had a 100% pass rate on discount questions, and shipped.
Now check the middle two lines. 15% of 240 is 36, not 32. 20% of 180 is 36, not 40. Both steps are wrong — one low by 4, one high by 4 — and they cancel. The right answer came from two errors that summed to zero.
Nothing about this assistant is reliable. Next time the numbers are 300 and 150 the errors will not cancel, and nothing in the test suite would have warned you. The harness validated an outcome; what needed validating was a process.
That is the whole subject in one example. Checking the answer is necessary and nowhere near sufficient: a chain can give a right answer through broken steps, a wrong answer through sound steps and one bad input, or a right answer to a question nobody asked. Telling those apart is what reasoning evaluation is for.
Why fluency is not evidence
Human readers use fluency as a proxy for correctness, and in human-written text it is decent: producing confident, well-organised, specific prose about a subject you do not understand is hard for a person.
It is not hard for a language model. The objective rewards likely text, and coherence, confident register, specific-sounding detail and clean structure are properties of likely text — optimised directly. Truth is optimised only where true statements were common in the corpus: fine for well-covered facts, useless where you need discrimination most — rare facts, recent facts, novel combinations, edge cases, anything specific to your organisation.
Fluency and correctness are produced by different mechanisms and correlate only weakly at the margin. The reader's confidence-detector is calibrated on humans and gives no useful signal here.
A second-order version catches those who know the first: a chain of thought is also generated text. Writing steps out buys real computation, but does not make them a faithful trace of how the answer was produced — a model can reach an answer and then generate a plausible justification. So a chain of reasoning is evidence to be checked, not an explanation to be trusted.
This has been measured. In a 2025 Anthropic study, reasoning models (Claude 3.7 Sonnet and DeepSeek R1) were given hints that changed their answers, and their written reasoning mentioned the hint only 25% and 39% of the time on average. There is also a practical limit: current Claude models do not return their raw thinking at all, only a summary or an empty block. What you can evaluate is the working written into the answer, the tool calls and their results, and the answer itself.
Where this matters most
| Setting | What makes it dangerous | What outcome-only checking misses |
|---|---|---|
| Any output a human will act on without re-deriving it | The reasoning is the deliverable | Everything — there is no separate "answer" to check |
| Chains longer than about five steps | Per-step error compounds; see below | Which step failed, so nothing can be fixed |
| Rare or novel cases | Fluency stays constant while accuracy drops | The drop, because the output still reads well |
| Anything with regulatory or safety exposure | You must show why, not just what | The audit trail |
| Cases with no ground truth available | You cannot check the answer at all | Process is the only thing left to check |
That second row deserves numbers. If each step is independently correct with probability p, the whole chain is correct with probability pn:
| Per-step accuracy | 3 steps | 6 steps | 10 steps |
|---|---|---|---|
| 0.90 | 72.9% | 53.1% | 34.9% |
| 0.95 | 85.7% | 73.5% | 59.9% |
| 0.98 | 94.1% | 88.6% | 81.7% |
| 0.99 | 97.0% | 94.1% | 90.4% |
A model right 95% of the time per step is right 73.5% of the time on a six-step chain, because 0.956=0.735. End-to-end measurement shows a 73.5% system without telling you whether to attack the model, the prompt or one step. Step-level measurement does. (Errors are not perfectly independent — a misread question corrupts every step — but the direction holds: long chains are fragile, short ones are not.)
Six dimensions worth scoring separately
"Is this good reasoning?" is not answerable. Six narrower questions are, and separating them tells you what to fix.
| Dimension | The question it answers | Best checker |
|---|---|---|
| 1. Logical consistency | Does each conclusion follow from what precedes it? Does anything contradict anything else? | Model, with a "quote both statements" instruction |
| 2. Factual accuracy | Is every stated fact and every computed number correct? | Code for arithmetic; retrieval for facts |
| 3. Completeness | Was every part of the question addressed? Any material consideration omitted? | Itemised checklist from the question |
| 4. Evidentiary support | Is each claim traceable to the input, or merely asserted? | String match of claims against sources |
| 5. Freedom from fallacies | Does any inference rely on a known invalid move? | Model, against a named list of fallacies |
| 6. Clarity | Can a competent reader follow and check the argument? | Model or human |
The ordering is deliberate. Dimensions 1 and 2 ask whether the reasoning is sound, 3 and 4 whether it is grounded, 5 whether the inference moves are valid, 6 whether any of it is visible. Clarity is last because it best predicts human approval and worst predicts correctness. Weight it accordingly.
Fallacies that actually show up in generated reasoning
Not the classical list — these appear repeatedly in practice:
- Correlation stated as causation. "Users of feature X retain better, so X drives retention" — the alternative, that engaged users adopt more features, goes unnamed.
- Base-rate neglect. A 95%-accurate test for a condition affecting 1 in 1,000. Of 100,000 people: 100 have it and about 95 test positive; 99,900 do not and about 4,995 false-positive. A positive result means roughly 95/5,090=1.9% — not 95%. Generated reasoning reaches for the 95% almost every time.
- Authority as evidence. "Industry best practice is..." with no mechanism and no source. The commonest in business analysis.
- Single-cause explanation. A multi-causal outcome pinned on one factor, because one factor makes a cleaner story.
- Equivocation across steps. "Users" means registered accounts in step 2 and monthly actives in step 5, and the ratio silently changes meaning.
Three evaluation frameworks
Component-based: verify each step
Split the chain into atomic claims and check each independently. On the opening example:
| Step | Claim | Checker | Result |
|---|---|---|---|
| 1 | 15% of 240 = 32 | 0.15 * 240 | FAIL — 36 |
| 2 | 20% of 180 = 40 | 0.20 * 180 | FAIL — 36 |
| 3 | 32 + 40 = 72 | 32 + 40 | PASS — internally valid |
| 4 | Total discount = 72 | Compare to correct 36 + 36 | PASS on value, on wrong inputs |
Step-level accuracy is 50%. Outcome accuracy is 100%. The gap between those numbers is the entire diagnostic.
Each step must be checked in isolation, without the surrounding narrative. In context, step 1 sits in a fluent calculation and reads fine. Alone, as "is 15% of 240 equal to 32?", it is trivially wrong. Context supplies plausibility that suppresses scrutiny, in the checker as much as the reader.
1import re23ARITH = re.compile(r"([\d.]+)\s*%\s*of\s*([\d,]+)\s*=\s*([\d,.]+)")45def check_percentages(chain):6 failures = []7 for pct, base, claimed in ARITH.findall(chain):8 actual = float(pct) / 100 * float(base.replace(",", ""))9 stated = float(claimed.replace(",", ""))10 if abs(actual - stated) > 0.01:11 failures.append({"claim": f"{pct}% of {base} = {claimed}",12 "actual": round(actual, 2)})13 return failuresTwelve deterministic lines catch a class of error no model critique reliably catches, because the critique needs the same multiplication the answer got wrong. Check in code whatever is checkable in code: free, exact, repeatable.
Rubric-based: score anchored dimensions and combine
Weight the six dimensions by how much each matters for the task. For an analytical report:
| Dimension | Weight | Score | Contribution |
|---|---|---|---|
| Factual accuracy | 0.30 | 6 | 1.80 |
| Logical consistency | 0.25 | 9 | 2.25 |
| Completeness | 0.15 | 7 | 1.05 |
| Evidentiary support | 0.15 | 3 | 0.45 |
| Freedom from fallacies | 0.10 | 8 | 0.80 |
| Clarity | 0.05 | 9 | 0.45 |
| Weighted total | 1.00 | — | 6.80 |
6.80 out of 10 reads as "reasonable, needs work". It is not. Evidentiary support scored 3: the claims are not traceable to sources, and a report whose claims cannot be traced does not need polish — it cannot be used. A 9 for clarity bought off a 3 for grounding.
So a weighted mean always needs gates:
1WEIGHTS = {"factual": 0.30, "logic": 0.25, "complete": 0.15,2 "support": 0.15, "fallacies": 0.10, "clarity": 0.05}3GATES = {"factual": 5, "logic": 5, "support": 4} # dimensions that cannot be traded away45def rubric_score(scores):6 total = sum(WEIGHTS[k] * scores[k] for k in WEIGHTS)7 breached = [k for k, floor in GATES.items() if scores[k] < floor]8 if breached:9 return {"score": min(total, 5.0), "verdict": "fail",10 "reason": f"gate breached on: {', '.join(breached)}", "raw": round(total, 2)}11 return {"score": round(total, 2), "verdict": "pass" if total >= 7 else "review"}With support = 3 below its gate of 4, this returns a capped 5.0 and a fail, keeping the raw 6.80 for diagnosis. The gate encodes what a mean cannot: some failures disqualify rather than cost.
The scale must be anchored or the scores are decorative: an unanchored 1–10 collapses to 7 and 8 for anything coherent, making the weighted total a constant. Define each band:
FACTUAL ACCURACY 9-10 Every checkable fact verified correct. No unverifiable specifics asserted. 7-8 All material facts correct. Minor imprecision that does not affect the conclusion. 5-6 One material fact wrong or unverifiable, conclusion survives. 3-4 A fact the conclusion depends on is wrong. 0-2 Multiple fabricated specifics (invented figures, sources, quotations).Branch validation: check what was not chosen
When an answer is a recommendation, the reasoning includes rejections — options considered and dropped. They are invisible in the output, and where the failure hides. Three questions:
- Were the real alternatives on the table? A recommendation that weighed two options when four existed has a completeness failure no amount of internal consistency fixes.
- Was each rejection justified against a stated criterion? "Option B is not a good fit" is a rejection with no content.
- Would the recommendation survive a different weighting? If ranking A above B depends entirely on weighting cost over speed, that dependency is the finding, and should be stated.
Prompting an evaluation
The self-evaluation prompt, and its ceiling
Asking a model to evaluate its own output has a hard limit: the critique comes from the same weights, so it cannot add a fact the answer lacked, and it is conditioned on the answer in context, which pulls towards agreement. Worth doing only when narrow and witness-bearing:
BADEvaluate the quality of your reasoning above.GOODCheck your answer against each item. Answer PASS or FAIL for each.A FAIL requires a quoted line and a concrete demonstration; if youcannot produce one, answer PASS.1. ARITHMETIC. Recompute every calculation independently, showing the working. Do not copy your earlier figures.2. SOURCING. List every factual claim you made. For each, quote the text in the provided documents that supports it, or mark it UNSUPPORTED.3. COVERAGE. Restate the question as a numbered list of sub-questions. For each, quote the sentence of your answer that addresses it.4. LOAD-BEARING ASSUMPTIONS. Name any assumption which, if false, would change your conclusion. State whether you made it explicit.5. CONTRADICTIONS. Quote any two statements of yours that cannot both be true.Every clause counters a mechanism. "Recompute independently, do not copy" stops the check being a re-read of the same tokens. "Quote the supporting text or mark UNSUPPORTED" makes fabrication detectable: an unquotable claim must be labelled. "Restate as numbered sub-questions" turns a vague completeness judgement into a matching exercise. And "if you cannot produce a demonstration, answer PASS" licenses PASS — without it, "identify the problems" manufactures them.
Adversarial evaluation
Rather than ask whether the answer is right, tell the evaluator it is wrong and make it find where. "Find the flaw" makes a flaw the likely continuation — the bias you want against uncritical agreement.
This analysis contains at least one significant error. Your job isto find it. Work in this order:1. Identify the single claim on which the conclusion most depends.2. Construct the strongest concrete case in which that claim is false. Use specific numbers or a specific scenario.3. State what the conclusion becomes if it is false.4. Only if you genuinely cannot construct such a case, say: "No load-bearing claim could be falsified" and explain why.You may not object to style, tone, hedging or completeness.Substantive errors only.Two guards keep this honest. A concrete falsifying case makes the objection constructible, not merely assertable. Forbidding style objections stops the evaluator discharging its instruction with a cost-free complaint about tone, which it will do otherwise.
Independent peer review
The strongest configuration, because it removes self-conditioning: a fresh call seeing the question and the answer but not the reasoning trace, ideally solving the problem itself before comparing. Two independent derivations agreeing is evidence; one approving of itself is not.
Judge biases you must design around
| Bias | Effect | Mitigation |
|---|---|---|
| Length | Longer answers score higher; detail reads as rigour | Cap length symmetrically; ask "is any of this padding?" |
| Position | In a pairwise comparison, one slot is systematically favoured | Run both orderings; a flipped verdict means "tie", not "close" |
| Self-preference | Text closer to the judge's own distribution scores higher | Unavoidable with one model; treat scores as evidence, not findings |
| Sycophancy to the framing | "Confirm this is correct" gets confirmation; "find the error" gets errors | Neutral instruction plus a required witness per finding |
| Score compression | Unanchored scales collapse to 7–8 | Anchored bands, or pairwise instead of absolute scores |
| Confidence | Assertive phrasing reads as competence | Score claims against sources, not against tone |
Models are far more reliable at pairwise comparison than at absolute scoring. Ranking two concrete items needs only the ordering to be right; scoring one item needs a calibrated internal scale, and there is no such scale.
Evaluate the evaluator
Your judge is a model, with an accuracy you do not know until you measure it. Build 50 carefully human-labelled items — known-good, known-broken-process-right-answer, known-subtly-wrong — and measure judge agreement against them. Then measure two humans against each other on the same set.
If human-human agreement is 0.85 and judge-human agreement is 0.68, your judge is a noisy instrument: useful for ranking a hundred outputs, useless for one borderline case. Report ranges and route disagreements to a person. If the rates are close, lean on the judge. Either way you know which — and almost nobody checks.
Domain checklists
Generic dimensions catch generic failures. Each domain has characteristic ones.
| Domain | Check |
|---|---|
| Clinical | Were dangerous alternatives ruled out explicitly, not just the likely one? Are dose, route and frequency present and consistent? Are contraindications and interactions addressed? Is uncertainty stated and escalation named? Is the evidence population the one being reasoned about? |
| Financial | Are figures nominal or real, and said to be? Is the period explicit and consistent across every figure? Are the discount rate and its justification given? Is the base of every percentage identified? Do balance-sheet identities balance? Are one-off items separated from recurring? |
| Legal | Is the jurisdiction named? Is the cited authority real, current and on point? Is a statute quoted or paraphrased — and if paraphrased, faithfully? Are contrary authorities acknowledged? Is the analysis applied to these facts, not the general rule? |
| Scientific | Is the sample size and its source given? Significance and effect size both reported? Is the control condition described? Is the causal claim supported by a design that can support causality? Is the confidence interval given, not just the point estimate? Are competing explanations named? |
Red flags, with the question that exposes each
| Pattern | What it looks like | Exposing question |
|---|---|---|
| Oversimplification | "Churn rose because the price increased." | "Name three other factors that could produce this, and what evidence would separate them." |
| Unfounded confidence | "This will increase conversion by roughly 15%." | "What is that 15% derived from? If nothing, restate it as a range with assumptions." |
| Missing context | "Revenue grew 40% this quarter." | "Compared to what baseline, over what period, and what was the prior-year figure?" |
| Logical jump | "Best practice is to migrate, so we should migrate." | "State the mechanism by which migrating produces the outcome we want, here." |
| Silent unit or definition shift | "Users" means accounts in one step, actives in another | "Define every quantity once, and confirm each use matches that definition." |
| Precision beyond the input | "Payback in 7.3 months" from two estimated inputs | "How many significant figures do the inputs justify?" |
Two worked evaluations
A business recommendation
"We should raise the price of the Pro plan from 29 to 39.Competitors charge 45 on average, we are clearly underpriced, anda 34% price increase will grow revenue by 34%. Churn risk is lowbecause our users love the product."| Dimension | Score | Finding |
|---|---|---|
| Factual accuracy | 4 | 29 → 39 is a 34.5% increase, which is right. "Competitors charge 45 on average" is unverified and names no competitors. |
| Logical consistency | 2 | The core inference is invalid. Revenue = price × quantity; a 34% price rise grows revenue by 34% only if quantity is unchanged, which is the question. Break-even is a volume loss of 1−1/1.345=25.7%; anything worse and revenue falls. |
| Completeness | 3 | No mention of existing versus new customers, grandfathering, feature parity with the 45 comparison, or elasticity evidence. |
| Evidentiary support | 2 | "Users love the product" as churn evidence — satisfaction and price sensitivity are different quantities. No retention or elasticity data. |
| Fallacies | 3 | Authority framing ("competitors charge more" as a reason), single-cause reasoning, and an unstated ceteris paribus assumption doing the work. |
| Clarity | 8 | Perfectly clear. This is the problem — it is clear, brief and wrong. |
Weighted total: 0.30(4)+0.25(2)+0.15(3)+0.15(2)+0.10(3)+0.05(8)=1.20+0.50+0.45+0.30+0.30+0.40=3.15. Gates breached on logic and support: fail. The most useful output is the break-even figure, 25.7% volume loss, which reframes the question from "are we underpriced?" to "will we lose more than a quarter of Pro subscribers?"
A scientific claim
"A study found that people who take vitamin D have 20% fewerrespiratory infections. Therefore vitamin D supplementationprevents respiratory infections."Six failures, each independent:
- Design unstated. "People who take" describes an observational cohort, which cannot support "prevents". Only a randomised trial licenses that verb.
- Confounding unaddressed. Supplement takers differ systematically — income, healthcare access, outdoor time, baseline health. Any could produce the association.
- "20% fewer" is unanchored. Relative or absolute? A drop from 5% to 4% is a 20% relative reduction and a 1-point absolute one, implying very different numbers needed to treat.
- No sample size, no interval. A 20% reduction with a 95% interval of 2% to 35% is a different finding from one spanning −5% to 40%.
- No dose, form, duration or baseline status. The effect in deficient populations differs from replete ones; conflating them is the standard error here.
- Single study treated as settled. No mention of replication or contrary results.
A defensible restatement: "One observational study reported a 20% lower rate of respiratory infection among supplement users. Because participants were not randomised, this cannot distinguish an effect of vitamin D from differences between people who do and do not supplement. Absolute reduction, confidence interval, dose and baseline status would all be needed before concluding anything."
What this means when you build something
Evaluation is not a review at the end but a layered filter, and the layers differ enormously in cost per error caught. Order them cheapest-first and most errors never reach an expensive one.
| Layer | Catches | Cost | Run on |
|---|---|---|---|
| Deterministic checks in code | Arithmetic, format, constraint violations, schema failures, unit mismatches | Effectively zero | Every output |
| Retrieval-backed claim checking | Fabricated facts, unsupported assertions | Low | Every output with factual claims |
| Rubric scoring with gates | Logic gaps, incompleteness, fallacies | Moderate | Samples, plus anything a cheap layer flagged |
| Adversarial or peer review | Subtle unsupported inferences | High | High-stakes outputs only |
| Human review | Everything else, including whether the question was right | Highest | Gate failures, low-confidence cases, a fixed audit sample |
Three disciplines separate a harness that improves a system from one that produces reassuring numbers.
Build the failure set before the success set. Test suites fill with cases the system already handles, because those are the ones you think of. Collect the other kind: right answer with broken working (the opening example), wrong answer with sound working, right answer to the wrong question, confident answer where "insufficient information" was correct. Without those, a suite measures nothing you cannot see already.
Score dimensions separately; never publish only the mean. One number cannot distinguish "clear but ungrounded" from "sound but unreadable", and those need opposite fixes. Report the vector, gate what cannot be traded away, and keep the raw total beside the gated verdict.
Treat every score as an instrument reading with a known error bar. At 0.68 judge-human agreement, the gap between 7.1 and 7.4 is noise, and acting on it is superstition. Scores from a model evaluator are evidence about quality, not quality. The moment a team optimises the number rather than the thing, the evaluator has become the problem it was built to solve.