Advanced Prompting and Reasoning

Evaluating Reasoning Quality


An analytics assistant was asked for the total discount on two orders: 15% off 240, and 20% off 180. It replied:

Text
15% of 240 = 3220% of 180 = 40Total discount = 32 + 40 = 72

The harness checked the final number against the expected 72 and marked it correct. It had a 100% pass rate on discount questions, and shipped.

Now check the middle two lines. 15% of 240 is 36, not 32. 20% of 180 is 36, not 40. Both steps are wrong — one low by 4, one high by 4 — and they cancel. The right answer came from two errors that summed to zero.

Nothing about this assistant is reliable. Next time the numbers are 300 and 150 the errors will not cancel, and nothing in the test suite would have warned you. The harness validated an outcome; what needed validating was a process.

That is the whole subject in one example. Checking the answer is necessary and nowhere near sufficient: a chain can give a right answer through broken steps, a wrong answer through sound steps and one bad input, or a right answer to a question nobody asked. Telling those apart is what reasoning evaluation is for.

Six dimensions, scored separatelyWas thereasoning sound?Faithful to the stated stepsEach step valid on its ownArithmetic checked by a toolNothing relevant left outConclusion follows the stepsConfidence matches evidence
A fluent chain with a wrong multiplication scores well on five of these and fails the one that decides the answer.

Why fluency is not evidence

Human readers use fluency as a proxy for correctness, and in human-written text it is decent: producing confident, well-organised, specific prose about a subject you do not understand is hard for a person.

It is not hard for a language model. The objective rewards likely text, and coherence, confident register, specific-sounding detail and clean structure are properties of likely text — optimised directly. Truth is optimised only where true statements were common in the corpus: fine for well-covered facts, useless where you need discrimination most — rare facts, recent facts, novel combinations, edge cases, anything specific to your organisation.

Fluency and correctness are produced by different mechanisms and correlate only weakly at the margin. The reader's confidence-detector is calibrated on humans and gives no useful signal here.

A second-order version catches those who know the first: a chain of thought is also generated text. Writing steps out buys real computation, but does not make them a faithful trace of how the answer was produced — a model can reach an answer and then generate a plausible justification. So a chain of reasoning is evidence to be checked, not an explanation to be trusted.

This has been measured. In a 2025 Anthropic study, reasoning models (Claude 3.7 Sonnet and DeepSeek R1) were given hints that changed their answers, and their written reasoning mentioned the hint only 25% and 39% of the time on average. There is also a practical limit: current Claude models do not return their raw thinking at all, only a summary or an empty block. What you can evaluate is the working written into the answer, the tool calls and their results, and the answer itself.

Where this matters most

SettingWhat makes it dangerousWhat outcome-only checking misses
Any output a human will act on without re-deriving itThe reasoning is the deliverableEverything — there is no separate "answer" to check
Chains longer than about five stepsPer-step error compounds; see belowWhich step failed, so nothing can be fixed
Rare or novel casesFluency stays constant while accuracy dropsThe drop, because the output still reads well
Anything with regulatory or safety exposureYou must show why, not just whatThe audit trail
Cases with no ground truth availableYou cannot check the answer at allProcess is the only thing left to check

That second row deserves numbers. If each step is independently correct with probability pp, the whole chain is correct with probability pnp^n:

Per-step accuracy3 steps6 steps10 steps
0.9072.9%53.1%34.9%
0.9585.7%73.5%59.9%
0.9894.1%88.6%81.7%
0.9997.0%94.1%90.4%

A model right 95% of the time per step is right 73.5% of the time on a six-step chain, because 0.956=0.7350.95^6 = 0.735. End-to-end measurement shows a 73.5% system without telling you whether to attack the model, the prompt or one step. Step-level measurement does. (Errors are not perfectly independent — a misread question corrupts every step — but the direction holds: long chains are fragile, short ones are not.)

Six dimensions worth scoring separately

"Is this good reasoning?" is not answerable. Six narrower questions are, and separating them tells you what to fix.

DimensionThe question it answersBest checker
1. Logical consistencyDoes each conclusion follow from what precedes it? Does anything contradict anything else?Model, with a "quote both statements" instruction
2. Factual accuracyIs every stated fact and every computed number correct?Code for arithmetic; retrieval for facts
3. CompletenessWas every part of the question addressed? Any material consideration omitted?Itemised checklist from the question
4. Evidentiary supportIs each claim traceable to the input, or merely asserted?String match of claims against sources
5. Freedom from fallaciesDoes any inference rely on a known invalid move?Model, against a named list of fallacies
6. ClarityCan a competent reader follow and check the argument?Model or human

The ordering is deliberate. Dimensions 1 and 2 ask whether the reasoning is sound, 3 and 4 whether it is grounded, 5 whether the inference moves are valid, 6 whether any of it is visible. Clarity is last because it best predicts human approval and worst predicts correctness. Weight it accordingly.

Fallacies that actually show up in generated reasoning

Not the classical list — these appear repeatedly in practice:

  • Correlation stated as causation. "Users of feature X retain better, so X drives retention" — the alternative, that engaged users adopt more features, goes unnamed.
  • Base-rate neglect. A 95%-accurate test for a condition affecting 1 in 1,000. Of 100,000 people: 100 have it and about 95 test positive; 99,900 do not and about 4,995 false-positive. A positive result means roughly 95/5,090=1.9%95 / 5{,}090 = 1.9\% — not 95%. Generated reasoning reaches for the 95% almost every time.
  • Authority as evidence. "Industry best practice is..." with no mechanism and no source. The commonest in business analysis.
  • Single-cause explanation. A multi-causal outcome pinned on one factor, because one factor makes a cleaner story.
  • Equivocation across steps. "Users" means registered accounts in step 2 and monthly actives in step 5, and the ratio silently changes meaning.

Three evaluation frameworks

Component-based: verify each step

Split the chain into atomic claims and check each independently. On the opening example:

StepClaimCheckerResult
115% of 240 = 320.15 * 240FAIL — 36
220% of 180 = 400.20 * 180FAIL — 36
332 + 40 = 7232 + 40PASS — internally valid
4Total discount = 72Compare to correct 36 + 36PASS on value, on wrong inputs

Step-level accuracy is 50%. Outcome accuracy is 100%. The gap between those numbers is the entire diagnostic.

Each step must be checked in isolation, without the surrounding narrative. In context, step 1 sits in a fluent calculation and reads fine. Alone, as "is 15% of 240 equal to 32?", it is trivially wrong. Context supplies plausibility that suppresses scrutiny, in the checker as much as the reader.

Python
import reARITH = re.compile(r"([\d.]+)\s*%\s*of\s*([\d,]+)\s*=\s*([\d,.]+)")def check_percentages(chain):    failures = []    for pct, base, claimed in ARITH.findall(chain):        actual = float(pct) / 100 * float(base.replace(",", ""))        stated = float(claimed.replace(",", ""))        if abs(actual - stated) > 0.01:            failures.append({"claim": f"{pct}% of {base} = {claimed}",                             "actual": round(actual, 2)})    return failures

Twelve deterministic lines catch a class of error no model critique reliably catches, because the critique needs the same multiplication the answer got wrong. Check in code whatever is checkable in code: free, exact, repeatable.

Rubric-based: score anchored dimensions and combine

Weight the six dimensions by how much each matters for the task. For an analytical report:

DimensionWeightScoreContribution
Factual accuracy0.3061.80
Logical consistency0.2592.25
Completeness0.1571.05
Evidentiary support0.1530.45
Freedom from fallacies0.1080.80
Clarity0.0590.45
Weighted total1.00—6.80

6.80 out of 10 reads as "reasonable, needs work". It is not. Evidentiary support scored 3: the claims are not traceable to sources, and a report whose claims cannot be traced does not need polish — it cannot be used. A 9 for clarity bought off a 3 for grounding.

So a weighted mean always needs gates:

Python
WEIGHTS = {"factual": 0.30, "logic": 0.25, "complete": 0.15,           "support": 0.15, "fallacies": 0.10, "clarity": 0.05}GATES   = {"factual": 5, "logic": 5, "support": 4}   # dimensions that cannot be traded awaydef rubric_score(scores):    total = sum(WEIGHTS[k] * scores[k] for k in WEIGHTS)    breached = [k for k, floor in GATES.items() if scores[k] < floor]    if breached:        return {"score": min(total, 5.0), "verdict": "fail",                "reason": f"gate breached on: {', '.join(breached)}", "raw": round(total, 2)}    return {"score": round(total, 2), "verdict": "pass" if total >= 7 else "review"}

With support = 3 below its gate of 4, this returns a capped 5.0 and a fail, keeping the raw 6.80 for diagnosis. The gate encodes what a mean cannot: some failures disqualify rather than cost.

The scale must be anchored or the scores are decorative: an unanchored 1–10 collapses to 7 and 8 for anything coherent, making the weighted total a constant. Define each band:

Text
FACTUAL ACCURACY 9-10  Every checkable fact verified correct. No unverifiable       specifics asserted. 7-8   All material facts correct. Minor imprecision that does not       affect the conclusion. 5-6   One material fact wrong or unverifiable, conclusion survives. 3-4   A fact the conclusion depends on is wrong. 0-2   Multiple fabricated specifics (invented figures, sources,       quotations).

Branch validation: check what was not chosen

When an answer is a recommendation, the reasoning includes rejections — options considered and dropped. They are invisible in the output, and where the failure hides. Three questions:

  • Were the real alternatives on the table? A recommendation that weighed two options when four existed has a completeness failure no amount of internal consistency fixes.
  • Was each rejection justified against a stated criterion? "Option B is not a good fit" is a rejection with no content.
  • Would the recommendation survive a different weighting? If ranking A above B depends entirely on weighting cost over speed, that dependency is the finding, and should be stated.

Prompting an evaluation

The self-evaluation prompt, and its ceiling

Asking a model to evaluate its own output has a hard limit: the critique comes from the same weights, so it cannot add a fact the answer lacked, and it is conditioned on the answer in context, which pulls towards agreement. Worth doing only when narrow and witness-bearing:

Text
BADEvaluate the quality of your reasoning above.
Text
GOODCheck your answer against each item. Answer PASS or FAIL for each.A FAIL requires a quoted line and a concrete demonstration; if youcannot produce one, answer PASS.1. ARITHMETIC. Recompute every calculation independently, showing   the working. Do not copy your earlier figures.2. SOURCING. List every factual claim you made. For each, quote the   text in the provided documents that supports it, or mark it   UNSUPPORTED.3. COVERAGE. Restate the question as a numbered list of sub-questions.   For each, quote the sentence of your answer that addresses it.4. LOAD-BEARING ASSUMPTIONS. Name any assumption which, if false,   would change your conclusion. State whether you made it explicit.5. CONTRADICTIONS. Quote any two statements of yours that cannot   both be true.

Every clause counters a mechanism. "Recompute independently, do not copy" stops the check being a re-read of the same tokens. "Quote the supporting text or mark UNSUPPORTED" makes fabrication detectable: an unquotable claim must be labelled. "Restate as numbered sub-questions" turns a vague completeness judgement into a matching exercise. And "if you cannot produce a demonstration, answer PASS" licenses PASS — without it, "identify the problems" manufactures them.

Adversarial evaluation

Rather than ask whether the answer is right, tell the evaluator it is wrong and make it find where. "Find the flaw" makes a flaw the likely continuation — the bias you want against uncritical agreement.

Text
This analysis contains at least one significant error. Your job isto find it. Work in this order:1. Identify the single claim on which the conclusion most depends.2. Construct the strongest concrete case in which that claim is   false. Use specific numbers or a specific scenario.3. State what the conclusion becomes if it is false.4. Only if you genuinely cannot construct such a case, say:   "No load-bearing claim could be falsified" and explain why.You may not object to style, tone, hedging or completeness.Substantive errors only.

Two guards keep this honest. A concrete falsifying case makes the objection constructible, not merely assertable. Forbidding style objections stops the evaluator discharging its instruction with a cost-free complaint about tone, which it will do otherwise.

Independent peer review

The strongest configuration, because it removes self-conditioning: a fresh call seeing the question and the answer but not the reasoning trace, ideally solving the problem itself before comparing. Two independent derivations agreeing is evidence; one approving of itself is not.

Judge biases you must design around

BiasEffectMitigation
LengthLonger answers score higher; detail reads as rigourCap length symmetrically; ask "is any of this padding?"
PositionIn a pairwise comparison, one slot is systematically favouredRun both orderings; a flipped verdict means "tie", not "close"
Self-preferenceText closer to the judge's own distribution scores higherUnavoidable with one model; treat scores as evidence, not findings
Sycophancy to the framing"Confirm this is correct" gets confirmation; "find the error" gets errorsNeutral instruction plus a required witness per finding
Score compressionUnanchored scales collapse to 7–8Anchored bands, or pairwise instead of absolute scores
ConfidenceAssertive phrasing reads as competenceScore claims against sources, not against tone

Models are far more reliable at pairwise comparison than at absolute scoring. Ranking two concrete items needs only the ordering to be right; scoring one item needs a calibrated internal scale, and there is no such scale.

Evaluate the evaluator

Your judge is a model, with an accuracy you do not know until you measure it. Build 50 carefully human-labelled items — known-good, known-broken-process-right-answer, known-subtly-wrong — and measure judge agreement against them. Then measure two humans against each other on the same set.

If human-human agreement is 0.85 and judge-human agreement is 0.68, your judge is a noisy instrument: useful for ranking a hundred outputs, useless for one borderline case. Report ranges and route disagreements to a person. If the rates are close, lean on the judge. Either way you know which — and almost nobody checks.

Domain checklists

Generic dimensions catch generic failures. Each domain has characteristic ones.

DomainCheck
ClinicalWere dangerous alternatives ruled out explicitly, not just the likely one? Are dose, route and frequency present and consistent? Are contraindications and interactions addressed? Is uncertainty stated and escalation named? Is the evidence population the one being reasoned about?
FinancialAre figures nominal or real, and said to be? Is the period explicit and consistent across every figure? Are the discount rate and its justification given? Is the base of every percentage identified? Do balance-sheet identities balance? Are one-off items separated from recurring?
LegalIs the jurisdiction named? Is the cited authority real, current and on point? Is a statute quoted or paraphrased — and if paraphrased, faithfully? Are contrary authorities acknowledged? Is the analysis applied to these facts, not the general rule?
ScientificIs the sample size and its source given? Significance and effect size both reported? Is the control condition described? Is the causal claim supported by a design that can support causality? Is the confidence interval given, not just the point estimate? Are competing explanations named?

Red flags, with the question that exposes each

PatternWhat it looks likeExposing question
Oversimplification"Churn rose because the price increased.""Name three other factors that could produce this, and what evidence would separate them."
Unfounded confidence"This will increase conversion by roughly 15%.""What is that 15% derived from? If nothing, restate it as a range with assumptions."
Missing context"Revenue grew 40% this quarter.""Compared to what baseline, over what period, and what was the prior-year figure?"
Logical jump"Best practice is to migrate, so we should migrate.""State the mechanism by which migrating produces the outcome we want, here."
Silent unit or definition shift"Users" means accounts in one step, actives in another"Define every quantity once, and confirm each use matches that definition."
Precision beyond the input"Payback in 7.3 months" from two estimated inputs"How many significant figures do the inputs justify?"

Two worked evaluations

A business recommendation

Text
"We should raise the price of the Pro plan from 29 to 39.Competitors charge 45 on average, we are clearly underpriced, anda 34% price increase will grow revenue by 34%. Churn risk is lowbecause our users love the product."
DimensionScoreFinding
Factual accuracy429 → 39 is a 34.5% increase, which is right. "Competitors charge 45 on average" is unverified and names no competitors.
Logical consistency2The core inference is invalid. Revenue = price × quantity; a 34% price rise grows revenue by 34% only if quantity is unchanged, which is the question. Break-even is a volume loss of 1−1/1.345=25.7%1 - 1/1.345 = 25.7\%; anything worse and revenue falls.
Completeness3No mention of existing versus new customers, grandfathering, feature parity with the 45 comparison, or elasticity evidence.
Evidentiary support2"Users love the product" as churn evidence — satisfaction and price sensitivity are different quantities. No retention or elasticity data.
Fallacies3Authority framing ("competitors charge more" as a reason), single-cause reasoning, and an unstated ceteris paribus assumption doing the work.
Clarity8Perfectly clear. This is the problem — it is clear, brief and wrong.

Weighted total: 0.30(4)+0.25(2)+0.15(3)+0.15(2)+0.10(3)+0.05(8)=1.20+0.50+0.45+0.30+0.30+0.40=3.150.30(4) + 0.25(2) + 0.15(3) + 0.15(2) + 0.10(3) + 0.05(8) = 1.20 + 0.50 + 0.45 + 0.30 + 0.30 + 0.40 = 3.15. Gates breached on logic and support: fail. The most useful output is the break-even figure, 25.7% volume loss, which reframes the question from "are we underpriced?" to "will we lose more than a quarter of Pro subscribers?"

A scientific claim

Text
"A study found that people who take vitamin D have 20% fewerrespiratory infections. Therefore vitamin D supplementationprevents respiratory infections."

Six failures, each independent:

  1. Design unstated. "People who take" describes an observational cohort, which cannot support "prevents". Only a randomised trial licenses that verb.
  2. Confounding unaddressed. Supplement takers differ systematically — income, healthcare access, outdoor time, baseline health. Any could produce the association.
  3. "20% fewer" is unanchored. Relative or absolute? A drop from 5% to 4% is a 20% relative reduction and a 1-point absolute one, implying very different numbers needed to treat.
  4. No sample size, no interval. A 20% reduction with a 95% interval of 2% to 35% is a different finding from one spanning −5% to 40%.
  5. No dose, form, duration or baseline status. The effect in deficient populations differs from replete ones; conflating them is the standard error here.
  6. Single study treated as settled. No mention of replication or contrary results.

A defensible restatement: "One observational study reported a 20% lower rate of respiratory infection among supplement users. Because participants were not randomised, this cannot distinguish an effect of vitamin D from differences between people who do and do not supplement. Absolute reduction, confidence interval, dose and baseline status would all be needed before concluding anything."

What this means when you build something

Evaluation is not a review at the end but a layered filter, and the layers differ enormously in cost per error caught. Order them cheapest-first and most errors never reach an expensive one.

LayerCatchesCostRun on
Deterministic checks in codeArithmetic, format, constraint violations, schema failures, unit mismatchesEffectively zeroEvery output
Retrieval-backed claim checkingFabricated facts, unsupported assertionsLowEvery output with factual claims
Rubric scoring with gatesLogic gaps, incompleteness, fallaciesModerateSamples, plus anything a cheap layer flagged
Adversarial or peer reviewSubtle unsupported inferencesHighHigh-stakes outputs only
Human reviewEverything else, including whether the question was rightHighestGate failures, low-confidence cases, a fixed audit sample

Three disciplines separate a harness that improves a system from one that produces reassuring numbers.

Build the failure set before the success set. Test suites fill with cases the system already handles, because those are the ones you think of. Collect the other kind: right answer with broken working (the opening example), wrong answer with sound working, right answer to the wrong question, confident answer where "insufficient information" was correct. Without those, a suite measures nothing you cannot see already.

Score dimensions separately; never publish only the mean. One number cannot distinguish "clear but ungrounded" from "sound but unreadable", and those need opposite fixes. Report the vector, gate what cannot be traded away, and keep the raw total beside the gated verdict.

Treat every score as an instrument reading with a known error bar. At 0.68 judge-human agreement, the gap between 7.1 and 7.4 is noise, and acting on it is superstition. Scores from a model evaluator are evidence about quality, not quality. The moment a team optimises the number rather than the thing, the evaluator has become the problem it was built to solve.