Course Content
Advanced Prompting and Reasoning
3 sections · 7 lessons
Self-Consistency & Ensemble Reasoning
Here is a problem an analyst gave to a model, and the model got it right. Then they ran it again, and it got it wrong. Same prompt, same model, same day.
A shop buys 240 units at 7.50 each. It sells 60% of them at a 40%markup on cost, and the remaining units at cost plus 10%.What is the total profit?The correct answer is 504. Total cost is 240 × 7.50 = 1,800. The 144 units sold at a 40% markup go out at 7.50 × 1.40 = 10.50, earning 144 × 3.00 = 432 of profit. The other 96 units go at 8.25, earning 96 × 0.75 = 72. Profit is 432 + 72 = 504.
Run the prompt ten times and you will see something like: 504, 792, 504, 504, 468, 504, 792, 504, 504, 504. Seven correct, two that confused a 40% markup on cost with a 40% margin on price (dividing by 0.6 instead of multiplying by 1.4), one arithmetic slip. Any single run is a coin weighted at about 0.7. Ship that as a single call and three customers in ten get a wrong number, with no signal that anything went wrong.
Now notice what the distribution of ten runs tells you that any one run cannot: 504 appeared seven times and no other answer appeared more than twice. That structure is free information you were throwing away. Self-consistency is the technique that picks it up.
Why the same prompt gives different answers
This is not randomness bolted on for variety. It falls out of how a token is chosen.
At each step the model produces a logit — an unnormalised score — for every token in its vocabulary. Those scores become a probability distribution through the softmax function, with a temperature parameter T that controls how sharp it is:
Suppose that at the moment the model is about to emit the profit figure, three candidate numbers carry the logits below. Here is what temperature actually does to them:
| Candidate | Logit z | P at T=0.5 | P at T=1.0 | P at T=1.5 |
|---|---|---|---|---|
| 504 (correct) | 3.2 | 0.727 | 0.549 | 0.478 |
| 792 (margin/markup slip) | 2.6 | 0.219 | 0.301 | 0.321 |
| 468 (arithmetic slip) | 1.9 | 0.054 | 0.150 | 0.201 |
Work one cell to see it. At T=1: e3.2=24.53, e2.6=13.46, e1.9=6.69; they sum to 44.68, so 504 gets 24.53/44.68=0.549. Halve the temperature and the exponents double to 6.4, 5.2 and 3.8, giving 601.8, 181.3 and 44.7 — a sum of 827.8 and a probability of 0.727 for 504. The gaps between logits are unchanged; only their contrast has been stretched.
Two consequences follow, and they are the reason this whole technique exists:
- Lower temperature does not add knowledge. It sharpens whatever ranking the model already had. If the wrong answer has the highest logit, T→0 makes you wrong every time instead of most of the time.
- Divergence compounds. One different token early — "margin" instead of "markup" — changes the context for every token after it. Two runs that differ at step 12 are on different trajectories by step 40. This is why you get genuinely distinct reasoning paths rather than paraphrases.
Sampling variance is not noise contaminating a true answer. It is the model's uncertainty, made visible. Self-consistency reads that uncertainty instead of ignoring it.
A note on the temperature knob itself
On current frontier Claude models — Opus 5, Sonnet 5, the 4.7 and 4.8 family — temperature, top_p and top_k have been removed from the Messages API and passing them returns a 400 error. Sampling still happens; the server manages it. That changes the practical shape of self-consistency in two ways. You get path diversity by default rather than by dialling temperature up, and where you want more diversity than the default sampler gives, you have to create it in the prompt rather than in the sampler. On older models that still accept the parameter, roughly 0.7–1.0 is the usual range for this technique.
The method
Three steps, and the third is where the design decisions live.
- Sample k independent reasoning paths for the same question. Independent means separate API calls, not one call asked to produce five answers — more on why that distinction is critical below.
- Extract a comparable answer from each path. Usually a number, a label, or a short canonical string.
- Aggregate: take the most frequent extracted answer, and use the vote distribution as a confidence signal.
The formal justification is marginalisation. What you want is the most probable answer, but greedy decoding gives you the answer at the end of the most probable reasoning path. Those are not the same thing. If four distinct valid derivations each lead to 504, each with modest individual probability, their combined mass can easily exceed one slick, high-probability path that lands on 792. Voting sums probability mass over paths and thereby approximates
— where r ranges over reasoning paths and x is the prompt. Greedy decoding maximises over (a,r) jointly; self-consistency marginalises r away.
Why it works — three separate reasons
Reason one: correct paths converge, wrong paths scatter. There is usually one right answer and many distinct ways to be wrong. Ten runs of the shop problem produced 504 seven times; the errors split across two different values. Correct reasoning is an attractor and errors are diffuse, so the mode of the distribution is biased towards truth.
Reason two: majority voting beats an individual voter — provided the individual is better than chance. With k independent samples each correct with probability p, the probability the majority is correct is
At p=0.6:
| Samples k | Majority correct | Gain over one sample | Relative cost |
|---|---|---|---|
| 1 | 60.0% | — | 1× |
| 3 | 64.8% | +4.8 pts | 3× |
| 5 | 68.3% | +8.3 pts | 5× |
| 7 | 71.0% | +11.0 pts | 7× |
Check the k=5 row by hand: (35)(0.6)3(0.4)2=10×0.216×0.16=0.3456, plus (45)(0.6)4(0.4)=5×0.1296×0.4=0.2592, plus (0.6)5=0.0778. Total 0.6826.
Look at the shape of that column. The first extra pair of samples buys 4.8 points; the next pair buys 3.5; the next 2.7. Cost is linear, benefit is concave. There is no setting of k at which voting is free.
Reason three — and this one is a warning, not a benefit. Run the same formula at p=0.4 and the majority is correct only 31.7% of the time. You paid five times as much to become worse than a single sample. Voting amplifies whatever the model's dominant tendency is. If the model systematically misreads "markup" as "margin", every sample makes the same mistake and the vote is unanimous and wrong.
Majority voting is a variance reducer, not a bias corrector. Below 50% accuracy it actively hurts, and it never fixes an error the model makes consistently.
What the research says, and what it does not
The self-consistency result (Wang et al., 2022) — sample many chains, take the majority — produced large, repeatable gains on arithmetic and commonsense reasoning benchmarks, in the region of ten to twenty accuracy points on grade-school maths word problems, using tens of samples. That is a real and important finding, but it was measured on models without built-in thinking.
Five things it does not establish, all of which get assumed anyway:
- It does not say the majority answer is correct. It says it is correct more often. A 9-out-of-10 vote on a question the model misunderstands is 9 identical mistakes.
- The gains are benchmark-specific. They were measured on tasks with a single short verifiable answer. On open-ended generation there is nothing to count.
- The independence assumption is false. The binomial formula assumes voters fail independently. Your samples share weights, share a prompt, and share whatever bias the training data instilled. Real gains are consistently smaller than the table predicts — treat that table as an upper bound.
- Most of the benefit arrives early. The curve is steep from 1 to about 5–10 samples and close to flat beyond 20–40. Paying for 40 samples to gain the last point is rarely defensible.
- It does not say the gain carries over to reasoning models. A model with built-in thinking already explores and checks alternatives inside one call, so its single-sample accuracy p starts higher and there is less left for voting to recover. Measure it on your own task rather than assuming the original gains.
Prompting for it
The single most common implementation error:
BADSolve this problem five different ways and tell me which answerappears most often.A shop buys 240 units at 7.50 each...This is one generation, not five. Attempt 2 is written in a context that already contains attempt 1, so it is conditioned on it — and the most probable continuation after a completed derivation ending in 504 is another derivation ending in 504. You have not sampled five paths; you have sampled one path and four echoes of it. The "vote" that follows is theatre, and the reported agreement is meaningless because the samples were never independent.
GOOD (one prompt, sent k times as separate requests)Solve this problem. Work through it step by step, showing eachcalculation. Then give your answer on a final line in exactlythis form, with no units, symbols or commas:ANSWER: <number>A shop buys 240 units at 7.50 each...Two design choices are doing work. Separate requests give genuinely independent draws. The rigid ANSWER: line makes extraction a one-line regular expression rather than a parsing problem — and if extraction is unreliable, your vote counts are garbage regardless of how good the reasoning was. Specifying "no units, symbols or commas" prevents 504, "504", "£504" and "504.00" from being counted as four different answers.
Deliberate diversity injection
When the default sampler gives paths that are too similar — and on a well-trained model asked a familiar question, they often are — vary the prompt instead of the sampler. Each variant nudges the model into a different region of its own distribution:
1FRAMINGS = [2 "Work through it step by step, showing each calculation.",3 "Start by listing every quantity you are given and what is being asked. Then compute.",4 "Solve it, then check your answer by computing the total a second, different way.",5 "Define variables for each unknown, write the equations, then solve them.",6 "Reason about the units at each step to make sure they are consistent.",7]89def sample_paths(problem, k=5):10 outs = []11 for i in range(k):12 resp = client.messages.create(13 model="claude-opus-5",14 max_tokens=4000, # leaves room for thinking tokens15 messages=[{"role": "user", "content":16 f"{FRAMINGS[i % len(FRAMINGS)]}\n\n{problem}\n\n"17 f"End with a final line: ANSWER: <number>"}],18 )19 outs.append("".join(b.text for b in resp.content if b.type == "text"))20 return outsThis is a strictly better source of diversity than temperature for one reason: high temperature produces variation by flattening the whole distribution, which makes every token noisier including the ones that were right. Prompt variation changes the approach while leaving each individual path decoded at normal sharpness. You get different routes, not sloppier ones.
Aggregating
1import re2from collections import Counter34def extract(text):5 m = re.search(r"ANSWER:\s*(-?[\d.]+)", text)6 if not m:7 return None8 return round(float(m.group(1)), 2) # canonicalise 504 / 504.0 / 504.00910def vote(paths):11 answers = [a for a in map(extract, paths) if a is not None]12 if not answers:13 return {"answer": None, "reason": "no parseable answers"}14 counts = Counter(answers)15 (top, n_top), = counts.most_common(1)16 runner_up = counts.most_common(2)[1][1] if len(counts) > 1 else 017 return {18 "answer": top,19 "agreement": n_top / len(answers), # share of valid votes20 "margin": n_top - runner_up, # lead over second place21 "parse_rate": len(answers) / len(paths),22 "distribution": dict(counts),23 }On our ten runs this returns answer 504, agreement 0.7, margin 5, distribution {504: 7, 792: 2, 468: 1}. Keep all four fields. margin is the honest confidence signal — 7-versus-2 is a different situation from 4-versus-3 even though both are majorities. And parse_rate is the one people forget: if it drops below 1.0 your prompt's output contract is leaking, and a silent extraction failure looks exactly like a rare answer.
Aggregation strategies compared
| Strategy | How it works | Use when | Failure mode |
|---|---|---|---|
| Exact-match majority | Count canonicalised strings or numbers | Numeric or single-label answers | Formatting variants split the vote |
| Tolerance majority | Cluster values within a tolerance, e.g. ±1% | Continuous estimates, forecasts | Tolerance too wide merges genuinely different answers |
| Semantic clustering | Embed each answer, cluster, take the largest cluster's medoid | Short free-text answers | Embedding similarity is not logical equivalence |
| Self-reported confidence weighting | Ask each path to rate its confidence, weight votes by it | Rarely — see below | Stated confidence is poorly calibrated and correlates with fluency, not correctness |
| Judge selection | Show all k paths to a fresh call, ask which reasoning is soundest | Open-ended answers with no countable form | Judges favour longer and more confident text; position bias |
| Consistency selection | Show all k answers, ask which is most consistent with the others | Free text where you still want the modal view | Reduces to "pick the most typical", which is not "pick the best" |
The confidence-weighting row is worth dwelling on, because it is intuitively appealing and usually harmful. A model's stated confidence is generated by the same process as its answer — it is a prediction of what a confident-sounding sentence looks like here, not a read-out of internal probability. In the shop problem the 792 paths tend to sound more assured than the 504 paths, because the margin interpretation produces a cleaner-looking calculation. Weight by stated confidence and you can push the vote towards the wrong answer. The raw count is dumber and more honest.
Self-consistency with a tool-using loop
You can vote over whole trajectories, not just final numbers: run k independent agent loops on the same question and take the majority final answer. Two constraints make this practical or impractical.
Side effects. Voting requires running every path to completion. If a path sends an email, charges a card or writes a row, running five of them does that five times. Trajectory-level voting is safe only over read-only tools — searches, lookups, calculators. If any tool mutates state, either mock it during the voting phase or restrict voting to the planning stage and execute only the winning plan.
Cost multiplies twice. Five trajectories at eight steps each is forty model calls and forty tool calls, against eight of each for a single run. If the tools are rate-limited or billed per call, the tool bill may dominate the model bill.
A cheaper hybrid, worth knowing: vote at the decision points rather than over whole trajectories. Run one loop, but at each step where the model must choose a tool, sample three candidate actions and take the majority. You get most of the stability at a fraction of the cost, because you are only paying triple on short action-selection generations, not on the entire transcript.
Deciding whether to pay for it
Take concrete numbers. Suppose a query is 400 input tokens and 900 output tokens, and you are billed 5 dollars per million input tokens and 25 dollars per million output tokens.
- Input: 400×5/106=0.0020 dollars.
- Output: 900×25/106=0.0225 dollars.
- One call: 0.0245 dollars. Five calls: 0.1225 dollars. Added cost: 0.098 dollars per query.
Now put a price on being wrong. Self-consistency at k=5 moved accuracy from 60% to 68.3%, so it prevents an error on 8.3% of queries. Let C be what one wrong answer costs you — a support ticket, a refund, a re-run, an apology.
| Cost of one error, C | Expected saving per query | Added cost | Verdict |
|---|---|---|---|
| 0.50 dollars | 0.083 × 0.50 = 0.042 | 0.098 | Not worth it — you lose 0.056 per query |
| 1.18 dollars | 0.098 | 0.098 | Break-even |
| 12 dollars | 0.083 × 12 = 1.00 | 0.098 | Clearly worth it — roughly 10× return |
The break-even point is 0.098/0.083=1.18 dollars per error. That single number replaces a great deal of arguing. Below it, ship the single call. Above it, sample.
On a reasoning model there is a third option to price before you pay for k calls: one call at a higher effort setting. It spends extra thinking on the same question, adds output tokens to one call rather than multiplying the whole call by k, and on hard questions it can close much of the same gap. Measure both on your eval set. Voting still gives you one thing a single call cannot: the agreement signal, which is what the routing below depends on.
Adaptive sampling: stop when the vote is decided
Fixed k is wasteful because easy questions are settled after two identical answers. Draw incrementally and stop on a margin:
1def adaptive_vote(problem, k_min=3, k_max=9, need_margin=2):2 answers = []3 while len(answers) < k_max:4 a = extract(one_sample(problem))5 if a is not None:6 answers.append(a)7 if len(answers) >= k_min:8 c = Counter(answers).most_common(2)9 lead = c[0][1] - (c[1][1] if len(c) > 1 else 0)10 if lead >= need_margin:11 break12 counts = Counter(answers)13 top, n = counts.most_common(1)[0]14 return top, n / len(answers), len(answers)On a question the model finds easy, the first three samples agree, the margin is 3, and you stop at k=3. On a genuinely contested question you keep drawing to 9. If a typical workload is 70% easy and 30% hard, the average is 0.7×3+0.3×9=4.8 samples rather than a flat 9 — roughly half the cost for nearly the same accuracy. The low margin on the hard cases is also exactly the signal you want for routing to a human.
Where people get this wrong
| Belief | Why it is wrong |
|---|---|
| "High agreement means the answer is right." | Agreement measures how sharp the model's distribution is, not whether it points at the truth. A systematically misunderstood question yields unanimous nonsense. |
| "More samples is always better." | Gains are concave and cost is linear; and if p<0.5 more samples make it worse. There is an optimum, and it is usually 3–7. |
| "Ask for five answers in one response." | Later answers are conditioned on earlier ones, so they are correlated by construction. You get one sample plus four echoes. |
| "It works on any task." | It needs a discrete, extractable, comparable answer. For an essay or a design, "the most frequent output" does not exist. |
| "It is an ensemble, so it has ensemble properties." | A real ensemble combines models with different failure modes. Here every voter is the same weights on the same prompt, so errors are strongly correlated and the independence maths overstates the gain. |
| "Set temperature to 0 for accuracy instead." | Greedy decoding returns the most likely path, not the most likely answer — and on the latest Claude models the parameter is not even accepted. |
What this means when you build something
The lasting value of self-consistency is not the accuracy bump. It is that you get a confidence signal you can act on, computed from the model's own behaviour rather than from its claims about itself. That signal is what makes an unreliable component safe to put in a pipeline.
Concretely, this is what it lets you build:
- Three-way routing. Unanimous vote → auto-approve. Clear majority → approve with the distribution logged. Split vote → escalate to a human. You have converted "the model is 70% accurate" into "the model handles 70% of cases and flags the rest", which is a completely different operational proposition.
- A live regression detector. Track mean agreement per query type over time. A drop from 0.85 to 0.6 on invoice questions tells you something changed — a new document format, a prompt edit, a model update — before anyone files a complaint.
- An honest budget conversation. "Five samples costs 0.098 dollars more per query and breaks even if an error costs more than 1.18 dollars" is a sentence a finance team can act on. "It seems more reliable" is not.
Two boundaries to hold firmly. Never present agreement to a user as a probability of correctness — it is not calibrated and treating it as calibrated is how a 90%-agreement wrong answer ends up unchallenged in a report. And never let self-consistency stand in for verification when verification is available. If you can check the answer — run the arithmetic, query the database, execute the test — a single check beats any number of votes, because it is grounded in something outside the model. Vote when you cannot verify; verify when you can.