Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

What is self-consistency prompting?


Five runs of the same dispute classifierdupdupfrauddupdup01234one runslippeddup is duplicate_charge. 4 of 5 agree, so it is handled automatically; 3 of 5 would go to an analyst.
Errors scatter across runs while the right answer repeats, so the agreement rate doubles as a confidence score.

What you need to know

How it works

  1. Send the same prompt N times (often 3 to 10), with sampling that allows variety.
  2. Extract only the final answer from each run.
  3. Vote: the most frequent answer wins.
  4. Use the agreement rate as confidence: 5 out of 5 is strong; 3 out of 5 is weak.

It needs answers that can be compared exactly — a label, a number, a yes/no. For free text, a variant called universal self-consistency asks a model to pick the response most consistent with the others.

Sampling on current models

The classic recipe says "non-zero temperature". On models that reject a temperature setting, the default sampling already varies between runs, so you simply make N calls.

Costs and trade-offs

  • Cost and latency scale with N. Run the calls in parallel to keep latency near one call.
  • Diminishing returns — most of the gain comes from the first 5 samples.
  • Reasoning models already explore several lines of thought internally, so gains are smaller. Its best modern use is as a confidence signal for routing, not as an accuracy trick.

A cheaper pattern

Run one call first. Run extra samples only when the case is high-value or the first answer's confidence is low.

A real-life example

A fintech app classifies disputes into fraud, merchant_dispute, duplicate_charge and other. Wrong fraud labels freeze customer cards, which is expensive. They run the classifier five times per dispute:

Text
Classify the dispute. Explain briefly, then on the last line writeLABEL: <one of fraud, merchant_dispute, duplicate_charge, other><dispute>I see two charges of Rs 1,499 from the same shop at 7:02 pm.I only bought once. I did not share my OTP.</dispute>

The five final labels come back duplicate_charge, duplicate_charge, fraud, duplicate_charge, duplicate_charge. Majority: duplicate_charge, agreement 4 of 5. The team's rule: act automatically at 5 of 5 or 4 of 5, and send 3 of 5 or worse to a human analyst. About 12% of disputes go to humans, and wrong automatic fraud labels drop by more than half compared with a single call.

Follow-up questions to expect

  • "How is this different from asking the model to double-check itself?" — Self-checking uses one path that can repeat its own error; self-consistency uses independent paths and votes.
  • "How many samples?" — Start at 5 and measure; accuracy usually flattens after that while cost keeps rising.
  • "Can you use it for summaries?" — Not with plain voting. Use universal self-consistency or a judge that picks the most consistent candidate.