Course Content
Prompt Engineering for LLMs
3 sections · 8 lessons
Chain-of-Thought & Advanced Reasoning Techniques
Here is a small billing problem. Work it out yourself first, on paper, and notice how many separate steps you take.
A supplier sells a part at £14 per unit. A customer orders 23 units.Orders above 20 units get 15% off the goods total. Shipping is a flat£9.50. VAT of 20% applies to the discounted goods but not to shipping.What does the customer pay in total?You almost certainly did this: multiply, subtract the discount, take VAT on the result, add shipping. Four operations, each depending on the one before it, with an intermediate number written down at every stage.
Now ask a language model for just the number — "Reply with the total only, no working." You will get a confident figure. Sometimes it is right. Often it is £339.84, which is what you get if you charge VAT on the shipping as well. Sometimes it is £395.90, which is the discount forgotten entirely. (The correct total is £337.94.) The model does not hesitate or hedge. It produces a well-formed, plausible, wrong amount.
Then ask the same model to show its working, and the accuracy climbs sharply. Same model, same weights, same question. The only thing that changed is that it was allowed to write down intermediate results before committing to an answer.
That is not a quirk. It follows directly from how the machine computes, and once you see the mechanism you will know exactly when this technique helps and when it is a waste of tokens.
Why a direct answer fails on multi-step problems
A language model generates one token at a time. To produce each token it runs the input through a fixed stack of layers — the same number of layers every time, no matter how hard the question is. There is no loop that runs longer for a difficult problem. There is no scratch memory that persists between tokens other than the text already generated.
Consider what that means for "Reply with the total only". The very next token after your question has to be the first character of the final answer. All four arithmetic steps — multiply, discount, VAT, add shipping — must be resolved inside a single forward pass through a fixed number of layers, with no intermediate value ever written anywhere.
Asking for the answer with no working is asking the model to solve a four-step problem entirely in its head, in one pass, with no scratch paper. What comes out is the most plausible-looking total, which is not the same thing as the correct one.
Now consider what happens when the model writes "23 × 14 = 322" first. That string is now in the context. Every subsequent token can attend to it. The number 322 no longer has to be recomputed or held internally — it is sitting there as text, as reliable as anything else in the prompt. The next step needs only one operation, not four. Then the next result gets written down too, and so on.
Two distinct things are happening, and it helps to keep them apart:
- More computation. Each generated token is another full forward pass. Forty tokens of working is roughly forty times the serial computation of a single-token answer. The model is not "concentrating harder"; it is genuinely getting more compute applied to the problem.
- Externalised working memory. Intermediate results become text in the context. The model reads them back rather than having to maintain them internally, which is the part it is worst at.
This is chain-of-thought prompting: asking the model to produce its intermediate reasoning as text before it produces the final answer. The name makes it sound psychological. The mechanism is not.
The simplest version, and what it is really doing
BAD:A supplier sells a part at £14 per unit. A customer orders 23 units.Orders above 20 units get 15% off the goods total. Shipping is a flat£9.50. VAT of 20% applies to the discounted goods but not to shipping.Total? Give the number only.GOOD:A supplier sells a part at £14 per unit. A customer orders 23 units.Orders above 20 units get 15% off the goods total. Shipping is a flat£9.50. VAT of 20% applies to the discounted goods but not to shipping.Work through this one step at a time. After each step, write therunning figure. Do not state the total until every step is done.Then give the final line as:TOTAL: £<amount>Why the second works. "Work through this one step at a time" makes step-by-step text the likely continuation, which buys the extra forward passes. "Write the running figure" forces each intermediate value into the context in a form the next step can attend to. "Do not state the total until every step is done" is the load-bearing line — without it, models frequently open with "The total is £339.84" and then produce working that leads somewhere else, because a token already generated cannot be revised. And the TOTAL: line gives you something a program can extract with one regular expression, instead of hunting for a number in prose.
On magic phrases
You will see "Let's think step by step" quoted as if it were an incantation. It is not. It is a phrase that, in the training distribution, is very frequently followed by enumerated reasoning — so it raises the probability of enumerated reasoning. Any phrasing that does the same job works comparably: "Work through this in order", "Show each calculation", "First establish X, then Y".
The phrase is not doing the work. Producing reasoning tokens before the answer is doing the work. Judge any prompting trick by asking what it changes about the computation, not by whether it sounds like a spell.
This matters because folklore is expensive. People carry around phrases like "you are a world-class expert" or "take a deep breath" and add them to prompts by habit, without ever measuring whether they help on their task. Some of them do help on some tasks; several published results are real. But a phrase that helped on a maths benchmark in one paper with one model is not a general law, and you can only find out by testing on your own inputs.
Showing reasoning instead of asking for it
For domain-specific reasoning, an instruction to "think step by step" leaves the shape of the steps up to the model. If the steps have a shape that matters to you, demonstrate one.
Decide whether each transaction should be flagged for review.Transaction: £4,200 to a new payee, 03:14, customer's usualmaximum single payment is £600.Reasoning: Amount is 7x the customer's usual maximum. Payee isnew. Time is outside the customer's normal activity window.Three independent risk signals.Decision: FLAGTransaction: £180 to a saved payee, 13:20, customer pays thispayee most months.Reasoning: Amount is typical. Payee is established. Time iswithin normal activity. No risk signals.Decision: ALLOWTransaction: £2,000 to a saved payee, 19:40, largest previouspayment to this payee was £1,900.Reasoning: Amount is slightly above the previous maximum but notanomalous. Payee is established. Time is normal. One weak signal.Decision: ALLOWTransaction: {new_transaction}Reasoning:This does more than the generic instruction. It fixes the features that get examined — amount relative to history, payee status, timing — so the model checks the same three things on every transaction instead of noticing whatever stands out. It fixes the label vocabulary. And the third example is deliberately a near-miss that resolves to ALLOW, which teaches the threshold rather than just the procedure.
It is also worth noting the ordering: Reasoning: comes before Decision: in every example, so the pattern the model completes puts the reasoning first. Had the examples listed the decision first, you would get decisions first, and the reasoning would be a post-hoc justification of a token already committed to. That would be strictly worse than no reasoning at all, because it would look thorough while being decorative.
Self-consistency: sample several, take the majority
Reasoning chains are generated by sampling, so with a non-zero temperature the same prompt produces different chains. A chain that goes wrong usually goes wrong in an idiosyncratic way, whereas correct chains tend to converge on the same answer. That asymmetry is exploitable.
Self-consistency means running the same chain-of-thought prompt several times at a moderate temperature, extracting the final answer from each run, and returning the most common one.
1import re2from collections import Counter34ANSWER = re.compile(r"TOTAL:\s*£?\s*([\d,]+\.\d{2})")56def solve(prompt: str, call, n: int = 5, temperature: float = 0.7):7 answers = []8 for _ in range(n):9 m = ANSWER.search(call(prompt, temperature=temperature))10 if m:11 answers.append(m.group(1).replace(",", ""))12 if not answers:13 return None, 0.014 value, count = Counter(answers).most_common(1)[0]15 return value, count / len(answers) # agreement as confidenceThe returned agreement figure is the most useful part. If five runs produce 337.94 five times, you can route the result straight through. If they produce 337.94 twice, 339.84 twice and 328.44 once, the problem is ambiguous or the prompt is underspecified, and that item deserves a human. A single call gives you an answer with no way to tell those two situations apart.
The cost is straightforwardly linear. Five samples cost five times as much and, if run sequentially, take five times as long.
| Approach | Relative cost | Gives you | Sensible for |
|---|---|---|---|
| Direct answer | 1x | An answer | Lookup, classification, simple extraction |
| Single chain of thought | 3–10x output tokens | An answer plus an auditable trace | Multi-step reasoning, arithmetic, constraint checks |
| Self-consistency, n=5 | ~5x the above | An answer plus an agreement score | High-value decisions where being wrong is expensive |
Decomposition: several calls instead of one long chain
Chain of thought keeps everything in one generation. For genuinely complex work that is often the wrong shape, because an error in step two silently corrupts steps three through eight and you find out only at the end.
Decomposition splits the task into separate calls, each with its own prompt and its own validation point.
Single long chain: "Read this 40-page contract and tell me our exposure."Decomposed: Call 1 Extract every clause that mentions liability, indemnity or termination. Return as a JSON list with clause numbers. Call 2 For each extracted clause, state the obligation it places on us in one sentence. [validate: count matches] Call 3 Given the obligations from step 2, list the three with the highest financial exposure and explain the ranking.What you gain is inspectability. After call 1 you can check whether the clause list is plausible before spending anything on the rest. If the analysis is wrong you know which stage produced the error. Each prompt is short enough to hold one clear job, so each can be tuned separately.
What you pay is orchestration: more calls, more latency, error handling between stages, and the risk of information lost at each boundary. Reach for it when a single chain is failing in ways you cannot diagnose, not before.
The reasoning you read is not necessarily the reasoning that happened
This is the caveat that keeps people honest, and it is regularly missed.
When a model writes out a chain of thought, that text is generated the same way as any other text: token by token, as a plausible continuation. There is no guarantee that it faithfully describes the computation that produced the final answer. Researchers have demonstrated this directly (Turpin et al., 2023, "Language Models Don't Always Say What They Think"): introduce a bias into the prompt (for example, always marking option (A) as correct in the few-shot examples), and models will shift towards answering (A) while producing reasoning that never mentions the pattern and instead constructs a respectable-sounding case for (A).
A chain of thought is a useful artefact and a real accuracy improvement. It is not a transcript of the model's internal process, and it is not evidence that the answer is correct.
Two practical implications. Do not present model reasoning to users as an explanation of why the system decided something, in any setting where that explanation carries weight. And when you audit outputs, check the answer against ground truth — a chain that reads beautifully and lands on the wrong number is a common and particularly convincing failure.
When chain of thought makes things worse
It is not free and it is not universally good. Four situations where it hurts.
| Situation | What goes wrong |
|---|---|
| Simple classification or retrieval | Nothing to decompose; the reasoning becomes rationalisation, and models can talk themselves out of a correct first instinct |
| Latency-sensitive paths | Reasoning tokens are generated serially — a 400-token chain is several seconds of user-visible wait |
| Strict machine-readable output | Reasoning text contaminates the payload unless you separate the channels explicitly |
| Subjective or aesthetic judgements | Verbose justification of an arbitrary preference; more words, no more accuracy |
The third one has a clean fix. Keep the reasoning, but fence it so your parser can discard it.
Work through the decision inside <scratch> tags. Then give thefinal answer inside <answer> tags as a single JSON object.Nothing outside the <answer> tags will be read by the system.<scratch>your reasoning here</scratch><answer>{"decision": "...", "confidence": 0.0}</answer>You get the accuracy benefit of reasoning tokens, a trace you can log for debugging, and a payload you can parse deterministically. This pattern is for models that do not reason on their own. On the reasoning models described next, the built-in reasoning replaces the scratch block, and the API's structured-output mode is a stronger way to get a parseable answer than tags.
Models that reason before answering by default
Most frontier models in 2026 are reasoning models. They produce hidden reasoning tokens before the visible answer, and an API setting controls how much, not your prompt text. On Anthropic's current models this is adaptive thinking with an effort level (low up to max) in output_config. On OpenAI's reasoning models it is reasoning.effort. The mechanism from the start of this lesson still applies: the reasoning is generated tokens, it costs time, and you pay for it as output tokens. What changes is who writes the scratch paper.
That has three practical effects:
- Drop the incantations. "Think step by step, then check your work, then reconsider" is redundant on these models and can get in the way, because you are hand-writing a procedure the model already runs. Describe the problem and what a good answer looks like; turn effort up if the task needs more thought.
- You usually cannot read the reasoning. Providers return a summary of the reasoning, or nothing, not the raw tokens. The faithfulness caveat above applies even more strongly to a summary.
- Structured reasoning examples still help when the features to check are specific to your domain, as in the fraud example above. That is a statement of policy, not a request to think.
The rule is the same as everywhere else in this subject: check the guidance for the specific model you are calling, and measure on your own task rather than assuming a technique transfers. A prompt tuned for a model that needed explicit chain-of-thought scaffolding may be the wrong prompt for one that does not.
When you build something with this
Decide by asking a single question about the task: does getting this right require holding an intermediate result?
If a competent human would reach for a pen — because there is arithmetic, or several constraints to satisfy at once, or a chain of dependencies — the model needs the equivalent, and the equivalent is generated tokens. If a competent human would answer instantly from recognition, the reasoning is decoration and you are paying for words.
Then get four details right, because they are where the technique is usually botched:
- Reasoning strictly before the answer. Check your examples and your output format. Any layout that lets the answer appear first has thrown away the entire benefit, since a generated token cannot be taken back.
- Structure the steps when the task has a known structure. "Check the amount against history, then the payee, then the timing" produces consistent coverage across every input. "Think step by step" produces whatever steps occurred to the model this time.
- Separate the channels. Reasoning in one fenced block, the machine-readable answer in another. Log the first, parse the second.
- Measure the trade, do not assume it. Run your test set with and without reasoning and record accuracy, tokens and latency for both. Chain of thought that costs six times the tokens for a one-point gain is a bad deal on high-volume traffic and an obvious one on a decision worth thousands of pounds. You cannot tell which you have without the numbers.