Course Content
Advanced Prompting and Reasoning
3 sections · 7 lessons
ReAct — Reasoning and Acting in Language Models
A company builds an HR assistant. An employee asks: "How many holiday days do I have left this year?" The assistant, prompted to think step by step, replies:
Let me work through this. The standard annual allowance at mostUK companies is 25 days. You joined in March, so you accruepro-rata: 25 × (10/12) = 20.8, round to 21 days. Recordssuggest you have taken 12 days so far. That leaves 21 - 12 = 9days remaining. You should book them before 31 December.Every sentence is grammatical. The arithmetic is correct. The reasoning is legible and the tone is confident. And it is entirely fictional. This particular company gives 28 days, the employee joined in January not March, and they have taken 19 days, not 12. The true answer is 9 days — the assistant got the right number by coincidence, from three wrong inputs. Next month, with different numbers, it will be wrong and just as confident.
This is not a failure of reasoning. The reasoning was fine. It is a failure of grounding: nothing in the chain came from the world. Every token — including "25", "March" and "12" — was sampled from the model's own distribution. Chain-of-thought gives a model more computation; it does not give it more information. ReAct is the pattern that fixes exactly this gap, and understanding why it works requires being precise about what a language model can and cannot do inside a single forward pass.
The mechanism: fixed compute per token, no hidden scratchpad
Two properties of transformer decoding explain almost every reasoning technique worth knowing. Get these right and the rest follows.
Property one: compute per token is fixed
When a model generates one token, it runs the same computation every time: the same number of layers, the same matrix sizes, the same attention heads. A model with 80 layers runs 80 layers for the token "the" and 80 layers for the token that completes a hard integral. There is no internal dial labelled "think harder for this one".
So if a problem needs more computation than one forward pass provides, there is exactly one way to buy it: generate more tokens. Each generated token is another full forward pass, conditioned on everything written so far. Writing out "first I need the allowance, then subtract days taken" is not a stylistic flourish — it is the model allocating additional passes to the problem, with the intermediate results carried forward in the context.
Reasoning tokens are not an explanation of the computation. They are the computation. A model that answers in one token has spent one forward pass on the problem, no matter how hard the problem was.
Property two: the context is the only working memory
Between one generated token and the next, nothing persists except the sequence itself (and its cached key/value tensors, which are just a re-encoding of that sequence). There is no register file, no variable store, no private notepad. If an intermediate value is not written into the text, it is gone.
This is why "work it out silently and give me only the answer" is a bad instruction for a hard problem on a model that has no separate thinking step: you have forbidden it from using the only working memory it has.
Current reasoning models change the practical advice, not the mechanism. Models with built-in thinking (on the Claude API, adaptive thinking, with an effort setting from low to max) write their working into thinking blocks before the visible answer. The extra tokens are still the computation; they are just produced without you asking. On such a model, "think step by step" adds little, and asking for a bare final answer no longer starves it. The instruction still matters on models without thinking, or with thinking switched off.
What neither property gives you
Both properties are about computation and memory. Neither is about truth. Every token in a chain of thought is drawn from the model's learnt distribution over plausible continuations. A chain can therefore be long, internally consistent, arithmetically valid and factually invented — exactly what happened to the HR assistant.
The only way to inject something the model did not generate is to place it in the context from outside. That is the whole idea.
From a chain to a loop
Chain-of-thought is a straight line: prompt in, one long generation out, answer at the end. ReAct — short for Reason + Act — cuts that line into segments and splices external text in between.
The cycle has three roles:
- Thought — model-generated. What do I know, what is missing, what should I do next.
- Action — model-generated. A structured request to something outside the model: a search, a database query, a calculator, an API call.
- Observation — not model-generated. The literal result of executing that action, pasted into the context.
The loop repeats until the model emits a final answer instead of an action.
Thought 1 → Action 1 → [execute] → Observation 1 ↓Thought 2 → Action 2 → [execute] → Observation 2 ↓Thought 3 → Final AnswerThe critical property: an Observation enters the context as ordinary tokens, so every later forward pass attends to it exactly as it attends to the model's own words — but its content was never sampled. It came from a database, an index, an interpreter. That single asymmetry is what converts a plausible story into a checked one.
Chain-of-thought buys compute. ReAct buys compute and evidence. The Observation is the only part of the transcript the model could not have hallucinated.
ReAct versus plain function calling
Modern APIs let a model emit a structured tool call directly. People often assume this makes ReAct obsolete. It does not — they solve different halves of the problem.
| Plain function calling | ReAct | |
|---|---|---|
| What it standardises | The format of a tool request | The control loop around tool requests |
| Number of calls | Often one round, then answer | Many rounds, each informed by the last |
| Explicit reasoning | Optional, frequently absent | Required — a Thought precedes each Action |
| Recovery from a bad result | Undefined; often just answers anyway | Next Thought sees the failure and re-plans |
| Debuggability | You see calls, not motives | You see the stated reason for every call |
In production you normally use both: the structured tool-call mechanism for the wire format, and the ReAct discipline — think, act, observe, re-think, with a step budget and a stopping rule — for the loop. Function calling without a loop is a single lookup. ReAct without structured calls means parsing free text, which breaks.
Be clear about what survives. The original text protocol, with literal Thought:, Action: and Observation: lines, is now mostly a teaching device; production agents use native tool calling. On a reasoning model the Thought has moved into the model's thinking, which current Claude models do not return in raw form (you get a summary or an empty block, depending on the display setting). What remains of ReAct is the loop and its rules: evidence from tools, one decision per turn, errors as observations, a step budget and a stopping rule.
Anatomy of a real transcript
Here is the HR question done properly. Note that no number in the final answer originates in the model.
Question: How many holiday days do I have left this year?Thought 1: I need three facts: the employee's annual entitlement,their start date (for pro-rata accrual), and days already booked.I have none of them. Start with the employee record.Action 1: hr_lookupAction Input 1: {"employee_id": "E-4471", "fields": ["entitlement_days", "start_date"]}Observation 1: {"entitlement_days": 28, "start_date": "2019-01-14"}Thought 2: Entitlement is 28, and they started in 2019, so this isa full accrual year — no pro-rata. Now I need days taken.Action 2: leave_recordsAction Input 2: {"employee_id": "E-4471", "year": 2026, "status": "approved"}Observation 2: {"days_taken": 19, "days_pending_approval": 2}Thought 3: 28 entitlement minus 19 approved = 9. There are also 2days pending. Whether pending days count depends on policy, and Ishould not guess. Check the policy.Action 3: policy_lookupAction Input 3: {"topic": "pending_leave_balance"}Observation 3: "Pending requests are deducted from the displayedbalance to prevent over-booking."Thought 4: So the displayed balance deducts pending days too:28 - 19 - 2 = 7. I have all three facts from source systems.Final Answer: You have 7 days remaining. That is 28 days ofentitlement, less 19 already taken and 2 currently awaitingapproval. If either pending request is declined, your balancereturns to 9.Three things are worth pointing at.
Thought 2 does real work. It reads the start date, decides pro-rata does not apply, and moves on. That decision is one forward-pass-per-token of computation spent on a branch, and it is visible.
Thought 3 recognises an unknown and refuses to guess. The naive version invented "most UK companies give 25 days". Here the model notices a policy question and issues a third action instead. The behaviour is prompt-driven, and you have to ask for it explicitly.
The final answer is auditable. Every figure — 28, 19, 2 — traces to an Observation. If the balance is wrong, the bug is in leave_records, not in the model, and you can prove it from the transcript.
Designing the prompt
A ReAct prompt has to do three jobs: fix the output grammar, describe the tools, and set the rules of engagement. Most failures come from doing the third job badly or not at all.
The mistake almost everyone makes first
BADYou are a helpful assistant with access to tools. Use them whenneeded to answer the user's question accurately. Think step bystep and be thorough.Tools: search, calculator, databaseMechanically, this prompt is nearly empty. "Use them when needed" gives no decision rule, so the model falls back on its prior, which strongly favours answering directly — the overwhelming majority of question-answer text in training data has no tool call in it. "Tools: search, calculator, database" gives no argument shapes, so any call it does emit is a guess at a schema. And nothing defines what the model should do when a tool returns an error, which means the next token after a failure is sampled from "what usually follows an error message in text", and what usually follows is an apology and a plausible answer.
GOODYou answer questions by alternating reasoning and tool use.FORMAT — emit exactly one block at a time, then stop: Thought: <what you know, what is missing, what to do next> Action: <one tool name> Action Input: <a JSON object matching that tool's schema>...or, when you are finished: Thought: <why you now have enough> Final Answer: <the answer>RULES1. Any company-specific fact — balances, dates, policy, prices, names — must come from an Observation. If you have not observed it, you do not know it.2. One Action per turn. Do not write an Observation yourself; it will be supplied.3. If an Observation contradicts what you expected, the Observation wins. State the contradiction in your next Thought.4. If a tool errors, your next Thought must say what you will change — different arguments, different tool, or give up. Do not retry an identical call.5. Budget: 8 Actions. At step 8, answer with what you have and state plainly what remains unknown.TOOLShr_lookup(employee_id: str, fields: list[str]) → dict Employee master record. Valid fields: entitlement_days, start_date, manager_id, contract_type.leave_records(employee_id: str, year: int, status: str) → dict status is one of: approved, pending, all.policy_lookup(topic: str) → str Free-text HR policy search. Returns the matching clause.Each rule maps to a specific failure the loop would otherwise have. Rule 1 blocks the invented-25-days failure. Rule 2 stops the model writing its own Observations — a genuinely common and very dangerous behaviour, because a self-written Observation looks identical to a real one downstream. Rule 3 counteracts the prior overriding evidence. Rule 4 breaks retry loops. Rule 5 bounds cost and forces a graceful ending rather than a silent stall.
Rule 2 is the one people leave out and regret. If the model hallucinates an Observation, the transcript still looks perfect — you have lost the one signal that distinguished evidence from invention.
Tool descriptions are prompt, not documentation
The tool description is read by the model at selection time and by nobody else. It is competing for attention with everything else in the context. Judge it by whether it makes the right tool obviously right and the wrong tool obviously wrong.
| Weak description | Strong description | Failure it prevents |
|---|---|---|
| "Search the database." | "Search internal product catalogue by SKU or exact product name. Returns price, stock, warehouse. Does not cover discontinued lines." | Model uses it for customer data, gets nothing, concludes the customer does not exist |
| "Get user info." | "Fetch one user by numeric user_id. If you only have an email, call find_user_by_email first." | Model passes an email into an integer field and the call 400s |
| "Calculator." | "Evaluate one arithmetic expression. Use for any multiplication, division or percentage — do not compute these in your head." | Silent arithmetic slips inside a Thought |
That last one deserves a comment. Multi-digit arithmetic is a genuinely poor fit for fixed per-token compute: carrying digits is a serial algorithm and the model must approximate it in a fixed number of layers. Pushing arithmetic to a calculator tool is not a nicety; it moves a serial algorithm onto hardware that does serial algorithms exactly.
Implementing the loop
Two implementations follow. The first parses text and exposes the mechanism; the second uses structured tool calls and is what you should actually ship.
Text-mode ReAct, to see the machinery
1import json, re, anthropic23client = anthropic.Anthropic()45TOOLS = {6 "hr_lookup": lambda a: hr_db.fetch(a["employee_id"], a["fields"]),7 "leave_records": lambda a: leave_db.query(a["employee_id"], a["year"], a["status"]),8 "policy_lookup": lambda a: policy_index.search(a["topic"]),9}1011def run_react(question, system_prompt, max_steps=8):12 transcript = f"Question: {question}\n"13 for step in range(max_steps):14 resp = client.messages.create(15 model="claude-opus-5",16 max_tokens=4000, # thinking tokens count against this17 system=system_prompt,18 messages=[{"role": "user", "content": transcript}],19 stop_sequences=["Observation:"], # hand control back to us20 )21 # thinking is on by default, so a thinking block may come first22 chunk = "".join(b.text for b in resp.content if b.type == "text")23 transcript += chunk2425 if "Final Answer:" in chunk:26 return transcript, chunk.split("Final Answer:")[-1].strip()2728 m = re.search(r"Action:\s*(\w+)\s*\nAction Input:\s*(\{.*?\})", chunk, re.S)29 if not m:30 transcript += "\nObservation: Malformed action block. Re-emit exactly one Action and one Action Input.\n"31 continue3233 name, raw_args = m.group(1), m.group(2)34 try:35 result = TOOLS[name](json.loads(raw_args))36 except KeyError:37 result = f"ERROR: no tool named {name}. Available: {list(TOOLS)}"38 except json.JSONDecodeError as e:39 result = f"ERROR: Action Input was not valid JSON ({e}). Re-emit it as a JSON object."40 except Exception as e:41 result = f"ERROR: {type(e).__name__}: {e}"4243 transcript += f"\nObservation: {json.dumps(result, default=str)}\n"4445 return transcript, "Step budget exhausted before reaching an answer."The stop_sequences argument is doing the essential work. Without it the model happily continues past "Action Input" and writes its own "Observation:" line — a fluent, confident, fabricated tool result. Stopping generation at that exact string makes it structurally impossible.
Notice also that every error path returns a string into the transcript rather than raising. An exception that escapes the loop teaches the model nothing. An error placed in the Observation slot is context the next forward pass can condition on, and a well-written error message ("Action Input was not valid JSON... re-emit it as a JSON object") is a corrective instruction delivered at precisely the moment it is needed.
Structured-tool-call ReAct, to ship
1import json, anthropic23client = anthropic.Anthropic()45tool_specs = [{6 "name": "leave_records",7 "description": ("Approved and pending leave for one employee in one calendar year. "8 "Use this for any question about days taken or booked."),9 "input_schema": {10 "type": "object",11 "properties": {12 "employee_id": {"type": "string", "description": "e.g. E-4471"},13 "year": {"type": "integer"},14 "status": {"type": "string", "enum": ["approved", "pending", "all"]},15 },16 "required": ["employee_id", "year", "status"],17 "additionalProperties": False,18 },19 "strict": True,20}]2122def agent(question, system_prompt, max_steps=8):23 messages = [{"role": "user", "content": question}]24 for _ in range(max_steps):25 resp = client.messages.create(26 model="claude-opus-5",27 max_tokens=4000,28 system=system_prompt,29 tools=tool_specs,30 thinking={"type": "adaptive"},31 messages=messages,32 )33 messages.append({"role": "assistant", "content": resp.content})3435 if resp.stop_reason != "tool_use":36 return "".join(b.text for b in resp.content if b.type == "text")3738 results = []39 for block in resp.content:40 if block.type != "tool_use":41 continue42 try:43 out = TOOLS[block.name](block.input) # block.input is parsed JSON44 results.append({"type": "tool_result", "tool_use_id": block.id,45 "content": json.dumps(out, default=str)})46 except Exception as e:47 results.append({"type": "tool_result", "tool_use_id": block.id,48 "content": f"{type(e).__name__}: {e}", "is_error": True})49 messages.append({"role": "user", "content": results}) # all results, one message50 return "Step budget exhausted."Three details matter here and are routinely got wrong.
strict: TruewithadditionalProperties: falseconstrains decoding so the emitted arguments are guaranteed to validate against your schema. This removes an entire class of parse failures rather than handling them.- If the model emits several tool calls in one turn, all the results must go back in a single user message. Splitting them across messages is accepted by the API but quietly teaches the model to stop making parallel calls.
- A failing tool returns a
tool_resultwithis_error: true— it is never dropped. A missing result for an emitted call leaves a dangling reference the model cannot reason about.
Adaptive thinking is worth keeping on for the loop (on Claude Opus 5 it is the default; on some earlier models you must set it). It lets the model spend extra internal reasoning tokens on the hard steps — deciding which tool, interpreting a surprising Observation — without you having to guess a fixed budget in advance. It is the same fixed-compute-per-token mechanism, allocated automatically. To trade depth for cost, lower effort in output_config rather than switching thinking off.
Failure modes, and why each one happens
| Symptom | Mechanical cause | Fix |
|---|---|---|
Model writes its own Observation: line with invented data | Nothing stopped generation; "Observation:" is a highly predictable continuation of "Action Input:" | Stop sequence in text mode; structured tool blocks in production; explicit "never write an Observation" rule |
| Model narrates the call in prose — "I'll now look up the leave records" — but emits no tool call, and the turn ends | The tool-call block was never produced, so nothing executed and no error was raised. Much more likely with internal reasoning switched off | Leave adaptive thinking on; check stop_reason == "tool_use" rather than trusting the text; treat a turn with intent-but-no-call as a retryable event |
| Observation returns 28 days; answer says 25 | The parametric prior for "typical holiday allowance" is stronger than one short retrieved token span | Rule 3 (Observation wins, state the contradiction); put observations close to the answer; keep them short and structured, not buried in prose |
| Same action repeated four times with identical arguments | Nothing in the context changed, so the next-token distribution is nearly identical to last time | Detect repeats in code and inject "You already ran this exact call and got X. Try something different." |
| 15 steps of activity, no answer | No budget and no stopping rule; "continue investigating" is always a locally plausible continuation | Hard step cap, plus a prompt rule for what to do when the cap is hit |
| Cost per query ten times the estimate | Every turn resends the whole transcript | See the arithmetic below — cache the stable prefix, trim old observations |
The cost arithmetic nobody does in advance
A ReAct loop resends the entire conversation on every turn. If your system prompt plus tool schemas is S tokens and each cycle adds about T tokens of thought, action and observation, then after n turns the total input billed is
Take a realistic case: S=1,500 (system prompt and four tool schemas), T=800, and n=12 turns.
- System prompt resent 12 times: 12×1,500=18,000 tokens.
- Growing transcript: 800×212×13=800×78=62,400 tokens.
- Total input: 80,400 tokens for one user question — against roughly 9,600 tokens of transcript actually produced.
You paid for 80,400 tokens of input to generate a 9,600-token transcript. That factor of eight is not a bug; it is what a loop costs. Two mitigations follow directly from the formula. Prompt caching attacks the nS term — the system prompt and tool schemas are byte-identical every turn, so cache them and the resends become cheap reads. Trimming or summarising old observations attacks the quadratic term, which is the one that hurts as n grows.
What this changes about how you build
The practical shift is this: once you adopt ReAct, your prompt stops being a request and becomes a specification for a control loop. You are no longer writing "please answer this well". You are writing down a decision procedure, a set of legal moves, an evidence policy, an error policy and a termination condition — and then letting a stochastic process execute it.
That reframing has consequences for the code around the model.
- Every tool needs a bad-path answer, not an exception. Timeouts, empty results, malformed arguments and permission denials all become Observations. Write those messages as instructions to the model, because that is what they are.
- Log the whole transcript, always. Every tool call, its arguments and its result, plus whatever reasoning the model shows. When something goes wrong, you can see whether the model chose the wrong tool, was given a bad result, or ignored a good one — three different bugs with three different owners. Treat the stated reasoning as a claim, not a record: in Anthropic's 2025 study of reasoning models, the models used a planted hint far more often than they mentioned it in their written reasoning. The calls and results are the facts; the rationale is a hint about why.
- Set the step budget by looking at real traces. If the median successful task takes three actions and the 95th percentile takes six, a cap of eight is generous and a cap of twenty is a licence to burn tokens.
- Treat "did the model use the right tool?" as a metric. Tool selection is next-token prediction over your descriptions. If selection accuracy is poor, the fix is nearly always in the descriptions, not the model.
- Do not reach for the loop when you do not need it. If a question is answerable from one lookup, a loop adds latency, cost and failure modes for nothing. ReAct earns its keep when the second action depends on the first one's result.
The HR assistant that invented a 25-day allowance was not badly reasoned. It was ungrounded, unbudgeted and unauditable. ReAct is what you build when those three properties actually matter.