Course Content
AI Agent Fundamentals
5 sections · 13 lessons
ReAct (Reason + Act) Loop
Someone asks an agent: "Did we overspend on cloud compute in July, and by how much?"
The first design plans everything up front, then executes the plan:
PLAN 1. query_spend("2026-07") 2. get_budget("platform", "2026-07") 3. calculator(spend - budget) 4. reportEXECUTE 1. query_spend -> { total_usd: 41820, note: "excludes reserved instances" } 2. get_budget -> 50000 3. calculator -> 41820 - 50000 = -8180 4. report -> "Under budget by 8,180 dollars."The answer is wrong. The real July spend was 54,220 dollars and the team was over budget by 4,220. The plan missed it because step 1 returned a note — "excludes reserved instances" — that nobody was in a position to read. The plan had been fixed before that note existed, and step 2 executed regardless.
A human analyst would have stopped dead at that note. That is the whole idea behind ReAct: put a thought between every action and the next one, so the agent gets a chance to notice.
What ReAct is
ReAct stands for Reason + Act. It interleaves two things that most systems keep apart:
- Reasoning traces — the model writes out what it is thinking before it does anything.
- Actions — the model calls a tool and receives a real observation from the world.
The unit of work is a triple: Thought → Action → Observation, repeated until the model produces a final answer instead of an action.
Each half fixes the other's weakness. Pure reasoning — a model thinking step by step with no tools — produces fluent chains that drift away from reality, because nothing ever contradicts them. Pure acting — a model emitting tool calls with no visible reasoning — cannot handle anything where the right next move depends on interpreting the last result. Interleave them and the reasoning stays anchored by observations while the actions stay directed by reasoning.
Reasoning without acting drifts. Acting without reasoning flails. ReAct is the loop where each one keeps the other honest.
The loop
Question: [the task]Thought 1: what do I know, what do I need, what will I doAction 1: tool_name[arguments]Observation 1: [what the tool actually returned]Thought 2: what did that tell me, does it change my planAction 2: tool_name[arguments]Observation 2: [what the tool actually returned]...Thought N: I now have everything I needAction N: finish[the answer]Three structural properties are doing the work.
The thought comes before the action, every time. Not once at the start. The model is forced to justify each action given everything observed so far, which is the mechanism that lets it change course.
Observations are appended verbatim. The model does not get a summary of what happened; it gets what happened. Summarising observations is a common optimisation that quietly destroys the loop, because the detail you compress away is exactly the detail that would have changed the plan.
Finishing is an action. finish[...] sits in the same slot as any tool. This means the decision to stop is made by the same reasoning that decides everything else, rather than being imposed by a counter.
What belongs in a reasoning trace
A trace is not the model narrating. It is a short piece of work with four jobs:
| Job | Question it answers | Weak version | Strong version |
|---|---|---|---|
| Assess | What do I now know? | "I have some data." | "Spend is 41,820 but flagged as excluding reserved instances." |
| Identify the gap | What is missing? | "I need more information." | "I am missing the reserved-instance figure for July." |
| Choose | Which tool, and why that one? | "I will search." | "query_reserved returns exactly that, and takes a month string." |
| Predict | What should I expect back? | — | "A dollar figure, probably 5,000–20,000. Anything outside that is suspicious." |
The prediction line is the most underused and the most valuable. An agent that says what it expects can recognise when reality disagrees. Without it, a tool returning 0 or null gets absorbed into the reasoning without comment and poisons everything downstream.
Contrast the two traces for the same situation:
WEAKThought: I have the spend. Now let me get the budget.STRONGThought: query_spend returned 41820 with note "excludes reserved instances". That means 41820 is NOT total spend, so comparing it to the budget now would understate us. I need the reserved-instance cost for the same month before any comparison. query_reserved(month) gives that. I expect a five-figure number; if it comes back 0 I should check whether the team uses reserved instances at all.The weak trace is what produced the -8,180 answer. It is not that the model failed to think — it is that it was never asked to look at the observation before choosing the next action.
Building one
Step 1 — the tools
1def query_spend(month: str) -> dict:2 """On-demand cloud spend for a month. month format: YYYY-MM."""3 data = {"2026-07": {"total_usd": 41820,4 "note": "excludes reserved instances"}}5 if month not in data:6 return {"error": f"no spend data for {month}"}7 return data[month]89def query_reserved(month: str) -> dict:10 """Reserved-instance cloud cost for a month. month: YYYY-MM."""11 data = {"2026-07": {"reserved_usd": 12400}}12 if month not in data:13 return {"error": f"no reserved data for {month}"}14 return data[month]1516def get_budget(team: str, month: str) -> dict:17 """Approved budget for a team in a month."""18 return {"budget_usd": 50000, "team": team, "month": month}1920def calculator(expression: str) -> dict:21 """Evaluate an arithmetic expression. Digits and + - * / ( ) only."""22 if not re.fullmatch(r"[\d\s+\-*/().]+", expression):23 return {"error": "expression contains disallowed characters"}24 try:25 return {"result": eval(expression, {"__builtins__": {}}, {})}26 except Exception as e:27 return {"error": f"{type(e).__name__}: {e}"}2829TOOLS = {"query_spend": query_spend, "query_reserved": query_reserved,30 "get_budget": get_budget, "calculator": calculator}Two details worth copying. Every tool returns a dict with either data or an error key, never a bare exception — that keeps failures inside the observation channel where the model can reason about them. And calculator validates its input against a character whitelist before evaluating, because a tool that runs arbitrary strings from a language model is a remote code execution hole.
Step 2 — the prompt
1PROMPT = """Answer the question by cycling through2Thought / Action / Observation.34Available actions:5{tool_docs}6 finish[answer] -- use when you have the complete answer78Rules:9- Exactly one Thought and one Action per turn. Then stop.10- Never write an Observation yourself; it will be given to you.11- If an observation is surprising, incomplete, or carries a caveat,12 say so in the next Thought and adjust.13- Do not compute arithmetic in your head. Use calculator.1415Question: {question}1617{trace}""""Never write an Observation yourself" earns its place. Without it, models happily hallucinate the tool's reply and continue as if it were real — producing a complete, confident, entirely fictional trace. The stop sequence in the API call enforces the same thing mechanically.
Step 3 — the loop
1import re23ACTION_RE = re.compile(r"Action:\s*(\w+)\[(.*?)\]\s*$", re.S)45def react(question, llm, tools, max_steps=8):6 trace = ""7 for step in range(1, max_steps + 1):8 out = llm(PROMPT.format(tool_docs=describe(tools),9 question=question, trace=trace),10 stop=["Observation:"])11 trace += out1213 m = ACTION_RE.search(out)14 if not m:15 trace += ("\nObservation: could not parse an Action. Reply "16 "with exactly: Action: tool_name[arguments]\n")17 continue1819 name, arg = m.group(1), m.group(2).strip()2021 if name == "finish":22 return arg, trace2324 if name not in tools:25 obs = (f"unknown tool '{name}'. Available: "26 f"{', '.join(tools)}")27 else:28 try:29 obs = tools[name](**parse_args(arg))30 except Exception as e:31 obs = {"error": f"{type(e).__name__}: {e}"}3233 trace += f"\nObservation: {obs}\n"3435 return "Step budget exhausted without an answer.", traceNotice that all four things that can go wrong — unparseable output, unknown tool, bad arguments, tool exception — end up as an Observation the model reads next turn. None of them raise. That single design choice is why a ReAct agent recovers from a typo instead of crashing on it.
Step 4 — a test double for the model
To develop the loop you need a stand-in for the LLM that behaves deterministically. A scripted one is enough:
1class ScriptedLLM:2 """Replays a fixed sequence of model turns, ignoring the prompt."""3 def __init__(self, turns):4 self.turns, self.i = turns, 056 def __call__(self, prompt, stop=None):7 turn = self.turns[self.i]8 self.i += 19 return turn1011llm = ScriptedLLM([12 "Thought: I need July on-demand spend first.\n"13 "Action: query_spend[2026-07]",1415 "Thought: 41820 is flagged as excluding reserved instances, so it "16 "is not the total. I need the reserved figure before comparing.\n"17 "Action: query_reserved[2026-07]",1819 "Thought: Total is 41820 + 12400. Compute it exactly.\n"20 "Action: calculator[41820 + 12400]",2122 "Thought: Total is 54220. Now the approved budget.\n"23 "Action: get_budget[platform, 2026-07]",2425 "Thought: 54220 against a 50000 budget. Compute the overage.\n"26 "Action: calculator[54220 - 50000]",2728 "Thought: Over by 4220. As a percentage of budget: 4220/50000.\n"29 "Action: calculator[4220 / 50000]",3031 "Thought: 0.0844, so 8.44 percent over. I can answer.\n"32 "Action: finish[July cloud spend was 54,220 USD (41,820 on-demand "33 "plus 12,400 reserved) against a 50,000 budget - over by 4,220 "34 "USD, or 8.44 percent.]",35])The trace, in full
Question: Did we overspend on cloud compute in July, and by how much?Thought 1: I need July on-demand spend first.Action 1: query_spend[2026-07]Observation 1: {'total_usd': 41820, 'note': 'excludes reserved instances'}Thought 2: 41820 is flagged as excluding reserved instances, so it is not the total. I need the reserved figure first.Action 2: query_reserved[2026-07]Observation 2: {'reserved_usd': 12400}Thought 3: Total is 41820 + 12400. Compute it exactly.Action 3: calculator[41820 + 12400]Observation 3: {'result': 54220}Thought 4: Total is 54220. Now the approved budget.Action 4: get_budget[platform, 2026-07]Observation 4: {'budget_usd': 50000, ...}Thought 5: 54220 against a 50000 budget. Compute the overage.Action 5: calculator[54220 - 50000]Observation 5: {'result': 4220}Thought 6: Over by 4220. As a fraction of budget: 4220/50000.Action 6: calculator[4220 / 50000]Observation 6: {'result': 0.0844}Thought 7: 8.44 percent over. Answer.Action 7: finish[...]Check the arithmetic: 41,820 + 12,400 = 54,220. 54,220 − 50,000 = 4,220. 4,220 ÷ 50,000 = 0.0844, so 8.44%. The plan-and-execute version reported "under budget by 8,180" — wrong in sign and off by 12,400 dollars, which is precisely the number it never went and fetched.
The whole difference is Thought 2. That one paragraph reads the caveat, recognises the figure is partial, and inserts a step the original plan did not contain. Everything downstream follows.
Why "use the calculator" is not fussiness
Steps 3, 5 and 6 all delegate arithmetic. Language models compute by pattern-matching over digit sequences, not by arithmetic, and they are unreliable on multi-digit sums in a way that is hard to predict — usually right, occasionally off by one digit, always equally confident. A calculator tool costs one extra loop iteration and removes the entire class of error. Any number that appears in a final answer should have come from a tool.
Errors inside the loop
ReAct's real advantage is that failure is just another observation. Four kinds, each with a distinct handling:
| Failure | Observation the agent should see | What it should do next |
|---|---|---|
| Unparseable action | could not parse an Action. Reply with exactly: Action: tool_name[args] | Reformat and retry |
| Unknown tool | unknown tool 'search_web'. Available: query_spend, query_reserved, get_budget, calculator | Pick a real tool |
| Bad arguments | TypeError: query_spend() missing required argument 'month' | Supply the argument |
| Tool returns no data | {'error': 'no spend data for 2026-13'} | Fix the input, or conclude it is unavailable |
A recovery in practice:
Thought 2: I need the reserved cost.Action 2: get_reserved_cost[July 2026]Observation 2: unknown tool 'get_reserved_cost'. Available: query_spend, query_reserved, get_budget, calculatorThought 3: Wrong name and wrong date format. The tool is query_reserved and months are YYYY-MM.Action 3: query_reserved[2026-07]Observation 3: {'reserved_usd': 12400}Two things made that recovery possible: the error listed the valid tools, and the loop kept going. An error message reading Error: invalid tool gives the model nothing to correct towards, and it will guess again — usually wrongly, usually repeatedly.
Write error observations as instructions, not as complaints. "Unknown tool X, available: A, B, C" is a recovery path. "Error" is a dead end.
The repeated-action failure
The most common ReAct pathology is the agent calling the same tool with the same arguments over and over. It happens when an observation is uninformative — an empty list, a generic error — so the model has no new information and its best next guess is identical to its last one.
Detect it and feed the detection back as an observation:
1seen = {}23key = (name, arg)4seen[key] = seen.get(key, 0) + 15if seen[key] > 1:6 obs = (f"You already ran {name}[{arg}] and got the same result. "7 f"Repeating it will not help. Either use a different tool, "8 f"different arguments, or finish with what you know.")9else:10 obs = tools[name](**parse_args(arg))Handling it as an observation rather than an exception matters: it gives the model an explicit chance to change strategy, and about nine times in ten it does.
ReAct against the alternatives
| Approach | Reasons? | Acts? | Adapts mid-run? | Cost | Fails by |
|---|---|---|---|---|---|
| Direct prompting | No | No | No | 1 call | Confident fabrication |
| Chain-of-thought | Yes | No | No | 1 call | Fluent reasoning from wrong facts |
| Act-only (function calling) | No | Yes | Barely | N calls | Wrong tool, no recovery |
| Plan-and-execute | Once, up front | Yes | No | 1 + N | Plan built on assumptions the data contradicts |
| ReAct | Every step | Yes | Yes | N calls | Loops; wanders on vague goals |
| Reflexion (ReAct + critique) | Every step, plus after | Yes | Yes, across attempts | N + M calls | Self-critique that agrees with itself |
The cost column deserves honesty. ReAct is expensive. Every step re-sends the whole trace, so if a step adds roughly 400 tokens, step 7's prompt carries about 2,800 tokens of accumulated trace on top of the system prompt. Costs grow faster than linearly in step count. Plan-and-execute makes one reasoning call and then runs tools cheaply — genuinely better when the steps really are independent of each other's results.
| Use plan-and-execute when | Use ReAct when |
|---|---|
| All steps are known before starting | Step n+1 depends on what step n returned |
| Tool results cannot invalidate the plan | Results can carry caveats, gaps or surprises |
| Steps can run in parallel | Steps are strictly sequential |
| Latency budget is tight | Correctness matters more than speed |
| The domain is stable and well-understood | The environment can push back |
Where people get this wrong
Letting the model write its own observations. Without a stop sequence, the model continues past Action: and invents Observation: too. The trace looks perfect. No tool ran. Always stop generation at the observation marker and always fill it from real execution.
Compressing observations to save tokens. Truncating a tool result to its first 200 characters is where a caveat like "excludes reserved instances" gets cut. If you must compress, keep every field that is not the main payload — flags, warnings, counts, notes — and compress the payload.
Treating the thought as decoration. Some implementations parse only the action and drop the thought from the trace to save tokens. This removes the entire mechanism: on the next turn the model cannot see why it did what it did, and it starts re-deriving, contradicting itself, and looping.
One thought, many actions. Models will happily emit three actions in one turn if you let them. Then you either execute all three blind — which is plan-and-execute with extra steps — or execute the first and discard the rest, wasting tokens. Parse exactly one action per turn and enforce it in the prompt.
Using ReAct where a single call would do. "Summarise this document" needs no tools and no loop. Wrapping it in ReAct pays several times the cost for a worse answer, plus a new way to fail.
What this means when you build one
Start with the trace format, not the code. Write out by hand, on paper, the exact Thought/Action/Observation sequence you want for two or three representative tasks — including one where something goes wrong. That document is simultaneously your prompt's few-shot examples, your test fixtures, and your specification for what the tools must return.
Then make your tools tell the truth about their limits. The entire lesson of the 12,400-dollar error is that query_spend did the right thing by including the note — the failure was in the architecture that could not read it. Every tool you write should surface its caveats in the returned data: what was excluded, what was approximate, what was truncated, how stale the data is. An agent can reason about a caveat it can see and cannot reason about one you left in the docstring.
Set a step cap around 8 to 12 and log the full trace of every run. When something goes wrong, the trace tells you exactly which Thought went astray, and the fix is almost always in the observation just above it rather than in the model.
And route arithmetic, dates and lookups through tools without exception. If a number appears in your agent's final answer and you cannot point at the observation it came from, it is not a fact — it is a plausible-looking string, which is precisely what the model is built to produce.