Course Content
AI Agent Fundamentals
5 sections · 13 lessons
Implementing a Simple ReAct Agent (Python)
Here is the agent almost everyone writes first. Thirty-one lines, and it works in the demo.
1def agent(question):2 trace = ""3 for _ in range(10):4 out = llm(f"Answer using tools.\n{question}\n{trace}")5 if "Final Answer:" in out:6 return out.split("Final Answer:")[1]7 action = out.split("Action:")[1].split("\n")[0]8 name, arg = action.split("[")9 result = TOOLS[name](arg.rstrip("]"))10 trace += f"{out}\nObservation: {result}\n"11 return "gave up"It survived one question. Then, over a week, it produced these:
| What happened | Line that did it |
|---|---|
IndexError: list index out of range | out.split("Action:")[1] — the model wrote prose instead |
KeyError: 'search_web ' | A trailing space in the tool name |
The model wrote its own Observation: and answered from fiction | No stop sequence on the LLM call |
| Same tool called eleven times with identical arguments | No repeat detection |
| One run cost 3.40 dollars | No token budget; the trace grew unbounded |
| A tool raised and the whole run died at step 7 | No try/except around execution |
| Two arguments needed; the parser only handles one | arg.rstrip("]") as an argument parser |
Every one of these is an engineering problem, not a prompting problem. This lesson builds the version that survives — around 150 lines, with each part justified by one of the failures above.
The five parts
┌──────────────────────────────────────────────────┐│ Agent ││ owns the loop, the budgets, the stop rules ││ ││ ┌────────────┐ ┌────────────┐ ┌────────────┐ ││ │ LLMClient │ │ Registry │ │ Trace │ ││ │ one call, │ │ tools + │ │ memory + │ ││ │ swappable │ │ schemas + │ │ compaction │ ││ │ for tests │ │ validation │ │ │ ││ └────────────┘ └────────────┘ └────────────┘ ││ ┌────────────────────────┐ ││ │ Parser │ ││ │ text -> (tool, args) │ ││ │ never raises │ ││ └────────────────────────┘ │└──────────────────────────────────────────────────┘| Failure from the table above | Component that prevents it | Mechanism |
|---|---|---|
IndexError on prose output | Parser | Returns an error tuple; the loop continues |
KeyError: 'search_web ' | Parser | Anchored regex plus .strip() on the name |
| Model invents its own observations | LLMClient | stop_sequences=["Observation:"] |
| Same call eleven times | Trace | repeated() returns a corrective observation |
| One run costs 3.40 dollars | Trace + Agent | Compaction plus an independent character budget |
| Tool exception kills the run | Registry | call() catches and returns the error as text |
| Two-argument call mis-parsed | Parser | Bracket-aware split_top_level and key=value pairs |
The separation is not architectural taste. Each boundary buys something specific: a swappable LLMClient makes the agent testable without spending money, a Registry puts validation in one place instead of scattered through the loop, a Trace object makes memory a thing you can measure and trim, and a Parser that never raises turns malformed output into a recoverable observation.
Tools and the registry
Tools are plain functions. The registry derives the schema from the signature and docstring so the description the model reads cannot drift from the code that runs.
1import inspect, json, re2from typing import Callable, Any34class Registry:5 def __init__(self):6 self.tools: dict[str, dict] = {}78 def register(self, fn: Callable) -> Callable:9 sig = inspect.signature(fn)10 params = {}11 for name, p in sig.parameters.items():12 ann = p.annotation13 params[name] = {14 "type": {str: "string", int: "integer",15 float: "number", bool: "boolean"16 }.get(ann, "string"),17 "required": p.default is inspect.Parameter.empty,18 }19 self.tools[fn.__name__] = {20 "fn": fn,21 "doc": inspect.getdoc(fn) or "",22 "params": params,23 }24 return fn2526 def describe(self) -> str:27 lines = []28 for name, spec in self.tools.items():29 args = ", ".join(30 f"{k}: {v['type']}" + ("" if v["required"] else "?")31 for k, v in spec["params"].items())32 lines.append(f" {name}({args})\n {spec['doc']}")33 return "\n".join(lines)3435 def validate(self, name: str, args: dict) -> str | None:36 """Return an error message, or None if the call is valid."""37 if name not in self.tools:38 return (f"Unknown tool '{name}'. Available: "39 f"{', '.join(self.tools)}.")40 spec = self.tools[name]["params"]41 missing = [k for k, v in spec.items()42 if v["required"] and k not in args]43 if missing:44 return (f"{name} is missing required argument(s) "45 f"{missing}. Expected: {list(spec)}.")46 unknown = [k for k in args if k not in spec]47 if unknown:48 return (f"{name} has no parameter(s) {unknown}. "49 f"Valid parameters: {list(spec)}.")50 return None5152 def call(self, name: str, args: dict) -> str:53 problem = self.validate(name, args)54 if problem:55 return problem56 try:57 return str(self.tools[name]["fn"](**args))58 except Exception as e:59 return (f"{name} raised {type(e).__name__}: {e}. "60 f"Check the arguments and try a different approach.")call returns a string on every path — success, invalid arguments, unknown tool, exception. Nothing escapes as an exception. That single property removes the "tool raised and the run died" failure permanently, and it is why the loop below has no error handling of its own.
Make one rule and hold it everywhere: nothing raises. Bad arguments, unknown tools, exceptions inside tools and unparseable model output all become strings that go back into the trace.
Some tools:
1registry = Registry()23@registry.register4def get_order(order_id: str) -> str:5 """Look up one order by ID. Format: ORD- followed by 4 digits."""6 db = {"ORD-4471": {"customer_id": 8842, "total_paise": 425000,7 "status": "delivered", "placed": "2026-07-12"}}8 if order_id not in db:9 return (f"No order '{order_id}'. IDs look like ORD-4471. "10 f"Known orders for this session: {list(db)}.")11 return json.dumps(db[order_id])1213@registry.register14def get_refund_policy(status: str) -> str:15 """Refund window in days for an order status: delivered,16 shipped, placed, or cancelled."""17 policy = {"delivered": 30, "shipped": 14,18 "placed": 90, "cancelled": 0}19 if status not in policy:20 return f"Unknown status '{status}'. Valid: {list(policy)}."21 return f"{status}: refundable within {policy[status]} days."2223@registry.register24def days_between(start: str, end: str) -> str:25 """Whole days between two ISO dates (YYYY-MM-DD)."""26 from datetime import date27 try:28 d0, d1 = date.fromisoformat(start), date.fromisoformat(end)29 except ValueError:30 return f"Dates must be YYYY-MM-DD. Got '{start}', '{end}'."31 return str((d1 - d0).days)Every error return names the fix. get_order even lists the valid IDs. This is not politeness — it is the difference between an agent that recovers on the next turn and one that guesses again.
The parser
Turning model text into a call is where the naive version broke three separate ways. The parser must handle every malformed shape and return an error rather than raising.
1ACTION_RE = re.compile(2 r"^\s*Action:\s*([A-Za-z_]\w*)\s*\[(.*)\]\s*$",3 re.MULTILINE | re.DOTALL)4FINAL_RE = re.compile(r"^\s*Final Answer:\s*(.+)", re.MULTILINE | re.DOTALL)56def parse(text: str):7 """-> ('final', answer) | ('action', name, args) | ('error', msg)"""8 final = FINAL_RE.search(text)9 if final:10 return ("final", final.group(1).strip())1112 m = ACTION_RE.search(text)13 if not m:14 return ("error",15 "No Action found. Reply with exactly one line of the "16 "form: Action: tool_name[arg=value, arg2=value] "17 "or: Final Answer: your answer")1819 name = m.group(1).strip()20 raw = m.group(2).strip()21 if not raw:22 return ("action", name, {})2324 args = {}25 for part in split_top_level(raw):26 if "=" not in part:27 return ("error",28 f"Argument '{part.strip()}' is not in key=value "29 f"form. Example: {name}[order_id=ORD-4471]")30 k, v = part.split("=", 1)31 args[k.strip()] = coerce(v.strip().strip("'\""))32 return ("action", name, args)3334def split_top_level(s: str) -> list[str]:35 """Split on commas not inside quotes or brackets."""36 out, depth, quote, buf = [], 0, None, []37 for ch in s:38 if quote:39 if ch == quote:40 quote = None41 elif ch in "'\"":42 quote = ch43 elif ch in "([{":44 depth += 145 elif ch in ")]}":46 depth -= 147 elif ch == "," and depth == 0:48 out.append("".join(buf)); buf = []; continue49 buf.append(ch)50 if buf:51 out.append("".join(buf))52 return out5354def coerce(v: str):55 if v.lower() in ("true", "false"):56 return v.lower() == "true"57 try:58 return int(v)59 except ValueError:60 pass61 try:62 return float(v)63 except ValueError:64 return vsplit_top_level exists because a naive raw.split(",") destroys search[query=orders, refunds]. Twenty lines that stop a whole family of silent argument corruption.
.strip() on the tool name kills the KeyError: 'search_web ' failure. It looks trivial and it accounted for about one failed run in twenty.
Memory
The trace is the agent's working memory, and it is the thing that made one run cost 3.40 dollars. Make it an object that knows its own size and can shrink.
1class Trace:2 def __init__(self, keep_recent: int = 4, max_chars: int = 8000):3 self.entries: list[dict] = []4 self.keep_recent = keep_recent5 self.max_chars = max_chars67 def add(self, thought, action=None, observation=None):8 self.entries.append({"thought": thought, "action": action,9 "observation": observation})1011 def repeated(self, action: str) -> bool:12 return any(e["action"] == action for e in self.entries)1314 def render(self) -> str:15 text = self._render(self.entries)16 if len(text) <= self.max_chars:17 return text18 head = self.entries[:1]19 recent = self.entries[-self.keep_recent:]20 middle = self.entries[1:-self.keep_recent]21 summary = ("[{} earlier steps compressed. Tools used: {}. "22 "Facts established: {}]").format(23 len(middle),24 ", ".join(sorted({e["action"].split("[")[0]25 for e in middle if e["action"]})),26 "; ".join(e["observation"][:80] for e in middle27 if e["observation"] and "rror" not in28 e["observation"])[:400])29 return self._render(head) + "\n" + summary + "\n" + \30 self._render(recent)3132 @staticmethod33 def _render(entries) -> str:34 out = []35 for i, e in enumerate(entries, 1):36 out.append(f"Thought: {e['thought']}")37 if e["action"]:38 out.append(f"Action: {e['action']}")39 if e["observation"] is not None:40 out.append(f"Observation: {e['observation']}")41 return "\n".join(out)The compaction keeps the first step (which usually contains the framing), summarises the middle, and keeps the last four steps verbatim. Recent steps stay whole because they are what the next decision depends on. Note that the summary explicitly excludes failed observations — repeating a list of errors into every subsequent prompt makes the model fixate on them.
The arithmetic on why this matters. A step of roughly 400 tokens, an unbounded trace, and a 20-step run: step 20's prompt carries about 8,000 tokens of trace. Summed across all 20 calls that is roughly 84,000 prompt tokens — about 25 cents at an illustrative 3 dollars per million input tokens, for one question. Capping the trace at 8,000 characters (~2,000 tokens) holds the total near 40,000 tokens, roughly halving it, and the agent behaves better because the relevant context is not buried.
A step budget does not bound cost. The trace grows with every step, so late steps are far more expensive than early ones - you need a token or character budget alongside it.
The LLM adapter
One interface, two implementations. The fake one is what lets you develop the loop at zero cost and write deterministic tests.
1class LLMClient:2 def complete(self, prompt: str, stop: list[str]) -> str:3 raise NotImplementedError45class ClaudeClient(LLMClient):6 def __init__(self, model: str | None = None):7 import os8 from anthropic import Anthropic9 self.client = Anthropic()10 # Read the model id from config; claude-sonnet-5 is a good default.11 self.model = model or os.environ.get("AGENT_MODEL", "claude-sonnet-5")12 self.input_tokens = self.output_tokens = 01314 def complete(self, prompt, stop):15 r = self.client.messages.create(16 model=self.model, max_tokens=512,17 stop_sequences=stop,18 messages=[{"role": "user", "content": prompt}])19 self.input_tokens += r.usage.input_tokens20 self.output_tokens += r.usage.output_tokens21 return "".join(b.text for b in r.content if b.type == "text")2223class ScriptedClient(LLMClient):24 """Deterministic stand-in. Replays turns in order."""25 def __init__(self, turns): self.turns, self.i = turns, 026 def complete(self, prompt, stop):27 turn = self.turns[min(self.i, len(self.turns) - 1)]28 self.i += 129 return turnstop_sequences=["Observation:"] is the fix for the model writing its own observations. Without it, generation runs straight past the action and invents the tool's reply, and the resulting trace looks flawless while being entirely fictional.
The agent
1SYSTEM = """Answer the question by cycling Thought / Action / Observation.23Available tools:4{tools}56Format - exactly one Thought and one Action, then stop:7 Thought: your reasoning8 Action: tool_name[arg=value, arg2=value]910When you have the answer:11 Thought: your reasoning12 Final Answer: the answer1314Rules:15- Never write an Observation yourself. It will be provided.16- Do not do arithmetic mentally; use a tool.17- If an observation is empty, zero, or surprising, say so in the18 next Thought and change approach rather than repeating the call.1920Question: {question}2122{trace}"""2324class Agent:25 def __init__(self, llm, registry, max_steps=10, max_chars=40000):26 self.llm, self.registry = llm, registry27 self.max_steps, self.max_chars = max_steps, max_chars2829 def run(self, question: str) -> dict:30 trace = Trace()31 chars_used = 03233 for step in range(1, self.max_steps + 1):34 prompt = SYSTEM.format(tools=self.registry.describe(),35 question=question,36 trace=trace.render())37 chars_used += len(prompt)38 if chars_used > self.max_chars:39 return self._stop(trace, "budget_exhausted",40 "Character budget exceeded.")4142 out = self.llm.complete(prompt, stop=["Observation:"])43 thought = self._thought(out)44 parsed = parse(out)4546 if parsed[0] == "final":47 trace.add(thought, action="Final Answer")48 return {"answer": parsed[1], "steps": step,49 "status": "ok", "trace": trace}5051 if parsed[0] == "error":52 trace.add(thought, action=None, observation=parsed[1])53 continue5455 _, name, args = parsed56 rendered = f"{name}[{', '.join(f'{k}={v}' for k, v in args.items())}]"5758 if trace.repeated(rendered):59 obs = (f"You already ran {rendered} and got the same "60 f"result. Repeating it cannot help. Use "61 f"different arguments, a different tool, or "62 f"give your Final Answer with what you know.")63 else:64 obs = self.registry.call(name, args)6566 trace.add(thought, action=rendered, observation=obs)6768 return self._stop(trace, "max_steps",69 f"Reached the {self.max_steps}-step limit.")7071 @staticmethod72 def _thought(text: str) -> str:73 m = re.search(r"Thought:\s*(.+?)(?=\n(?:Action:|Final Answer:)|$)",74 text, re.S)75 return m.group(1).strip() if m else text.strip()[:200]7677 @staticmethod78 def _stop(trace, status, msg):79 return {"answer": None, "status": status,80 "steps": len(trace.entries), "reason": msg,81 "trace": trace}Four things this loop does that the naive one did not:
- Parse errors continue the loop rather than crashing, and the error text tells the model the exact format to use.
trace.repeatedconverts a doom loop into an observation the model can respond to.- Two independent budgets — steps and characters. Steps alone do not bound cost, because a long trace makes late steps far more expensive than early ones.
- Failures return a structured result with a status, not a string saying "gave up". Callers can distinguish "no answer exists" from "we ran out of room".
A complete run
Question: "Order ORD-4471 was placed on 2026-07-12. Is it still within its refund window as of 2026-08-20?"
1llm = ScriptedClient([2 "Thought: I need the order's status before I can look up a policy.\n"3 "Action: get_order[order_id=ORD-4471]",45 "Thought: Status is delivered and it was placed 2026-07-12. "6 "Now the window for delivered orders.\n"7 "Action: get_refund_policy[status=delivered]",89 "Thought: 30 days. I need the elapsed days from 2026-07-12 to "10 "2026-08-20 - I must not compute that mentally.\n"11 "Action: days_between[start=2026-07-12, end=2026-08-20]",1213 "Thought: 39 days elapsed against a 30-day window, so it has "14 "expired by 9 days.\n"15 "Final Answer: No. ORD-4471 was delivered and placed on "16 "2026-07-12. Delivered orders are refundable for 30 days, but "17 "39 days have elapsed as of 2026-08-20 - the window closed 9 "18 "days ago, on 2026-08-11.",19])2021result = Agent(llm, registry).run(22 "Order ORD-4471 was placed on 2026-07-12. Is it still within "23 "its refund window as of 2026-08-20?")2425print(result["status"], result["steps"])26print(result["answer"])27print(result["trace"].render())ok 4Thought: I need the order's status before I can look up a policy.Action: get_order[order_id=ORD-4471]Observation: {"customer_id": 8842, "total_paise": 425000, "status": "delivered", "placed": "2026-07-12"}Thought: Status is delivered and it was placed 2026-07-12. Now the window for delivered orders.Action: get_refund_policy[status=delivered]Observation: delivered: refundable within 30 days.Thought: 30 days. I need the elapsed days from 2026-07-12 to 2026-08-20 - I must not compute that mentally.Action: days_between[start=2026-07-12, end=2026-08-20]Observation: 39Thought: 39 days elapsed against a 30-day window, so it has expired by 9 days.Action: Final AnswerVerify the dates by hand, because the agent's whole answer rests on them. July has 31 days, so 12 July to 31 July is 19 days; 31 July to 20 August is 20 days. 19 + 20 = 39. The window is 30 days, expiring 30 days after 12 July: 19 days takes you to 31 July, leaving 11 more, which lands on 11 August. From 11 to 20 August is 9 days overdue. The agent's answer checks out — and it checks out because days_between computed it rather than the model.
Watching it recover
Thought: Let me get the order.Action: get_order[id=ORD-4471]Observation: get_order has no parameter(s) ['id']. Valid parameters: ['order_id'].Thought: The parameter is order_id, not id.Action: get_order[order_id=ORD-447]Observation: No order 'ORD-447'. IDs look like ORD-4471. Known orders for this session: ['ORD-4471'].Thought: I truncated the ID.Action: get_order[order_id=ORD-4471]Observation: {"customer_id": 8842, ... "status": "delivered" ...}Two mistakes, two recoveries, zero crashes. Under the naive implementation the first line raises a TypeError and the run is over. The only difference is that each error message names the correction.
Where people get this wrong
Parsing with split(). It works on the happy path and fails on multi-line arguments, commas inside values, trailing whitespace, and prose. Use a regex with anchors and a bracket-aware splitter.
No stop sequence. The single highest-severity bug in this whole lesson, because it produces a complete, fluent, confident trace in which no tool ever ran.
Bounding steps but not tokens. Cost grows superlinearly with step count as the trace accumulates. Ten steps with an unbounded trace can cost more than thirty with a compacted one.
Letting tools raise. One unhandled exception discards every step of progress before it. Catch inside Registry.call and return the exception as text.
Testing only against the real API. Slow, expensive, and non-deterministic, so you cannot write a test that asserts what the loop does when a tool returns an error. A scripted client makes every one of those cases a fast unit test.
Throwing away the trace on success. The trace is the only artefact that explains an answer. Return it, store it, and you can debug a complaint from three weeks ago.
What this means when you build one
Build the ScriptedClient before the real one. Write out the exact turn sequence you expect for two or three tasks, including one that goes wrong, and get the loop working against those. You will find the parsing, budgeting and repeat-detection bugs in minutes rather than discovering them in production, and you will not spend anything doing it.
Then apply one rule everywhere: nothing raises except a genuine programming error. Bad arguments, unknown tools, exceptions inside tools, unparseable model output — all of them become strings that go back into the trace. That is what makes the loop resilient, and it is why the finished agent has no try/except in it at all: every failure was already converted at its source.
Set two budgets, not one, and log the full trace of every run with its status. When you later want to know why an agent gave a strange answer, the trace tells you which observation misled it, and the fix is almost always in the tool that produced that observation rather than in the prompt.