AI Agent Fundamentals

Implementing a Simple ReAct Agent (Python)


Here is the agent almost everyone writes first. Thirty-one lines, and it works in the demo.

Python
def agent(question):    trace = ""    for _ in range(10):        out = llm(f"Answer using tools.\n{question}\n{trace}")        if "Final Answer:" in out:            return out.split("Final Answer:")[1]        action = out.split("Action:")[1].split("\n")[0]        name, arg = action.split("[")        result = TOOLS[name](arg.rstrip("]"))        trace += f"{out}\nObservation: {result}\n"    return "gave up"

It survived one question. Then, over a week, it produced these:

What happenedLine that did it
IndexError: list index out of rangeout.split("Action:")[1] — the model wrote prose instead
KeyError: 'search_web 'A trailing space in the tool name
The model wrote its own Observation: and answered from fictionNo stop sequence on the LLM call
Same tool called eleven times with identical argumentsNo repeat detection
One run cost 3.40 dollarsNo token budget; the trace grew unbounded
A tool raised and the whole run died at step 7No try/except around execution
Two arguments needed; the parser only handles onearg.rstrip("]") as an argument parser

Every one of these is an engineering problem, not a prompting problem. This lesson builds the version that survives — around 150 lines, with each part justified by one of the failures above.

The five parts of the thirty-one-line agentTool registryParser forthoughtand actionMemory ofthe transcriptModeladapter,swappableThe loop,with a step capThe adapter is what lets a scripted double stand in for the model in tests.
The parser is where the demo agent breaks first: any output it cannot read becomes a silent dead end.

The five parts

Text
┌──────────────────────────────────────────────────┐│ Agent                                            ││   owns the loop, the budgets, the stop rules     ││                                                  ││  ┌────────────┐  ┌────────────┐  ┌────────────┐  ││  │ LLMClient  │  │  Registry  │  │   Trace    │  ││  │ one call,  │  │ tools +    │  │ memory +   │  ││  │ swappable  │  │ schemas +  │  │ compaction │  ││  │ for tests  │  │ validation │  │            │  ││  └────────────┘  └────────────┘  └────────────┘  ││         ┌────────────────────────┐               ││         │        Parser          │               ││         │ text -> (tool, args)   │               ││         │ never raises           │               ││         └────────────────────────┘               │└──────────────────────────────────────────────────┘
Failure from the table aboveComponent that prevents itMechanism
IndexError on prose outputParserReturns an error tuple; the loop continues
KeyError: 'search_web 'ParserAnchored regex plus .strip() on the name
Model invents its own observationsLLMClientstop_sequences=["Observation:"]
Same call eleven timesTracerepeated() returns a corrective observation
One run costs 3.40 dollarsTrace + AgentCompaction plus an independent character budget
Tool exception kills the runRegistrycall() catches and returns the error as text
Two-argument call mis-parsedParserBracket-aware split_top_level and key=value pairs

The separation is not architectural taste. Each boundary buys something specific: a swappable LLMClient makes the agent testable without spending money, a Registry puts validation in one place instead of scattered through the loop, a Trace object makes memory a thing you can measure and trim, and a Parser that never raises turns malformed output into a recoverable observation.

Tools and the registry

Tools are plain functions. The registry derives the schema from the signature and docstring so the description the model reads cannot drift from the code that runs.

Python
import inspect, json, refrom typing import Callable, Anyclass Registry:    def __init__(self):        self.tools: dict[str, dict] = {}    def register(self, fn: Callable) -> Callable:        sig = inspect.signature(fn)        params = {}        for name, p in sig.parameters.items():            ann = p.annotation            params[name] = {                "type": {str: "string", int: "integer",                         float: "number", bool: "boolean"                         }.get(ann, "string"),                "required": p.default is inspect.Parameter.empty,            }        self.tools[fn.__name__] = {            "fn": fn,            "doc": inspect.getdoc(fn) or "",            "params": params,        }        return fn    def describe(self) -> str:        lines = []        for name, spec in self.tools.items():            args = ", ".join(                f"{k}: {v['type']}" + ("" if v["required"] else "?")                for k, v in spec["params"].items())            lines.append(f"  {name}({args})\n      {spec['doc']}")        return "\n".join(lines)    def validate(self, name: str, args: dict) -> str | None:        """Return an error message, or None if the call is valid."""        if name not in self.tools:            return (f"Unknown tool '{name}'. Available: "                    f"{', '.join(self.tools)}.")        spec = self.tools[name]["params"]        missing = [k for k, v in spec.items()                   if v["required"] and k not in args]        if missing:            return (f"{name} is missing required argument(s) "                    f"{missing}. Expected: {list(spec)}.")        unknown = [k for k in args if k not in spec]        if unknown:            return (f"{name} has no parameter(s) {unknown}. "                    f"Valid parameters: {list(spec)}.")        return None    def call(self, name: str, args: dict) -> str:        problem = self.validate(name, args)        if problem:            return problem        try:            return str(self.tools[name]["fn"](**args))        except Exception as e:            return (f"{name} raised {type(e).__name__}: {e}. "                    f"Check the arguments and try a different approach.")

call returns a string on every path — success, invalid arguments, unknown tool, exception. Nothing escapes as an exception. That single property removes the "tool raised and the run died" failure permanently, and it is why the loop below has no error handling of its own.

Make one rule and hold it everywhere: nothing raises. Bad arguments, unknown tools, exceptions inside tools and unparseable model output all become strings that go back into the trace.

Some tools:

Python
registry = Registry()@registry.registerdef get_order(order_id: str) -> str:    """Look up one order by ID. Format: ORD- followed by 4 digits."""    db = {"ORD-4471": {"customer_id": 8842, "total_paise": 425000,                       "status": "delivered", "placed": "2026-07-12"}}    if order_id not in db:        return (f"No order '{order_id}'. IDs look like ORD-4471. "                f"Known orders for this session: {list(db)}.")    return json.dumps(db[order_id])@registry.registerdef get_refund_policy(status: str) -> str:    """Refund window in days for an order status: delivered,    shipped, placed, or cancelled."""    policy = {"delivered": 30, "shipped": 14,              "placed": 90, "cancelled": 0}    if status not in policy:        return f"Unknown status '{status}'. Valid: {list(policy)}."    return f"{status}: refundable within {policy[status]} days."@registry.registerdef days_between(start: str, end: str) -> str:    """Whole days between two ISO dates (YYYY-MM-DD)."""    from datetime import date    try:        d0, d1 = date.fromisoformat(start), date.fromisoformat(end)    except ValueError:        return f"Dates must be YYYY-MM-DD. Got '{start}', '{end}'."    return str((d1 - d0).days)

Every error return names the fix. get_order even lists the valid IDs. This is not politeness — it is the difference between an agent that recovers on the next turn and one that guesses again.

The parser

Turning model text into a call is where the naive version broke three separate ways. The parser must handle every malformed shape and return an error rather than raising.

Python
ACTION_RE = re.compile(    r"^\s*Action:\s*([A-Za-z_]\w*)\s*\[(.*)\]\s*$",    re.MULTILINE | re.DOTALL)FINAL_RE = re.compile(r"^\s*Final Answer:\s*(.+)", re.MULTILINE | re.DOTALL)def parse(text: str):    """-> ('final', answer) | ('action', name, args) | ('error', msg)"""    final = FINAL_RE.search(text)    if final:        return ("final", final.group(1).strip())    m = ACTION_RE.search(text)    if not m:        return ("error",                "No Action found. Reply with exactly one line of the "                "form: Action: tool_name[arg=value, arg2=value] "                "or: Final Answer: your answer")    name = m.group(1).strip()    raw = m.group(2).strip()    if not raw:        return ("action", name, {})    args = {}    for part in split_top_level(raw):        if "=" not in part:            return ("error",                    f"Argument '{part.strip()}' is not in key=value "                    f"form. Example: {name}[order_id=ORD-4471]")        k, v = part.split("=", 1)        args[k.strip()] = coerce(v.strip().strip("'\""))    return ("action", name, args)def split_top_level(s: str) -> list[str]:    """Split on commas not inside quotes or brackets."""    out, depth, quote, buf = [], 0, None, []    for ch in s:        if quote:            if ch == quote:                quote = None        elif ch in "'\"":            quote = ch        elif ch in "([{":            depth += 1        elif ch in ")]}":            depth -= 1        elif ch == "," and depth == 0:            out.append("".join(buf)); buf = []; continue        buf.append(ch)    if buf:        out.append("".join(buf))    return outdef coerce(v: str):    if v.lower() in ("true", "false"):        return v.lower() == "true"    try:        return int(v)    except ValueError:        pass    try:        return float(v)    except ValueError:        return v

split_top_level exists because a naive raw.split(",") destroys search[query=orders, refunds]. Twenty lines that stop a whole family of silent argument corruption.

.strip() on the tool name kills the KeyError: 'search_web ' failure. It looks trivial and it accounted for about one failed run in twenty.

Memory

The trace is the agent's working memory, and it is the thing that made one run cost 3.40 dollars. Make it an object that knows its own size and can shrink.

Python
class Trace:    def __init__(self, keep_recent: int = 4, max_chars: int = 8000):        self.entries: list[dict] = []        self.keep_recent = keep_recent        self.max_chars = max_chars    def add(self, thought, action=None, observation=None):        self.entries.append({"thought": thought, "action": action,                             "observation": observation})    def repeated(self, action: str) -> bool:        return any(e["action"] == action for e in self.entries)    def render(self) -> str:        text = self._render(self.entries)        if len(text) <= self.max_chars:            return text        head = self.entries[:1]        recent = self.entries[-self.keep_recent:]        middle = self.entries[1:-self.keep_recent]        summary = ("[{} earlier steps compressed. Tools used: {}. "                   "Facts established: {}]").format(            len(middle),            ", ".join(sorted({e["action"].split("[")[0]                              for e in middle if e["action"]})),            "; ".join(e["observation"][:80] for e in middle                      if e["observation"] and "rror" not in                      e["observation"])[:400])        return self._render(head) + "\n" + summary + "\n" + \               self._render(recent)    @staticmethod    def _render(entries) -> str:        out = []        for i, e in enumerate(entries, 1):            out.append(f"Thought: {e['thought']}")            if e["action"]:                out.append(f"Action: {e['action']}")            if e["observation"] is not None:                out.append(f"Observation: {e['observation']}")        return "\n".join(out)

The compaction keeps the first step (which usually contains the framing), summarises the middle, and keeps the last four steps verbatim. Recent steps stay whole because they are what the next decision depends on. Note that the summary explicitly excludes failed observations — repeating a list of errors into every subsequent prompt makes the model fixate on them.

The arithmetic on why this matters. A step of roughly 400 tokens, an unbounded trace, and a 20-step run: step 20's prompt carries about 8,000 tokens of trace. Summed across all 20 calls that is roughly 84,000 prompt tokens — about 25 cents at an illustrative 3 dollars per million input tokens, for one question. Capping the trace at 8,000 characters (~2,000 tokens) holds the total near 40,000 tokens, roughly halving it, and the agent behaves better because the relevant context is not buried.

A step budget does not bound cost. The trace grows with every step, so late steps are far more expensive than early ones - you need a token or character budget alongside it.

The LLM adapter

One interface, two implementations. The fake one is what lets you develop the loop at zero cost and write deterministic tests.

Python
class LLMClient:    def complete(self, prompt: str, stop: list[str]) -> str:        raise NotImplementedErrorclass ClaudeClient(LLMClient):    def __init__(self, model: str | None = None):        import os        from anthropic import Anthropic        self.client = Anthropic()        # Read the model id from config; claude-sonnet-5 is a good default.        self.model = model or os.environ.get("AGENT_MODEL", "claude-sonnet-5")        self.input_tokens = self.output_tokens = 0    def complete(self, prompt, stop):        r = self.client.messages.create(            model=self.model, max_tokens=512,            stop_sequences=stop,            messages=[{"role": "user", "content": prompt}])        self.input_tokens += r.usage.input_tokens        self.output_tokens += r.usage.output_tokens        return "".join(b.text for b in r.content if b.type == "text")class ScriptedClient(LLMClient):    """Deterministic stand-in. Replays turns in order."""    def __init__(self, turns): self.turns, self.i = turns, 0    def complete(self, prompt, stop):        turn = self.turns[min(self.i, len(self.turns) - 1)]        self.i += 1        return turn

stop_sequences=["Observation:"] is the fix for the model writing its own observations. Without it, generation runs straight past the action and invents the tool's reply, and the resulting trace looks flawless while being entirely fictional.

The agent

Python
SYSTEM = """Answer the question by cycling Thought / Action / Observation.Available tools:{tools}Format - exactly one Thought and one Action, then stop:  Thought: your reasoning  Action: tool_name[arg=value, arg2=value]When you have the answer:  Thought: your reasoning  Final Answer: the answerRules:- Never write an Observation yourself. It will be provided.- Do not do arithmetic mentally; use a tool.- If an observation is empty, zero, or surprising, say so in the  next Thought and change approach rather than repeating the call.Question: {question}{trace}"""class Agent:    def __init__(self, llm, registry, max_steps=10, max_chars=40000):        self.llm, self.registry = llm, registry        self.max_steps, self.max_chars = max_steps, max_chars    def run(self, question: str) -> dict:        trace = Trace()        chars_used = 0        for step in range(1, self.max_steps + 1):            prompt = SYSTEM.format(tools=self.registry.describe(),                                   question=question,                                   trace=trace.render())            chars_used += len(prompt)            if chars_used > self.max_chars:                return self._stop(trace, "budget_exhausted",                                  "Character budget exceeded.")            out = self.llm.complete(prompt, stop=["Observation:"])            thought = self._thought(out)            parsed = parse(out)            if parsed[0] == "final":                trace.add(thought, action="Final Answer")                return {"answer": parsed[1], "steps": step,                        "status": "ok", "trace": trace}            if parsed[0] == "error":                trace.add(thought, action=None, observation=parsed[1])                continue            _, name, args = parsed            rendered = f"{name}[{', '.join(f'{k}={v}' for k, v in args.items())}]"            if trace.repeated(rendered):                obs = (f"You already ran {rendered} and got the same "                       f"result. Repeating it cannot help. Use "                       f"different arguments, a different tool, or "                       f"give your Final Answer with what you know.")            else:                obs = self.registry.call(name, args)            trace.add(thought, action=rendered, observation=obs)        return self._stop(trace, "max_steps",                          f"Reached the {self.max_steps}-step limit.")    @staticmethod    def _thought(text: str) -> str:        m = re.search(r"Thought:\s*(.+?)(?=\n(?:Action:|Final Answer:)|$)",                      text, re.S)        return m.group(1).strip() if m else text.strip()[:200]    @staticmethod    def _stop(trace, status, msg):        return {"answer": None, "status": status,                "steps": len(trace.entries), "reason": msg,                "trace": trace}

Four things this loop does that the naive one did not:

  • Parse errors continue the loop rather than crashing, and the error text tells the model the exact format to use.
  • trace.repeated converts a doom loop into an observation the model can respond to.
  • Two independent budgets — steps and characters. Steps alone do not bound cost, because a long trace makes late steps far more expensive than early ones.
  • Failures return a structured result with a status, not a string saying "gave up". Callers can distinguish "no answer exists" from "we ran out of room".

A complete run

Question: "Order ORD-4471 was placed on 2026-07-12. Is it still within its refund window as of 2026-08-20?"

Python
llm = ScriptedClient([  "Thought: I need the order's status before I can look up a policy.\n"  "Action: get_order[order_id=ORD-4471]",  "Thought: Status is delivered and it was placed 2026-07-12. "  "Now the window for delivered orders.\n"  "Action: get_refund_policy[status=delivered]",  "Thought: 30 days. I need the elapsed days from 2026-07-12 to "  "2026-08-20 - I must not compute that mentally.\n"  "Action: days_between[start=2026-07-12, end=2026-08-20]",  "Thought: 39 days elapsed against a 30-day window, so it has "  "expired by 9 days.\n"  "Final Answer: No. ORD-4471 was delivered and placed on "  "2026-07-12. Delivered orders are refundable for 30 days, but "  "39 days have elapsed as of 2026-08-20 - the window closed 9 "  "days ago, on 2026-08-11.",])result = Agent(llm, registry).run(    "Order ORD-4471 was placed on 2026-07-12. Is it still within "    "its refund window as of 2026-08-20?")print(result["status"], result["steps"])print(result["answer"])print(result["trace"].render())
Text
ok 4Thought: I need the order's status before I can look up a policy.Action: get_order[order_id=ORD-4471]Observation: {"customer_id": 8842, "total_paise": 425000,              "status": "delivered", "placed": "2026-07-12"}Thought: Status is delivered and it was placed 2026-07-12. Now the         window for delivered orders.Action: get_refund_policy[status=delivered]Observation: delivered: refundable within 30 days.Thought: 30 days. I need the elapsed days from 2026-07-12 to         2026-08-20 - I must not compute that mentally.Action: days_between[start=2026-07-12, end=2026-08-20]Observation: 39Thought: 39 days elapsed against a 30-day window, so it has expired         by 9 days.Action: Final Answer

Verify the dates by hand, because the agent's whole answer rests on them. July has 31 days, so 12 July to 31 July is 19 days; 31 July to 20 August is 20 days. 19 + 20 = 39. The window is 30 days, expiring 30 days after 12 July: 19 days takes you to 31 July, leaving 11 more, which lands on 11 August. From 11 to 20 August is 9 days overdue. The agent's answer checks out — and it checks out because days_between computed it rather than the model.

Watching it recover

Text
Thought: Let me get the order.Action: get_order[id=ORD-4471]Observation: get_order has no parameter(s) ['id']. Valid             parameters: ['order_id'].Thought: The parameter is order_id, not id.Action: get_order[order_id=ORD-447]Observation: No order 'ORD-447'. IDs look like ORD-4471. Known             orders for this session: ['ORD-4471'].Thought: I truncated the ID.Action: get_order[order_id=ORD-4471]Observation: {"customer_id": 8842, ... "status": "delivered" ...}

Two mistakes, two recoveries, zero crashes. Under the naive implementation the first line raises a TypeError and the run is over. The only difference is that each error message names the correction.

Where people get this wrong

Parsing with split(). It works on the happy path and fails on multi-line arguments, commas inside values, trailing whitespace, and prose. Use a regex with anchors and a bracket-aware splitter.

No stop sequence. The single highest-severity bug in this whole lesson, because it produces a complete, fluent, confident trace in which no tool ever ran.

Bounding steps but not tokens. Cost grows superlinearly with step count as the trace accumulates. Ten steps with an unbounded trace can cost more than thirty with a compacted one.

Letting tools raise. One unhandled exception discards every step of progress before it. Catch inside Registry.call and return the exception as text.

Testing only against the real API. Slow, expensive, and non-deterministic, so you cannot write a test that asserts what the loop does when a tool returns an error. A scripted client makes every one of those cases a fast unit test.

Throwing away the trace on success. The trace is the only artefact that explains an answer. Return it, store it, and you can debug a complaint from three weeks ago.

What this means when you build one

Build the ScriptedClient before the real one. Write out the exact turn sequence you expect for two or three tasks, including one that goes wrong, and get the loop working against those. You will find the parsing, budgeting and repeat-detection bugs in minutes rather than discovering them in production, and you will not spend anything doing it.

Then apply one rule everywhere: nothing raises except a genuine programming error. Bad arguments, unknown tools, exceptions inside tools, unparseable model output — all of them become strings that go back into the trace. That is what makes the loop resilient, and it is why the finished agent has no try/except in it at all: every failure was already converted at its source.

Set two budgets, not one, and log the full trace of every run with its status. When you later want to know why an agent gave a strange answer, the trace tells you which observation misled it, and the fix is almost always in the tool that produced that observation rather than in the prompt.