Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Build an AI agent that uses tools like calculator and search.
What you need to know
The pattern is often called ReAct (reason + act): the model thinks, acts with a tool, observes the result, and thinks again. Three pieces make it work:
- A tool registry — a dictionary from tool name to Python function, so the model's choice is a lookup, not code execution.
- A structured action format — here, one JSON object per reply. In production, the provider's native tool calling does this for you (the next question).
- The transcript — the list of previous actions and observations, sent back on every step. It is the agent's only memory.
Never eval() model output. A calculator built on eval will run __import__('os').system(...) if the model (or a prompt injection) writes it. Parse the expression into a syntax tree with ast and allow only numbers and arithmetic operators.
1import ast, json, operator2from collections.abc import Callable34_OPS = {ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,5 ast.Div: operator.truediv, ast.Pow: operator.pow, ast.USub: operator.neg}67def calculator(expression: str) -> str:8 """Safely evaluate + - * / ** on numbers. Never eval() model output."""9 def ev(node: ast.AST) -> float:10 if isinstance(node, ast.Constant) and type(node.value) in (int, float):11 return node.value12 if isinstance(node, ast.BinOp) and type(node.op) in _OPS:13 left, right = ev(node.left), ev(node.right)14 if isinstance(node.op, ast.Pow) and abs(right) > 100:15 raise ValueError("exponent too large")16 return _OPS[type(node.op)](left, right)17 if isinstance(node, ast.UnaryOp) and type(node.op) in _OPS:18 return _OPS[type(node.op)](ev(node.operand))19 raise ValueError(f"unsupported syntax: {type(node).__name__}")20 return str(ev(ast.parse(expression, mode="eval").body))2122INDEX = {"iphone 16 price": "The iPhone 16 costs Rs 79,900 in India."}2324def search(query: str) -> str:25 return INDEX.get(query.lower().strip(), "no results")2627TOOLS: dict[str, Callable[..., str]] = {"calculator": calculator, "search": search}The loop:
1SYSTEM = (2 "You are a tool-using agent. Reply with exactly one JSON object and nothing else:\n"3 '{"tool": "calculator", "args": {"expression": "2 + 2"}}\n'4 '{"tool": "search", "args": {"query": "..."}}\n'5 '{"tool": "final_answer", "answer": "..."}\n'6)78def run_agent(question: str, llm_fn: Callable[[str], str], max_steps: int = 6) -> str:9 """Ask the model for one action at a time until it answers or runs out of steps."""10 transcript = [f"Question: {question}"]11 for _ in range(max_steps):12 raw = llm_fn(SYSTEM + "\n".join(transcript) + "\nNext action:")13 try:14 action = json.loads(raw)15 except json.JSONDecodeError:16 action = None17 if not isinstance(action, dict):18 transcript.append("Observation: reply was not one JSON object. Try again.")19 continue20 if action.get("tool") == "final_answer":21 return str(action.get("answer", ""))22 fn = TOOLS.get(action.get("tool"))23 try:24 result = fn(**action.get("args", {})) if fn else f"unknown tool {action.get('tool')!r}"25 except Exception as exc: # includes bad arguments (TypeError)26 result = f"error: {type(exc).__name__}: {exc}"27 transcript.append(f"Action: {json.dumps(action)}\nObservation: {result}")28 return "Stopped: step limit reached."The tricky parts:
type(node.value) in (int, float), notisinstance:Trueis anintsubclass, and a calculator that acceptsTrue + 1is sloppy.- The exponent cap.
9 ** 9 ** 9is a valid arithmetic expression that would try to build a number with hundreds of millions of digits and hang the worker. Capping the exponent stops it. isinstance(action, dict)— the model can return valid JSON that is not an object, such as a list;.geton it would crash the loop.- Errors become observations, never exceptions that end the run. The model reads "ZeroDivisionError" and tries something else.
Complexity: each step is one model call plus the tool's cost. The prompt at step s contains all s earlier steps, so the total tokens over S steps are about 1 + 2 + … + S, which is O(S²). That quadratic growth is why max_steps is a hard number, not a suggestion.
A real-life example
A scripted fake model plays the part of the LLM, so the run is repeatable:
1script = iter([2 '{"tool": "search", "args": {"query": "iPhone 16 price"}}',3 '{"tool": "calculator", "args": {"expression": "79900 * 3 * 0.9"}}',4 '{"tool": "final_answer", "answer": "Three phones with 10% off cost Rs 2,15,730."}',5])6print(run_agent("What do 3 iPhone 16s cost with a 10% discount?", lambda _: next(script)))7# Three phones with 10% off cost Rs 2,15,730.| step | model's action | observation appended |
|---|---|---|
| 1 | search("iPhone 16 price") | The iPhone 16 costs Rs 79,900 in India. |
| 2 | calculator("79900 * 3 * 0.9") | 215730.0 |
| 3 | final_answer | loop returns the answer |
At step 2 the model's prompt already contains the search result, which is how it knew the price. At step 3 it contains both observations.
The same loop, with real tools, is what an e-commerce shopping assistant runs when it looks up a price, applies a coupon and quotes the total.
Follow-up questions to expect
- "Why not just use
evalwith a restricted namespace?" — Restrictingevalis famously leaky; attribute tricks reach builtins anyway. A whitelist over the syntax tree is small and closed by construction. - "How would you do this with a real provider?" — Use native tool calling: pass JSON schemas as
tools, and the API returns structuredtool_useblocks, so you stop parsing free text. - "The model calls the same search five times — how do you stop that?" — Detect repeated identical actions and stop or tell the model it is repeating itself; that is the stop-conditions question.