Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Build an AI agent that uses tools like calculator and search.


The loop behind the iPhone price questionquestion intranscriptmodelreturns one actionrun tool,catch errorsappendobservationsearch, thencalculatorastwhitelist, no evalStep 3 returns final_answer: 3 phones at 10 percent off cost Rs 2,15,730.
Every error comes back as an observation, so one bad call costs a step instead of the whole run.

What you need to know

The pattern is often called ReAct (reason + act): the model thinks, acts with a tool, observes the result, and thinks again. Three pieces make it work:

  • A tool registry — a dictionary from tool name to Python function, so the model's choice is a lookup, not code execution.
  • A structured action format — here, one JSON object per reply. In production, the provider's native tool calling does this for you (the next question).
  • The transcript — the list of previous actions and observations, sent back on every step. It is the agent's only memory.

Never eval() model output. A calculator built on eval will run __import__('os').system(...) if the model (or a prompt injection) writes it. Parse the expression into a syntax tree with ast and allow only numbers and arithmetic operators.

Python
import ast, json, operatorfrom collections.abc import Callable_OPS = {ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,        ast.Div: operator.truediv, ast.Pow: operator.pow, ast.USub: operator.neg}def calculator(expression: str) -> str:    """Safely evaluate + - * / ** on numbers. Never eval() model output."""    def ev(node: ast.AST) -> float:        if isinstance(node, ast.Constant) and type(node.value) in (int, float):            return node.value        if isinstance(node, ast.BinOp) and type(node.op) in _OPS:            left, right = ev(node.left), ev(node.right)            if isinstance(node.op, ast.Pow) and abs(right) > 100:                raise ValueError("exponent too large")            return _OPS[type(node.op)](left, right)        if isinstance(node, ast.UnaryOp) and type(node.op) in _OPS:            return _OPS[type(node.op)](ev(node.operand))        raise ValueError(f"unsupported syntax: {type(node).__name__}")    return str(ev(ast.parse(expression, mode="eval").body))INDEX = {"iphone 16 price": "The iPhone 16 costs Rs 79,900 in India."}def search(query: str) -> str:    return INDEX.get(query.lower().strip(), "no results")TOOLS: dict[str, Callable[..., str]] = {"calculator": calculator, "search": search}

The loop:

Python
SYSTEM = (    "You are a tool-using agent. Reply with exactly one JSON object and nothing else:\n"    '{"tool": "calculator", "args": {"expression": "2 + 2"}}\n'    '{"tool": "search", "args": {"query": "..."}}\n'    '{"tool": "final_answer", "answer": "..."}\n')def run_agent(question: str, llm_fn: Callable[[str], str], max_steps: int = 6) -> str:    """Ask the model for one action at a time until it answers or runs out of steps."""    transcript = [f"Question: {question}"]    for _ in range(max_steps):        raw = llm_fn(SYSTEM + "\n".join(transcript) + "\nNext action:")        try:            action = json.loads(raw)        except json.JSONDecodeError:            action = None        if not isinstance(action, dict):            transcript.append("Observation: reply was not one JSON object. Try again.")            continue        if action.get("tool") == "final_answer":            return str(action.get("answer", ""))        fn = TOOLS.get(action.get("tool"))        try:            result = fn(**action.get("args", {})) if fn else f"unknown tool {action.get('tool')!r}"        except Exception as exc:                       # includes bad arguments (TypeError)            result = f"error: {type(exc).__name__}: {exc}"        transcript.append(f"Action: {json.dumps(action)}\nObservation: {result}")    return "Stopped: step limit reached."

The tricky parts:

  • type(node.value) in (int, float), not isinstance: True is an int subclass, and a calculator that accepts True + 1 is sloppy.
  • The exponent cap. 9 ** 9 ** 9 is a valid arithmetic expression that would try to build a number with hundreds of millions of digits and hang the worker. Capping the exponent stops it.
  • isinstance(action, dict) — the model can return valid JSON that is not an object, such as a list; .get on it would crash the loop.
  • Errors become observations, never exceptions that end the run. The model reads "ZeroDivisionError" and tries something else.

Complexity: each step is one model call plus the tool's cost. The prompt at step s contains all s earlier steps, so the total tokens over S steps are about 1 + 2 + … + S, which is O(S²). That quadratic growth is why max_steps is a hard number, not a suggestion.

A real-life example

A scripted fake model plays the part of the LLM, so the run is repeatable:

Python
script = iter([    '{"tool": "search", "args": {"query": "iPhone 16 price"}}',    '{"tool": "calculator", "args": {"expression": "79900 * 3 * 0.9"}}',    '{"tool": "final_answer", "answer": "Three phones with 10% off cost Rs 2,15,730."}',])print(run_agent("What do 3 iPhone 16s cost with a 10% discount?", lambda _: next(script)))# Three phones with 10% off cost Rs 2,15,730.
stepmodel's actionobservation appended
1search("iPhone 16 price")The iPhone 16 costs Rs 79,900 in India.
2calculator("79900 * 3 * 0.9")215730.0
3final_answerloop returns the answer

At step 2 the model's prompt already contains the search result, which is how it knew the price. At step 3 it contains both observations.

The same loop, with real tools, is what an e-commerce shopping assistant runs when it looks up a price, applies a coupon and quotes the total.

Follow-up questions to expect

  • "Why not just use eval with a restricted namespace?" — Restricting eval is famously leaky; attribute tricks reach builtins anyway. A whitelist over the syntax tree is small and closed by construction.
  • "How would you do this with a real provider?" — Use native tool calling: pass JSON schemas as tools, and the API returns structured tool_use blocks, so you stop parsing free text.
  • "The model calls the same search five times — how do you stop that?" — Detect repeated identical actions and stop or tell the model it is repeating itself; that is the stop-conditions question.