Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Implement a function-calling handler for LLM APIs.


What you need to know

Function calling (also called tool use) means the model does not run anything. It returns a structured request — "call get_weather with {"city": "Pune"}" — and your code decides whether and how to run it. Four parts:

  • Schema. A name, a description the model reads to decide when to call it, and a JSON Schema for the arguments. additionalProperties: false plus a required list, with strict: true on the tool, makes the API guarantee the arguments match.
  • Dispatch. Name to function lookup, with errors caught.
  • Result protocol. For the Anthropic Messages API: the assistant turn contains tool_use blocks, each with an id. You reply with a user turn of tool_result blocks carrying the same tool_use_id.
  • Parallel calls. One reply can contain several tool_use blocks. All their results go back in one user message.

inspect.signature gives you each parameter's name, type annotation and default, which is everything a schema needs.

Python
import inspect, jsonfrom collections.abc import Callablefrom typing import AnyREGISTRY: dict[str, dict[str, Any]] = {}_JSON_TYPES = {str: "string", int: "integer", float: "number",               bool: "boolean", list: "array", dict: "object"}def tool(fn: Callable) -> Callable:    """Register fn as a tool; its schema comes from its signature and docstring."""    props, required = {}, []    for name, p in inspect.signature(fn).parameters.items():        props[name] = {"type": _JSON_TYPES.get(p.annotation, "string")}        if p.default is inspect.Parameter.empty:            required.append(name)    REGISTRY[fn.__name__] = {"fn": fn, "spec": {        "name": fn.__name__,        "description": inspect.getdoc(fn) or "",        "input_schema": {"type": "object", "properties": props,                         "required": required, "additionalProperties": False},        "strict": True,    }}    return fndef dispatch(name: str, args: dict) -> tuple[bool, Any]:    """Run a registered tool. Returns (ok, result_or_error_message)."""    entry = REGISTRY.get(name)    if entry is None:        return False, f"unknown tool: {name}"    try:        return True, entry["fn"](**args)    except TypeError as exc:        return False, f"bad arguments: {exc}"    except Exception as exc:        return False, f"{type(exc).__name__}: {exc}"@tooldef get_weather(city: str, unit: str = "C") -> str:    """Return the current weather for a city."""    return f"{city}: 31{unit}, humid"

The loop, with the Anthropic Python SDK:

Python
from anthropic import Anthropicclient = Anthropic()def chat(user_msg: str, max_turns: int = 8) -> str:    """Run the tool loop until the model stops asking for tools."""    messages = [{"role": "user", "content": user_msg}]    for _ in range(max_turns):        resp = client.messages.create(            model="claude-opus-5", max_tokens=16000,            tools=[t["spec"] for t in REGISTRY.values()], messages=messages,        )        messages.append({"role": "assistant", "content": resp.content})        if resp.stop_reason != "tool_use":            return "".join(b.text for b in resp.content if b.type == "text")        results = []        for block in resp.content:            if block.type == "tool_use":                ok, out = dispatch(block.name, block.input)                results.append({"type": "tool_result", "tool_use_id": block.id,                                "content": json.dumps(out, default=str), "is_error": not ok})        messages.append({"role": "user", "content": results})    # all results, one message    raise RuntimeError("tool loop did not finish within max_turns")

The tricky parts:

  • Append resp.content, not just its text. The assistant turn must contain the tool_use blocks (and any thinking blocks) exactly as returned, or the API cannot match your tool_result ids.
  • Every tool_use gets a tool_result, even on failure — a missing one makes the next request fail. is_error: True tells the model the call failed.
  • One user message for all results. Splitting parallel results across several messages works, but it teaches the model to stop making parallel calls.
  • json.dumps(out, default=str) so a tool returning a datetime or Decimal does not crash serialisation.

Complexity: building schemas is O(parameters) once at import. Dispatch is an O(1) dict lookup plus the tool's own cost. The loop is O(turns) API calls, capped by max_turns.

A real-life example

Python
print(json.dumps(REGISTRY["get_weather"]["spec"]["input_schema"]))# (one line, wrapped here)# {"type": "object", "properties": {"city": {"type": "string"}, "unit": {"type": "string"}},#  "required": ["city"], "additionalProperties": false}print(dispatch("get_weather", {"city": "Pune"}))# (True, 'Pune: 31C, humid')print(dispatch("get_weather", {}))# (False, "bad arguments: get_weather() missing 1 required positional argument: 'city'")print(dispatch("book_flight", {"to": "Goa"}))# (False, 'unknown tool: book_flight')

Trace of one full turn where the user asks "Weather in Pune and Delhi?":

messagecontent
userWeather in Pune and Delhi?
assistanttool_use id tu_1 get_weather(Pune), tool_use id tu_2 get_weather(Delhi) — stop_reason is tool_use
usertool_result for tu_1: "Pune: 31C, humid"; tool_result for tu_2: "Delhi: 31C, humid"
assistant"Both cities are at 31C and humid." — stop_reason is end_turn, loop returns

unit is not in required because it has a default, so the model may omit it.

A travel app's assistant uses exactly this handler for search_trains, check_pnr and get_fare, with the schemas generated from the Python functions its backend team already has.

Follow-up questions to expect

  • "How do you validate arguments beyond types?" — Use Pydantic models as the tool's parameter type and generate the schema from them (model_json_schema()), then validate before calling. strict: true guarantees shape; business rules (a date in the future) are still your code's job.
  • "How do you run parallel tool calls?" — Dispatch the blocks concurrently with a thread pool or asyncio.gather, then return all results in one message in any order; the ids do the matching.
  • "What if a tool takes 30 seconds?" — Give every tool a timeout, and return an error result on timeout rather than blocking the loop.