Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

How can smaller models be used effectively for tool delegation?


400,000 chats a day through a small routerSmallmodel: pick1 of 6 routestrack, cancel,refund: fixed handlerneeds_human_or_complexrouteLarge-modelagenthandles the rest68 percent of chats never reach the large model.
The router does not need to solve hard cases, only recognise them — the escape route is what keeps quality up.

What you need to know

Why it saves money

Small models such as Claude Haiku 4.5 cost a fraction of a top model per token and respond faster. In an agent, most calls are routine — pick a lookup, fill three fields, format a result. Paying top-model prices for those calls is waste. The skill is splitting the work so the small model gets only easy decisions.

Pattern 1: the router

Python
# config.py — model tiers live here, not in business codeMODELS = {"router": "claude-haiku-4-5", "main": "claude-opus-5"}
Python
ROUTES = [route_tool("track_order"), route_tool("cancel_order"),          route_tool("refund_status"), route_tool("needs_human_or_complex")]def handle(message):    choice = llm.call([{"role": "user", "content": message}], ROUTES,                      model=MODELS["router"], tool_choice={"type": "any"})    route = next(b for b in choice.content if b.type == "tool_use")    if route.name == "needs_human_or_complex":        return run_agent(message, model=MODELS["main"])   # escalate    return HANDLERS[route.name](**route.input)             # cheap, fixed path

The small model makes one decision with 4 options. The escape route needs_human_or_complex is what keeps quality high: the router does not need to solve hard cases, only recognise them.

Pattern 2: planner and executor

A large model writes a plan once. A small model (or plain code) executes each step: "call get_order with this id and return the status field". Each executor call is short and has one clear job.

Pattern 3: cascade

Try the small model first; validate its output (schema, business rules, confidence). If it fails, retry once with the error, then escalate to the large model.

What makes it work

  • Short menu: under about 10 clearly different tools per call.
  • Enums and strict schemas instead of free-text fields.
  • One decision per call, not a chain of reasoning.
  • An evaluation set: a few hundred real requests labelled with the correct tool and arguments. Compare small and large models on it before switching.

A real-life example

A food-delivery support bot handles 400,000 chats a day. Traces show 70% are "where is my order", "cancel my order" or "refund status" — each needs one lookup and a templated reply.

The team adds a small-model router in front of the large-model agent. On a labelled set of 2,000 chats, the router picks the right route 97% of the time; most of its mistakes go to needs_human_or_complex, which is a safe failure. In production, 68% of chats never reach the large model. Cost per chat drops by roughly 60%, and median reply time falls from 6 s to 2 s for routed chats.

One mistake they fixed: the router first had 14 routes, including near-duplicates like refund_status and refund_query. Accuracy was 89%. Merging to 6 clearly different routes lifted it to 97%.

Follow-up questions to expect

  • "How do you choose the escalation threshold?" — From the eval set: accept small-model decisions where its accuracy matches the large model's, and escalate the categories where it does not.
  • "Can the same model be both router and agent?" — Yes, at a lower effort setting for routing. Measure that first; it avoids maintaining two model tiers and keeps one prompt cache.
  • "What about latency?" — The router adds one short call, but saves several large-model calls on most traffic, so median latency usually falls.