Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
How can smaller models be used effectively for tool delegation?
What you need to know
Why it saves money
Small models such as Claude Haiku 4.5 cost a fraction of a top model per token and respond faster. In an agent, most calls are routine — pick a lookup, fill three fields, format a result. Paying top-model prices for those calls is waste. The skill is splitting the work so the small model gets only easy decisions.
Pattern 1: the router
# config.py — model tiers live here, not in business codeMODELS = {"router": "claude-haiku-4-5", "main": "claude-opus-5"}1ROUTES = [route_tool("track_order"), route_tool("cancel_order"),2 route_tool("refund_status"), route_tool("needs_human_or_complex")]34def handle(message):5 choice = llm.call([{"role": "user", "content": message}], ROUTES,6 model=MODELS["router"], tool_choice={"type": "any"})7 route = next(b for b in choice.content if b.type == "tool_use")8 if route.name == "needs_human_or_complex":9 return run_agent(message, model=MODELS["main"]) # escalate10 return HANDLERS[route.name](**route.input) # cheap, fixed pathThe small model makes one decision with 4 options. The escape route needs_human_or_complex is what keeps quality high: the router does not need to solve hard cases, only recognise them.
Pattern 2: planner and executor
A large model writes a plan once. A small model (or plain code) executes each step: "call get_order with this id and return the status field". Each executor call is short and has one clear job.
Pattern 3: cascade
Try the small model first; validate its output (schema, business rules, confidence). If it fails, retry once with the error, then escalate to the large model.
What makes it work
- Short menu: under about 10 clearly different tools per call.
- Enums and strict schemas instead of free-text fields.
- One decision per call, not a chain of reasoning.
- An evaluation set: a few hundred real requests labelled with the correct tool and arguments. Compare small and large models on it before switching.
A real-life example
A food-delivery support bot handles 400,000 chats a day. Traces show 70% are "where is my order", "cancel my order" or "refund status" — each needs one lookup and a templated reply.
The team adds a small-model router in front of the large-model agent. On a labelled set of 2,000 chats, the router picks the right route 97% of the time; most of its mistakes go to needs_human_or_complex, which is a safe failure. In production, 68% of chats never reach the large model. Cost per chat drops by roughly 60%, and median reply time falls from 6 s to 2 s for routed chats.
One mistake they fixed: the router first had 14 routes, including near-duplicates like refund_status and refund_query. Accuracy was 89%. Merging to 6 clearly different routes lifted it to 97%.
Follow-up questions to expect
- "How do you choose the escalation threshold?" — From the eval set: accept small-model decisions where its accuracy matches the large model's, and escalate the categories where it does not.
- "Can the same model be both router and agent?" — Yes, at a lower effort setting for routing. Measure that first; it avoids maintaining two model tiers and keeps one prompt cache.
- "What about latency?" — The router adds one short call, but saves several large-model calls on most traffic, so median latency usually falls.