Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
How can smaller models be used effectively for tool delegation beyond simple routing?
What you need to know
Lever 1: narrow the menu
A small model choosing among 30 tools makes many more mistakes than one choosing among 4. Ways to cut the menu:
- By phase. During "find the booking", only
search_bookingsandget_bookingare loaded; during "change the booking", onlyget_fare_rulesandmodify_booking. - Tool search. Declare rarely used tools with
defer_loading: trueand let the model search for them. Only matching schemas enter the context. - Merge near-duplicates. Two tools with overlapping descriptions are a routing bug waiting to happen.
Lever 2: constrain the output
strict: trueso every call matches the schema.enumfor every closed set (status, cabin class, currency).- IDs that come from earlier tool results, with the description saying so.
- Few parameters; supply defaults in code rather than asking the model.
Lever 3: validate, retry once, escalate
- Validate — does the order id exist, is the date in the future, is the amount within limits?
- Retry with the reason — return an
is_errorresult such as "date 2026-02-30 is not a real date; use YYYY-MM-DD". Small models usually fix a clear error on the next turn. - Escalate — if the second attempt fails, hand the whole step to the larger model.
A small model that is right 92% of the time, behind a validator that catches most of its mistakes and escalates them, can deliver close to large-model accuracy on most traffic at a fraction of the cost.
Lever 4: fine-tune when the tools stop changing
Log the large model's correct calls: (request, context) → (tool, arguments). A few thousand clean examples can noticeably lift a small model's tool-call accuracy on your tools. The catch: the fine-tuned model is tied to that schema. Every tool change means new data and retraining, so do this only for stable, high-volume tools.
What not to delegate
- Multi-step plans with dependencies.
- Decisions that need to weigh conflicting evidence.
- Anything irreversible without a check after it.
A real-life example
An airline's rebooking assistant uses a small model to fill modify_booking(pnr, new_flight_id, fare_diff_inr) after a large model has chosen the new flight.
At first, 7% of calls failed: wrong PNR format, a flight ID copied from the wrong search result, or a fare difference with the wrong sign. The team made three changes: a regex pattern on pnr in the schema with strict mode on; the executor sees only the one chosen flight, not the full search list; and the fare difference is computed by code, not by the model. Failures dropped to 0.8%, and every remaining failure was caught by the validator and escalated.
The lesson: the biggest gain came from removing decisions from the small model, not from better prompting.
Follow-up questions to expect
- "Do small models support parallel tool calls?" — Many do, but they are more likely to make mistakes in multi-call turns. For small models, one call per turn is often more reliable.
- "How do you measure 'tool-call accuracy'?" — Right tool, right arguments (exact or normalised match), scored per step on a labelled set.
- "Is fine-tuning better than prompting?" — Only once prompting, schemas and narrowing have been exhausted and the tool set is stable; otherwise it adds retraining cost for every change.