Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
How do you handle incorrect tool selection by an agent?
What you need to know
Measure before fixing
Build 200–500 real requests, each labelled with the right tool (and arguments). Run the agent's first decision on each and build a confusion table:
| Expected | Chosen instead | Count |
|---|---|---|
get_booking | search_hotels | 31 |
cancel_booking | modify_booking | 12 |
search_flights | (no tool, answered from memory) | 9 |
Now you know exactly which boundaries to fix.
Fixes, in order
- Sharpen the boundary. Add "Use when… Do not use for… — use X instead" to both tools in a confused pair.
- Shrink the menu. Accuracy drops as tool count grows. Merge near-duplicates; load only tools for the current phase; or mark rarely used tools with
defer_loading: trueso the model searches for them (tool search, supported by Anthropic and OpenAI). OpenAI'sallowed_toolstool choice can also restrict a turn to a subset. - Add examples. Two or three request → correct-call examples in the system prompt or tool definitions fix much of what remains.
- Validate and redirect. If
search_hotelsis called with a booking reference likePNR-8X2KQ, return an error: "This looks like a booking reference. Use get_booking." - Route explicitly. For flows where a wrong tool is costly, a small classifier or code decides the route, and the agent only gets that route's tools.
The "answered from memory" case
Sometimes the wrong choice is no tool — the model answers a live question from training data. Fix with trigger conditions in the description ("Always use for current prices; your own knowledge of fares is out of date") and, for critical facts, a validator that rejects answers containing numbers not seen in any tool result.
A real-life example
A travel assistant with 26 tools had 88% tool-call accuracy. The confusion table showed three problems: get_booking vs search_hotels (users saying "my hotel in Jaipur"), cancel_booking vs modify_booking ("change my flight to cancel the return leg"), and 40 fare questions answered with no tool.
Fixes: sharper descriptions for both pairs; tools grouped by phase so post-booking chats load only 7 booking-management tools; a validator that redirects booking-reference patterns; and "never quote fares without search_flights" in that tool's description. Accuracy rose to 97% on the same set, re-run after each change so they knew which fix helped most — it was the phase grouping.
Follow-up questions to expect
- "How many tools is too many?" — There is no fixed number; accuracy falls as tools overlap and grow. Measure on your eval set, and consider tool search once definitions take thousands of tokens.
- "Can fine-tuning fix tool selection?" — Yes for stable, high-volume tool sets, but try descriptions, menu size and examples first — they are cheaper to change.
- "What if two tools really do overlap?" — Merge them, or make one call the other; overlap in the menu is a design smell.