Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

How do you handle incorrect tool selection by an agent?


What you need to know

Measure before fixing

Build 200–500 real requests, each labelled with the right tool (and arguments). Run the agent's first decision on each and build a confusion table:

ExpectedChosen insteadCount
get_bookingsearch_hotels31
cancel_bookingmodify_booking12
search_flights(no tool, answered from memory)9

Now you know exactly which boundaries to fix.

Fixes, in order

  1. Sharpen the boundary. Add "Use when… Do not use for… — use X instead" to both tools in a confused pair.
  2. Shrink the menu. Accuracy drops as tool count grows. Merge near-duplicates; load only tools for the current phase; or mark rarely used tools with defer_loading: true so the model searches for them (tool search, supported by Anthropic and OpenAI). OpenAI's allowed_tools tool choice can also restrict a turn to a subset.
  3. Add examples. Two or three request → correct-call examples in the system prompt or tool definitions fix much of what remains.
  4. Validate and redirect. If search_hotels is called with a booking reference like PNR-8X2KQ, return an error: "This looks like a booking reference. Use get_booking."
  5. Route explicitly. For flows where a wrong tool is costly, a small classifier or code decides the route, and the agent only gets that route's tools.

The "answered from memory" case

Sometimes the wrong choice is no tool — the model answers a live question from training data. Fix with trigger conditions in the description ("Always use for current prices; your own knowledge of fares is out of date") and, for critical facts, a validator that rejects answers containing numbers not seen in any tool result.

A real-life example

A travel assistant with 26 tools had 88% tool-call accuracy. The confusion table showed three problems: get_booking vs search_hotels (users saying "my hotel in Jaipur"), cancel_booking vs modify_booking ("change my flight to cancel the return leg"), and 40 fare questions answered with no tool.

Fixes: sharper descriptions for both pairs; tools grouped by phase so post-booking chats load only 7 booking-management tools; a validator that redirects booking-reference patterns; and "never quote fares without search_flights" in that tool's description. Accuracy rose to 97% on the same set, re-run after each change so they knew which fix helped most — it was the phase grouping.

Follow-up questions to expect

  • "How many tools is too many?" — There is no fixed number; accuracy falls as tools overlap and grow. Measure on your eval set, and consider tool search once definitions take thousands of tokens.
  • "Can fine-tuning fix tool selection?" — Yes for stable, high-volume tool sets, but try descriptions, menu size and examples first — they are cheaper to change.
  • "What if two tools really do overlap?" — Merge them, or make one call the other; overlap in the menu is a design smell.