Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your agent has 200 internal APIs. Sending all tool descriptions costs $2/request and the model picks wrong 30% of the time. How do you route to the right tool without sending every spec to the model?
What you need to know
With 200 tools in the prompt, the model has to compare the request against 200 descriptions, many of them near-duplicates such as get_customer, get_customer_profile and fetch_customer_details. It picks a plausible one, and the prompt costs a lot on every call.
Retrieval over tools
1# built once, and whenever tools change2for tool in catalogue:3 text = f"{tool.name}: {tool.description}\nExamples: {'; '.join(tool.example_queries)}"4 tool_index.add(id=tool.name, vector=embed(text), payload={"domain": tool.domain})56# per request7candidates = tool_index.search(embed(user_message), k=12)8response = llm.invoke(messages, tools=[catalogue[c.id].schema for c in candidates])Example queries in the embedded text matter a lot. Users say "where's my parcel", not "retrieve shipment tracking status". Adding three real phrasings per tool helps retrieval match how people actually ask.
Some providers now offer a built-in tool-search feature that does this server-side; the principle is the same.
Hierarchical routing
- Group — organise tools into 10 to 15 domains: billing, CRM, logistics, HR.
- Pick the domain — a cheap, fast model call or a classifier chooses one or two domains.
- Offer that domain's tools — or retrieve within the domain if it is still large.
- Fall back — if confidence is low, retrieve across all domains.
This works better than pure similarity when names overlap across domains, for example create_invoice in both billing and procurement. If tools already sit behind MCP servers, each server is a natural domain.
Compare the options
| Approach | Tokens per request | Selection quality | Complexity |
|---|---|---|---|
| All 200 tools | Very high | Poor with near-duplicates | Lowest |
| Retrieve top 12 | About 6% of the above | Good if descriptions are good | Low |
| Domain router, then tools | Low | Best when domains overlap | Medium |
Fix the descriptions too
Routing reduces the 30% error rate but won't remove it, because part of it is vague or overlapping descriptions. Mine the misfires from logs, then rewrite:
get_order_status: Current delivery status of ONE order by order_id.Use for: "where is my order", "has it shipped".Do NOT use for refunds (use get_refund_status) or order history (use list_orders).Measure it
A tool-selection eval: several hundred real queries labelled with the correct tool. Score two things: router recall@k (is the right tool in the candidates?) and final selection accuracy. If recall@k is low, fix retrieval; if recall is high but selection is wrong, fix descriptions.
A real-life example
Scenario (illustrative numbers). A bank's internal operations assistant exposes 214 APIs. Each request carries about 60,000 tokens of tool specs, and on a 500-query eval, the model picks the right tool 70% of the time.
The team adds tool retrieval with three example queries per tool, returning the top 12. Router recall@12 is 96%, selection accuracy rises to 86%, and tool-spec tokens fall to about 4,000 per request. Mining the remaining errors shows 30 pairs of confusable tools; rewriting those descriptions with "do not use for" lines lifts accuracy to 93%. A domain router is added later for the "accounts" and "cards" tools, which share many names.
Follow-up questions to expect
- "What if the right tool isn't retrieved?" — Log requests where the model says it lacks a tool, add them as eval cases, and improve that tool's description and examples. Consider always including a small set of core tools.
- "How do you keep the index in sync?" — Re-embed a tool whenever its definition changes, as part of the deployment that changes it.
- "Does retrieval add latency?" — One embedding call and a vector search, typically tens of milliseconds, and the smaller prompt usually saves more than that.