Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your agent has 200 internal APIs. Sending all tool descriptions costs $2/request and the model picks wrong 30% of the time. How do you route to the right tool without sending every spec to the model?


Retrieval over 214 toolsUser messageSearch thetool index,embedded onceTop 12tools,recall 96%Model picksone of 12Selection evalon every releaseTool-spec tokens fall from about 60K to 4K.
Fewer, sharper choices raise accuracy as well as cutting cost; rewriting confusable descriptions does the rest.

What you need to know

With 200 tools in the prompt, the model has to compare the request against 200 descriptions, many of them near-duplicates such as get_customer, get_customer_profile and fetch_customer_details. It picks a plausible one, and the prompt costs a lot on every call.

Retrieval over tools

Python
# built once, and whenever tools changefor tool in catalogue:    text = f"{tool.name}: {tool.description}\nExamples: {'; '.join(tool.example_queries)}"    tool_index.add(id=tool.name, vector=embed(text), payload={"domain": tool.domain})# per requestcandidates = tool_index.search(embed(user_message), k=12)response = llm.invoke(messages, tools=[catalogue[c.id].schema for c in candidates])

Example queries in the embedded text matter a lot. Users say "where's my parcel", not "retrieve shipment tracking status". Adding three real phrasings per tool helps retrieval match how people actually ask.

Some providers now offer a built-in tool-search feature that does this server-side; the principle is the same.

Hierarchical routing

  1. Group — organise tools into 10 to 15 domains: billing, CRM, logistics, HR.
  2. Pick the domain — a cheap, fast model call or a classifier chooses one or two domains.
  3. Offer that domain's tools — or retrieve within the domain if it is still large.
  4. Fall back — if confidence is low, retrieve across all domains.

This works better than pure similarity when names overlap across domains, for example create_invoice in both billing and procurement. If tools already sit behind MCP servers, each server is a natural domain.

Compare the options

ApproachTokens per requestSelection qualityComplexity
All 200 toolsVery highPoor with near-duplicatesLowest
Retrieve top 12About 6% of the aboveGood if descriptions are goodLow
Domain router, then toolsLowBest when domains overlapMedium

Fix the descriptions too

Routing reduces the 30% error rate but won't remove it, because part of it is vague or overlapping descriptions. Mine the misfires from logs, then rewrite:

Text
get_order_status: Current delivery status of ONE order by order_id.Use for: "where is my order", "has it shipped".Do NOT use for refunds (use get_refund_status) or order history (use list_orders).

Measure it

A tool-selection eval: several hundred real queries labelled with the correct tool. Score two things: router recall@k (is the right tool in the candidates?) and final selection accuracy. If recall@k is low, fix retrieval; if recall is high but selection is wrong, fix descriptions.

A real-life example

Scenario (illustrative numbers). A bank's internal operations assistant exposes 214 APIs. Each request carries about 60,000 tokens of tool specs, and on a 500-query eval, the model picks the right tool 70% of the time.

The team adds tool retrieval with three example queries per tool, returning the top 12. Router recall@12 is 96%, selection accuracy rises to 86%, and tool-spec tokens fall to about 4,000 per request. Mining the remaining errors shows 30 pairs of confusable tools; rewriting those descriptions with "do not use for" lines lifts accuracy to 93%. A domain router is added later for the "accounts" and "cards" tools, which share many names.

Follow-up questions to expect

  • "What if the right tool isn't retrieved?" — Log requests where the model says it lacks a tool, add them as eval cases, and improve that tool's description and examples. Consider always including a small set of core tools.
  • "How do you keep the index in sync?" — Re-embed a tool whenever its definition changes, as part of the deployment that changes it.
  • "Does retrieval add latency?" — One embedding call and a vector search, typically tens of milliseconds, and the smaller prompt usually saves more than that.