Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Implement tool selection logic based on user query.
What you need to know
With three tools you can put every schema in the prompt and let the model choose. With 200 tools (common once you connect a few MCP servers), that fails: the prompt is huge, costs more, and selection accuracy drops because similar tools confuse the model. Tool routing narrows the menu first.
| tier | how | cost | good for |
|---|---|---|---|
| Rules | regex or keyword match | microseconds | order ids, arithmetic, slash commands |
| Embedding shortlist | cosine between query and tool descriptions | one embedding call | narrowing 200 tools to 5 |
| LLM choice | model picks from the short menu | one model call | the final, fuzzy decision |
Tool descriptions are the most important input. A description should say what the tool does, when to use it, and when not to. Adding two or three example user queries to the embedded text raises shortlist accuracy a lot for little effort.
1import re2from collections.abc import Callable3import numpy as np45class ToolRouter:6 """Embeds tool descriptions once; shortlists tools for a query by cosine."""78 def __init__(self, tools: list[dict], embed_fn: Callable, threshold: float = 0.25):9 self.tools, self.embed_fn, self.threshold = tools, embed_fn, threshold10 self.matrix = None11 if tools:12 texts = [f"{t['name']}: {t['description']} Examples: {'; '.join(t.get('examples', []))}"13 for t in tools]14 v = np.asarray(embed_fn(texts), dtype=np.float32)15 self.matrix = v / (np.linalg.norm(v, axis=1, keepdims=True) + 1e-10)1617 def shortlist(self, query: str, k: int = 5) -> list[tuple[dict, float]]:18 if self.matrix is None:19 return []20 q = np.asarray(self.embed_fn([query])[0], dtype=np.float32)21 scores = self.matrix @ (q / (np.linalg.norm(q) + 1e-10))22 return [(self.tools[i], float(scores[i]))23 for i in np.argsort(-scores, kind="stable")[:k] if scores[i] >= self.threshold]2425RULES = [26 (re.compile(r"^(?=.*\d)[\d\s+\-*/().]+$"), "calculator"), # pure arithmetic27 (re.compile(r"\border\s*#?\d{5,}\b", re.I), "order_lookup"), # "order #123456"28]2930def route(query: str, router: ToolRouter, llm_fn: Callable[[str], str]) -> str | None:31 """Return a tool name, or None to answer without tools."""32 for pattern, name in RULES: # tier 1: free and deterministic33 if pattern.search(query):34 return name35 candidates = router.shortlist(query) # tier 2: narrow the menu36 if not candidates:37 return None38 if len(candidates) == 1:39 return candidates[0][0]["name"]40 menu = "\n".join(f"- {t['name']}: {t['description']}" for t, _ in candidates)41 choice = llm_fn(f"Tools:\n{menu}\n\nQuery: {query}\nReply with one tool name, or NONE.").strip()42 return choice if choice in {t["name"] for t, _ in candidates} else NoneThe tricky parts:
(?=.*\d)is a lookahead that requires at least one digit, so a query of only spaces or brackets does not match "arithmetic".- The threshold in
shortlistis what makesNonepossible. Without it, "hello" is always routed to the least-bad tool. - Validating the model's choice against the menu. Models sometimes answer with a tool that exists elsewhere, or a slightly different spelling; anything not in the set becomes
None, never a blind call.
Complexity: rules are O(r × query length). Construction embeds n tool texts once. Each shortlist is one embedding call plus O(n·d) and an O(n log n) sort. The model call is one request with a menu of at most k tools.
A real-life example
A toy embedder counting three words (order, refund, weather) is enough to trace it:
1VOCAB = ["order", "refund", "weather"]2toy_embed = lambda ts: [[float(t.lower().count(w)) for w in VOCAB] for t in ts]3tools = [{"name": "order_status", "description": "Track an order", "examples": ["where is my order"]},4 {"name": "refund_request", "description": "Start a refund for an order",5 "examples": ["I want a refund"]},6 {"name": "weather", "description": "Weather for a city", "examples": ["weather in Pune"]}]7router = ToolRouter(tools, toy_embed)8llm = lambda prompt: "refund_request"910for q in ["23 * 4", "status of order #445566", "refund my order please", "tell me a joke"]:11 print(q, "->", route(q, router, llm))12# 23 * 4 -> calculator13# status of order #445566 -> order_lookup14# refund my order please -> refund_request15# tell me a joke -> None| query | tier that decided | why |
|---|---|---|
23 * 4 | rules | only digits and operators |
| status of order #445566 | rules | "order" followed by a 6-digit number |
| refund my order please | model | shortlist had refund_request (0.894) and order_status (0.707); model picked from those two |
| tell me a joke | shortlist | query vector is all zeros, every score is 0, below 0.25 → None |
A food-delivery app with separate tools for order tracking, refunds, coupons and restaurant search routes millions of messages this way, and most of them never need the model's choice at all.
Follow-up questions to expect
- "How do you measure routing quality?" — Keep a labelled set of (query, expected tool) pairs and track accuracy on every change; adding one new tool can silently steal queries from an old one.
- "Two tools keep getting confused?" — Merge them, or rewrite both descriptions to say explicitly when to use one and not the other.
- "Does the provider have something built in?" — Several APIs now support tool search: tools are marked as deferred, and the model searches the catalogue for the ones it needs. It is the same shortlist idea, run server-side.