Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Implement tool selection logic based on user query.


What you need to know

With three tools you can put every schema in the prompt and let the model choose. With 200 tools (common once you connect a few MCP servers), that fails: the prompt is huge, costs more, and selection accuracy drops because similar tools confuse the model. Tool routing narrows the menu first.

tierhowcostgood for
Rulesregex or keyword matchmicrosecondsorder ids, arithmetic, slash commands
Embedding shortlistcosine between query and tool descriptionsone embedding callnarrowing 200 tools to 5
LLM choicemodel picks from the short menuone model callthe final, fuzzy decision

Tool descriptions are the most important input. A description should say what the tool does, when to use it, and when not to. Adding two or three example user queries to the embedded text raises shortlist accuracy a lot for little effort.

Python
import refrom collections.abc import Callableimport numpy as npclass ToolRouter:    """Embeds tool descriptions once; shortlists tools for a query by cosine."""    def __init__(self, tools: list[dict], embed_fn: Callable, threshold: float = 0.25):        self.tools, self.embed_fn, self.threshold = tools, embed_fn, threshold        self.matrix = None        if tools:            texts = [f"{t['name']}: {t['description']} Examples: {'; '.join(t.get('examples', []))}"                     for t in tools]            v = np.asarray(embed_fn(texts), dtype=np.float32)            self.matrix = v / (np.linalg.norm(v, axis=1, keepdims=True) + 1e-10)    def shortlist(self, query: str, k: int = 5) -> list[tuple[dict, float]]:        if self.matrix is None:            return []        q = np.asarray(self.embed_fn([query])[0], dtype=np.float32)        scores = self.matrix @ (q / (np.linalg.norm(q) + 1e-10))        return [(self.tools[i], float(scores[i]))                for i in np.argsort(-scores, kind="stable")[:k] if scores[i] >= self.threshold]RULES = [    (re.compile(r"^(?=.*\d)[\d\s+\-*/().]+$"), "calculator"),     # pure arithmetic    (re.compile(r"\border\s*#?\d{5,}\b", re.I), "order_lookup"),   # "order #123456"]def route(query: str, router: ToolRouter, llm_fn: Callable[[str], str]) -> str | None:    """Return a tool name, or None to answer without tools."""    for pattern, name in RULES:                        # tier 1: free and deterministic        if pattern.search(query):            return name    candidates = router.shortlist(query)               # tier 2: narrow the menu    if not candidates:        return None    if len(candidates) == 1:        return candidates[0][0]["name"]    menu = "\n".join(f"- {t['name']}: {t['description']}" for t, _ in candidates)    choice = llm_fn(f"Tools:\n{menu}\n\nQuery: {query}\nReply with one tool name, or NONE.").strip()    return choice if choice in {t["name"] for t, _ in candidates} else None

The tricky parts:

  • (?=.*\d) is a lookahead that requires at least one digit, so a query of only spaces or brackets does not match "arithmetic".
  • The threshold in shortlist is what makes None possible. Without it, "hello" is always routed to the least-bad tool.
  • Validating the model's choice against the menu. Models sometimes answer with a tool that exists elsewhere, or a slightly different spelling; anything not in the set becomes None, never a blind call.

Complexity: rules are O(r × query length). Construction embeds n tool texts once. Each shortlist is one embedding call plus O(n·d) and an O(n log n) sort. The model call is one request with a menu of at most k tools.

A real-life example

A toy embedder counting three words (order, refund, weather) is enough to trace it:

Python
VOCAB = ["order", "refund", "weather"]toy_embed = lambda ts: [[float(t.lower().count(w)) for w in VOCAB] for t in ts]tools = [{"name": "order_status", "description": "Track an order", "examples": ["where is my order"]},         {"name": "refund_request", "description": "Start a refund for an order",          "examples": ["I want a refund"]},         {"name": "weather", "description": "Weather for a city", "examples": ["weather in Pune"]}]router = ToolRouter(tools, toy_embed)llm = lambda prompt: "refund_request"for q in ["23 * 4", "status of order #445566", "refund my order please", "tell me a joke"]:    print(q, "->", route(q, router, llm))# 23 * 4 -> calculator# status of order #445566 -> order_lookup# refund my order please -> refund_request# tell me a joke -> None
querytier that decidedwhy
23 * 4rulesonly digits and operators
status of order #445566rules"order" followed by a 6-digit number
refund my order pleasemodelshortlist had refund_request (0.894) and order_status (0.707); model picked from those two
tell me a jokeshortlistquery vector is all zeros, every score is 0, below 0.25 → None

A food-delivery app with separate tools for order tracking, refunds, coupons and restaurant search routes millions of messages this way, and most of them never need the model's choice at all.

Follow-up questions to expect

  • "How do you measure routing quality?" — Keep a labelled set of (query, expected tool) pairs and track accuracy on every change; adding one new tool can silently steal queries from an old one.
  • "Two tools keep getting confused?" — Merge them, or rewrite both descriptions to say explicitly when to use one and not the other.
  • "Does the provider have something built in?" — Several APIs now support tool search: tools are marked as deferred, and the model searches the catalogue for the ones it needs. It is the same shortlist idea, run server-side.