Course Content
LLMOps & Deployment
6 sections · 40 lessons
What is semantic routing, and how does it work in multi-model setups?
What you need to know
Three ways to build a router
| Router | Latency | Accuracy | When |
|---|---|---|---|
| Embedding similarity to example queries | ~5–20 ms | Good for clear intents | Start here |
| Small trained classifier on your labelled traffic | ~5–20 ms | Best per millisecond | Once you have a few thousand labels |
| A small LLM as router | Hundreds of ms | Handles nuance | Complex or overlapping intents |
How the embedding router works
For each route, collect about 50 example queries and average their embeddings into a centroid. For a new query, compute its embedding and the cosine similarity to each centroid. Take the best match if it passes a threshold; otherwise use the default route.
1import math23def cosine(a, b):4 dot = sum(x * y for x, y in zip(a, b))5 return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))67# Centroids: the average embedding of ~50 labelled example queries per route8ROUTES = {9 "application_status": [0.82, 0.10, 0.05],10 "scheme_eligibility": [0.08, 0.85, 0.20],11 "grievance": [0.05, 0.25, 0.90],12}1314def route(query_vec, threshold=0.80):15 name, score = max(((n, cosine(query_vec, c)) for n, c in ROUTES.items()),16 key=lambda pair: pair[1])17 return (name, round(score, 2)) if score >= threshold else ("default_llm", round(score, 2))1819print(route([0.79, 0.15, 0.02])) # ('application_status', 1.0)20print(route([0.40, 0.45, 0.45])) # ('default_llm', 0.77)Real embeddings have hundreds of dimensions; three are used here to keep it readable. The second query is ambiguous, scores below the threshold, and goes to the default route rather than guessing.
What routing is used for
- Cost cascade — simple questions to a small model, hard ones to a large model.
- Specialised prompts — billing, technical and legal prompts with their own tools.
- Non-LLM paths — status lookups served from an API with a template.
- Language — route to the model that is best in that language.
- Skip retrieval when the question does not need it.
A real-life example
A state government chatbot gets 2 million questions a month in 12 languages. Analysis shows 38% are "what is the status of my application?", 30% are eligibility questions, 12% are grievances, and 20% are everything else.
The router sends status questions to a deterministic path: extract the application number, call the status API, fill a template in the user's language — no large-model call at all. Eligibility questions go to a RAG prompt with the scheme rules. Grievances go to a structured-output prompt that files a ticket. Everything else, and anything under the threshold, goes to the general model.
Model cost falls by about 40%, and status answers are now always correct because they come from the database. The router is evaluated monthly on 1,000 labelled queries; when a new scheme launches, its questions first land in the default route, and the team adds a new route after a week of examples.
Follow-up questions to expect
- "How do you evaluate a router?" — Separately from the answer: route accuracy on labelled queries, plus the cost of each kind of mistake (a hard question on a small model costs quality; an easy one on a large model costs money).
- "How do you pick the threshold?" — On a labelled set, choose the value that keeps wrong routings low for high-risk routes; send uncertain cases to the default.
- "Is this the same as a model cascade?" — A cascade tries the cheap model first and escalates on low confidence; a router decides up front. Many systems combine both.