LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What is semantic routing, and how does it work in multi-model setups?


Where a citizen's question goesEmbed thequestionCosine to eachroute centroidBest score ator above 0.80?Yes: statusAPI, RAG orticket promptNo: defaultgeneral model38 percent of questions are status checks that never need a large model.
The default route is what keeps a brand-new kind of question from being forced into the nearest wrong one.

What you need to know

Three ways to build a router

RouterLatencyAccuracyWhen
Embedding similarity to example queries~5–20 msGood for clear intentsStart here
Small trained classifier on your labelled traffic~5–20 msBest per millisecondOnce you have a few thousand labels
A small LLM as routerHundreds of msHandles nuanceComplex or overlapping intents

How the embedding router works

For each route, collect about 50 example queries and average their embeddings into a centroid. For a new query, compute its embedding and the cosine similarity to each centroid. Take the best match if it passes a threshold; otherwise use the default route.

Python
import mathdef cosine(a, b):    dot = sum(x * y for x, y in zip(a, b))    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))# Centroids: the average embedding of ~50 labelled example queries per routeROUTES = {    "application_status": [0.82, 0.10, 0.05],    "scheme_eligibility": [0.08, 0.85, 0.20],    "grievance":          [0.05, 0.25, 0.90],}def route(query_vec, threshold=0.80):    name, score = max(((n, cosine(query_vec, c)) for n, c in ROUTES.items()),                      key=lambda pair: pair[1])    return (name, round(score, 2)) if score >= threshold else ("default_llm", round(score, 2))print(route([0.79, 0.15, 0.02]))   # ('application_status', 1.0)print(route([0.40, 0.45, 0.45]))   # ('default_llm', 0.77)

Real embeddings have hundreds of dimensions; three are used here to keep it readable. The second query is ambiguous, scores below the threshold, and goes to the default route rather than guessing.

What routing is used for

  • Cost cascade — simple questions to a small model, hard ones to a large model.
  • Specialised prompts — billing, technical and legal prompts with their own tools.
  • Non-LLM paths — status lookups served from an API with a template.
  • Language — route to the model that is best in that language.
  • Skip retrieval when the question does not need it.

A real-life example

A state government chatbot gets 2 million questions a month in 12 languages. Analysis shows 38% are "what is the status of my application?", 30% are eligibility questions, 12% are grievances, and 20% are everything else.

The router sends status questions to a deterministic path: extract the application number, call the status API, fill a template in the user's language — no large-model call at all. Eligibility questions go to a RAG prompt with the scheme rules. Grievances go to a structured-output prompt that files a ticket. Everything else, and anything under the threshold, goes to the general model.

Model cost falls by about 40%, and status answers are now always correct because they come from the database. The router is evaluated monthly on 1,000 labelled queries; when a new scheme launches, its questions first land in the default route, and the team adds a new route after a week of examples.

Follow-up questions to expect

  • "How do you evaluate a router?" — Separately from the answer: route accuracy on labelled queries, plus the cost of each kind of mistake (a hard question on a small model costs quality; an easy one on a large model costs money).
  • "How do you pick the threshold?" — On a labelled set, choose the value that keeps wrong routings low for high-risk routes; send uncertain cases to the default.
  • "Is this the same as a model cascade?" — A cascade tries the cheap model first and escalates on low confidence; a router decides up front. Many systems combine both.