Course Content
Advanced RAG
3 sections · 38 lessons
How do you route queries between RAG, direct LLM responses, or other systems?
What you need to know
Why routing matters
Many "RAG gives bad answers" complaints are really routing failures: a question that needed a database query was sent to document search, or a simple rewrite request dragged in five irrelevant chunks. Each kind of question has a best path:
| Query type | Best route | Example |
|---|---|---|
| General or writing task | Direct LLM | "Make this email more formal" |
| Policy or document knowledge | RAG | "What is the re-KYC period for low-risk customers?" |
| Live data or counts | SQL or API tool | "How many STRs did we file last month?" |
| Actions or personal records | API tool with the user's auth | "Show my open cases" |
| Judgement or high risk | Human handoff | "Should we exit this client relationship?" |
Ways to build a router, cheapest first
- Rules and regex — greetings, "reset my password", known commands. Zero latency, fully predictable.
- Embedding router — for each route, embed 20–100 example queries; at run time, compare the query's embedding with each route's examples (or their centroid) and pick the best above a threshold. Milliseconds, no LLM call, and hard to beat for well-separated routes.
- Small LLM classifier — a tool call or structured output with an
enumof route names, not free text. Handles nuance; costs one small call. - Native tool calling — expose
search_policies,query_transactionsandcreate_case_noteas tools and let the main model choose. Most flexible, can combine routes in one answer, but hardest to predict and evaluate.
Real systems combine them: rules first, then embeddings, then an LLM only when the cheaper layers are unsure.
Design points
- Default route. When nothing matches confidently, send to the safest broad route (usually RAG with a "no answer" floor), never to an error.
- Fan out when unsure. If the top two routes are close, run both in parallel and let the generator use whichever returns relevant results. A redundant call costs little; a wrong route costs the answer.
- Different errors, different costs. Sending a transaction question to RAG produces a confident wrong number — worse than sending a policy question to SQL, which returns nothing. Weight the router accordingly.
- Log every decision with its confidence. Misroutes are invisible otherwise.
- Access control per route. Tools must run with the user's permissions; routing is not a security boundary.
A real-life example
A bank's compliance assistant serves 400 compliance officers. After launch, reviewers find that 20% of bad answers came from questions like "How many high-risk accounts in the Pune region haven't completed re-KYC?" The RAG pipeline answered from a policy document that mentioned re-KYC rules and made up a number.
The team adds a layered router:
- Rules: messages with "draft", "rewrite", "summarise this" and attached text go to the direct LLM route.
- Embedding router with five routes built from 60 labelled example queries each: policy (RAG), metrics (SQL over reporting views), case lookup (case API), drafting, and escalate.
- Small LLM classifier when the top two embedding scores are within 0.05 of each other.
The SQL route runs read-only queries on approved reporting views with the officer's data permissions. Questions asking for legal interpretation ("Is this structure a breach of FEMA?") go to the escalate route, which drafts a note to the legal team rather than answering.
On 600 labelled questions, the confusion matrix shows the main remaining error is "metrics questions phrased as policy questions" ("what's our current re-KYC backlog policy?"). The team adds 30 such examples to the metrics route. Made-up numbers in answers stop appearing in the weekly review.
Follow-up questions to expect
- "How do you get training examples for the router?" — Start with a few dozen hand-written per route, then add real queries from logs, especially misroutes found in review.
- "Why not let the main LLM decide everything with tools?" — It works and is flexible, but it adds variance and cost, and it is harder to test. Cheap layers in front handle the obvious majority predictably.
- "How does routing relate to adaptive retrieval?" — Adaptive retrieval is the "retrieve or not" slice of routing; a full router also chooses among tools, indexes and humans.