Course Content
Advanced RAG
3 sections · 38 lessons
How does adaptive retrieval decide if external data is needed?
What you need to know
Why not retrieve every time
Unconditional retrieval has two costs. It wastes time and money on messages like "thanks!" or "rewrite this more politely". Worse, it can hurt: the retriever always returns something, and irrelevant passages pull the answer off course. The goal is to retrieve when the answer depends on knowledge the model does not reliably have — private, recent or very specific facts.
The signals, cheapest first
| Signal | How it works | Cost |
|---|---|---|
| Rules | Greetings, thanks, pure maths skip retrieval | ~0 ms |
| Query classifier | Embedding or small-LLM classifier: chit-chat, general, domain, account-specific | 5–300 ms |
| Model confidence | Draft an answer; retrieve if token probabilities are low (the FLARE idea) | One extra draft |
| Trained self-check | Self-RAG emits a Retrieve token when it needs evidence | Needs a fine-tuned model |
| Complexity routing | Adaptive-RAG: no retrieval, single-step, or multi-step pipeline | A classifier call |
| Tool calling | The model has a search tool and decides itself | Built into the main call |
Adaptive-RAG (2024) trains a small classifier to label the query's complexity and sends it to the matching pipeline, so simple questions never pay for an iterative loop and hard multi-hop questions get one.
The modern default: let the model call a tool
With current models, the most common adaptive retrieval is agentic: describe a search_docs tool clearly ("use this for anything about our products, policies or account rules") and let the model decide. It works with any provider, needs no training, and can call search several times. The downsides are that the decision is harder to audit and models sometimes answer from memory when they should have searched. Strong instructions and evaluation fix most of that.
Evaluate it as a classifier
Label a few hundred real queries as "needs retrieval" or "doesn't". Then read the confusion matrix, not a single accuracy number:
- False skip (needed retrieval, skipped it) → the model answers from memory → a likely hallucination. Expensive.
- False retrieve (didn't need it, retrieved anyway) → extra latency and some distraction. Cheap.
So set the threshold to favour retrieving, and log every skip so you can review them.
A real-life example
An Indian telecom's support bot handles about 200,000 messages a day. The team samples 1,000 and finds:
- 18% are greetings, thanks or "ok";
- 9% are general questions ("what is 5G SA?");
- 31% are about the customer's own account ("why was I charged ₹49?"), which need a billing API, not documents;
- 42% need the help articles ("how to activate eSIM on my phone").
They add a small classifier in front of the pipeline with four routes: reply directly, answer from the model, call the account tool, or retrieve. Retrieval traffic falls to under half of what it was, and the account questions stop getting irrelevant help articles pasted into their prompts.
In the first week's review of logged skips, they find that "what is the validity of the ₹299 plan?" was classified as general knowledge. The model answered from memory with an old validity. They move plan questions to the retrieve route and lower the skip threshold. The rule they adopt: when the classifier is unsure, retrieve.
Follow-up questions to expect
- "How is this different from query routing?" — Adaptive retrieval is one routing decision (retrieve or not, and how hard). Query routing is the general version: RAG, a SQL tool, an API, a human.
- "Can the main model just decide by itself?" — Yes, with a search tool. It is simple and flexible; the trade-off is less predictable cost and a decision you must evaluate from logs.
- "What if the model is confidently wrong?" — Confidence signals miss that case. That is why domain-specific questions should retrieve by rule, and only generic ones use confidence.