Course Content
Advanced RAG
3 sections · 38 lessons
When is RAG unnecessary, and what are simpler alternatives?
What you need to know
A RAG pipeline is a real system: parsing, chunking, an embedding model, a vector store, retrieval tuning, reranking, freshness jobs, evaluation. Each part can fail. An interviewer asking this question wants to see that you would not build all that by reflex.
The alternatives, in order of cost
| Situation | Better tool | Why |
|---|---|---|
| General knowledge, reasoning, rewriting, code | Just prompt the model | Retrieval adds latency and distractors |
| Small, stable corpus (a handbook, a 30-page policy) | Put it all in the prompt, cached | No retrieval misses; simpler |
| Answer is a number or a record | A tool: SQL or an API | Exact and current; RAG over table dumps is fragile |
| Need a format, tone, schema or dialect | Fine-tuning or examples in the prompt | That is behaviour, not knowledge |
| Templated queries, small corpus | Keyword search or a lookup table | Deterministic and cheap |
A rule of thumb
- RAG is for knowledge that is too large, too private or too fresh for the model's weights.
- Fine-tuning is for behaviour — how the model writes and structures its output.
- Tools are for live data and actions — balances, order status, counts.
Many good systems combine them: a tool for the account data, a small cached prefix for the core policy, and RAG only for the long tail.
When RAG is the right answer
Choose RAG when at least one of these holds:
- The corpus is far larger than a context window.
- Content changes daily and you cannot keep re-sending it.
- Different users may see different documents (per-user access control).
- Answers must cite a specific passage for audit.
If none applies, the honest interview answer is "I would not build RAG here", followed by what you would build instead.
A real-life example
A platform team at a software company plans an "engineering-wiki assistant" for three kinds of question. Before building, they sample 300 real Slack questions and sort them:
- 38% about the API style guide — a single 22-page document that changes twice a year. They put the whole guide (about 15,000 tokens) in a cached system prompt. No retrieval, no chunking, no misses.
- 27% about live state — "is the staging deploy blocked?", "how many P1 incidents are open?". These go to tools that call the deploy system and the incident tracker's API. A RAG index over exported incident pages would always be out of date.
- 35% about the rest of the wiki — 30,000 pages of runbooks and design docs, updated daily, some restricted to certain teams. This part gets RAG, with hybrid search, reranking and access filters.
The result is a smaller RAG system aimed only at the questions that need it. The team avoids the classic trap of indexing everything, including the style guide, and then debugging why top-5 retrieval sometimes misses rule 14 on page 19.
Follow-up questions to expect
- "How big can the 'just put it in the prompt' option go?" — Test it. Accuracy, cost and latency all get worse as the prompt grows; many teams find the comfort zone ends well before the model's maximum window.
- "When would you fine-tune and use RAG?" — When you need both new knowledge and a specific behaviour — for example, a fixed report format over retrieved clinical documents. Fine-tune for the format; retrieve for the facts.
- "Why not RAG over database rows?" — Similarity search cannot count, filter exactly or join. "How many orders failed yesterday?" needs a query, not the five most similar rows.