Course Content
RAG Systems
12 sections · 66 lessons
When should you prefer RAG over fine-tuning?
What you need to know
The key question is: when the model fails, is it missing information, or is it using the information badly?
A quick decision table
| Situation | Choose | Why |
|---|---|---|
| Prices, policies or stock change weekly | RAG | Update by re-indexing, not retraining |
| Answers must cite a document and page | RAG | Retrieval knows the source |
| Users have different permissions | RAG | Filter at retrieval time |
| "Right to be forgotten" requests | RAG | Delete the chunk; weights cannot forget |
| Output must always be a fixed JSON schema | Structured output first, then fine-tune if needed | It is a behaviour |
| Brand tone the model keeps drifting from | Prompt, then fine-tune | It is a behaviour |
| Long prompt makes latency too high at scale | Fine-tune (distil) | Moves prompt instructions into weights |
| Small, static corpus under ~100k tokens | Maybe neither: long context with caching | Simplest thing that works |
How to tell which failure you have
Build an evaluation set of 50 to 100 real questions with known answers. For each wrong answer, check the prompt the model actually saw.
- The answer was not in the retrieved context. That is a retrieval failure. Fine-tuning will not help. Fix chunking, the embedding model, hybrid search or filters.
- The answer was in the context, and the model still got it wrong or used the wrong format. That is a generation failure. Try a clearer prompt first; if it persists across many cases, fine-tuning is now justified.
This check is what separates a considered answer from a guess in the interview.
A real-life example
An e-commerce marketplace wants a product Q&A box on every product page: "Does this mixer grinder work on 110 volts?", "Is the jar dishwasher safe?"
The catalogue has 2 million products. Sellers edit listings every day, and prices change hourly. A product manager suggests fine-tuning a model on the catalogue. The engineer shows why that fails: a training run every day would still be a day out of date, the model could mix up two similar products, and it could not show which listing it read.
They build RAG instead. Each query is filtered to the current product ID, so retrieval searches only that product's listing, specifications, seller answers and top reviews, about 3,000 tokens. The model answers and links the spec line it used.
Six months later they fine-tune a small model for one narrow job: rewriting answers into a short, friendly style in Hindi and English. That is a behaviour, and the fine-tuned small model is cheaper per call than the large one. Retrieval is unchanged.
Follow-up questions to expect
- "When would you fine-tune first?" — When there is no knowledge gap at all: for example, classifying support tickets into 40 categories where the model understands the text but keeps picking the wrong label.
- "How much data does a LoRA fine-tune need?" — Often a few hundred to a few thousand high-quality examples for a narrow behaviour. Quality matters more than volume.
- "What if the model ignores retrieved context?" — First make the prompt clearer and put the question after the context. If it still prefers its own memory, fine-tuning to read context (as RAFT does) is a reasonable next step.