Course Content
RAG Systems
12 sections · 66 lessons
How is RAG different from fine-tuning an LLM?
What you need to know
Side by side
| RAG | Fine-tuning | |
|---|---|---|
| What changes | The prompt | The weights |
| Best at | Facts, freshness, private data | Output format, tone, task behaviour |
| Updating one fact | Re-index one document (seconds) | New training run (hours) and new evaluation |
| Citations | Yes — you know which chunk was used | No |
| Delete data or filter per user | Delete a row, add a filter | Not possible without retraining |
| Cost per request | More input tokens, plus retrieval time | Shorter prompts, cheaper per call |
| Up-front cost | Build an ingestion pipeline and index | Collect labelled examples, train, evaluate |
Why fine-tuning is a poor way to add facts
Fine-tuning can make a model repeat facts it saw many times, but it is unreliable for rare or specific facts, like one clause in one contract. Research on this (for example, Gekhman et al., 2024) found that training on facts the model did not already know can make it hallucinate more, because it learns to answer confidently beyond what it knows. You also cannot see which training example produced an answer.
Where the two meet
- Fine-tune the generator to use retrieval better. RAFT (retrieval-augmented fine-tuning, 2024) trains a model on questions paired with a mix of useful and distracting documents, so it learns to quote the right one and ignore the rest.
- Fine-tune the retriever. Training the embedding model on your own query-passage pairs often improves retrieval on specialist text more than changing the chat model does.
A real-life example
A law firm builds a legal-contract search tool. Lawyers want two things.
- "What is the liability cap in our master services agreement with Client X?" The answer is in one of 12,000 contracts, and contracts are added every week. This is RAG: retrieve the limitation-of-liability clause, show it with the page number.
- Every answer must follow the firm's format: a one-line answer, the quoted clause, then a risk note in fixed categories. The base model follows this format about 80% of the time in their tests, even with examples in the prompt. The team fine-tunes a smaller model with LoRA on 1,500 answers written by associates. Format compliance improves, and the prompt no longer needs 2,000 tokens of examples.
The fine-tuned model still gets its facts from retrieval. If they had fine-tuned on the contracts instead, every new contract would need a new training run, and nobody could cite a page.
Follow-up questions to expect
- "Can fine-tuning and RAG be combined?" — Yes. Fine-tune for the answer format or domain style, and retrieve the facts. RAFT goes further and fine-tunes the model to read retrieved documents well.
- "What about continued pre-training on a domain?" — Training on large amounts of raw domain text (legal, medical) can improve vocabulary and style, but it still does not give citations or per-user filtering.
- "Which is cheaper?" — RAG is cheaper to start and to update. Fine-tuning can be cheaper per request at very high volume, because prompts get shorter and a smaller model may be enough.