RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How is RAG different from fine-tuning an LLM?


Where the new knowledge livesRAG changes what it sees• Facts arrive in the prompt per request• Update: re-index one document• Every answer can cite its chunk• Delete or filter a row per userFine-tuning changes what it is• Behaviour stored in the weights• Update: a new training run• No source to point to• Cannot forget or filter per user
Fine-tune for form, retrieve for facts — the contract tool does both and still cites every clause.

What you need to know

Side by side

RAGFine-tuning
What changesThe promptThe weights
Best atFacts, freshness, private dataOutput format, tone, task behaviour
Updating one factRe-index one document (seconds)New training run (hours) and new evaluation
CitationsYes — you know which chunk was usedNo
Delete data or filter per userDelete a row, add a filterNot possible without retraining
Cost per requestMore input tokens, plus retrieval timeShorter prompts, cheaper per call
Up-front costBuild an ingestion pipeline and indexCollect labelled examples, train, evaluate

Why fine-tuning is a poor way to add facts

Fine-tuning can make a model repeat facts it saw many times, but it is unreliable for rare or specific facts, like one clause in one contract. Research on this (for example, Gekhman et al., 2024) found that training on facts the model did not already know can make it hallucinate more, because it learns to answer confidently beyond what it knows. You also cannot see which training example produced an answer.

Where the two meet

  • Fine-tune the generator to use retrieval better. RAFT (retrieval-augmented fine-tuning, 2024) trains a model on questions paired with a mix of useful and distracting documents, so it learns to quote the right one and ignore the rest.
  • Fine-tune the retriever. Training the embedding model on your own query-passage pairs often improves retrieval on specialist text more than changing the chat model does.

A real-life example

A law firm builds a legal-contract search tool. Lawyers want two things.

  1. "What is the liability cap in our master services agreement with Client X?" The answer is in one of 12,000 contracts, and contracts are added every week. This is RAG: retrieve the limitation-of-liability clause, show it with the page number.
  2. Every answer must follow the firm's format: a one-line answer, the quoted clause, then a risk note in fixed categories. The base model follows this format about 80% of the time in their tests, even with examples in the prompt. The team fine-tunes a smaller model with LoRA on 1,500 answers written by associates. Format compliance improves, and the prompt no longer needs 2,000 tokens of examples.

The fine-tuned model still gets its facts from retrieval. If they had fine-tuned on the contracts instead, every new contract would need a new training run, and nobody could cite a page.

Follow-up questions to expect

  • "Can fine-tuning and RAG be combined?" — Yes. Fine-tune for the answer format or domain style, and retrieve the facts. RAFT goes further and fine-tunes the model to read retrieved documents well.
  • "What about continued pre-training on a domain?" — Training on large amounts of raw domain text (legal, medical) can improve vocabulary and style, but it still does not give citations or per-user filtering.
  • "Which is cheaper?" — RAG is cheaper to start and to update. Fine-tuning can be cheaper per request at very high volume, because prompts get shorter and a smaller model may be enough.