Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Leadership wants to fine-tune a model on the company's documents so it "knows" them. RAG or fine-tuning?
What you need to know
The request "make the model know our documents" mixes up two kinds of learning.
RAG changes what the model sees
- Facts arrive in the prompt for each request
- Update one document: re-index it in seconds
- Every answer can cite its source chunk
- Filter per user: each person sees only what they may see
Fine-tuning changes what the model is
- Behaviour is stored in the weights
- Update a fact: run training again
- No source to point to
- Cannot filter or delete per user
Fine-tuning on raw documents does teach some facts, but unreliably: the model may recall a fact in one phrasing and not another, blend two versions of a policy, or still invent details. And you cannot show the user where an answer came from.
A decision rule
| Situation | Choose | Why |
|---|---|---|
| Answers must cite sources, data changes often, or access differs per user | RAG | Facts stay editable, traceable and filterable |
| Facts are right but the output shape is wrong: tone, JSON, taxonomy, length | Fine-tune | That is a behaviour problem |
| A frontier model already works on a narrow, high-volume task, but costs too much | Fine-tune a small model on its outputs (distillation) | Same behaviour, lower cost and latency |
| Both a knowledge and a format problem | Both: fine-tune for format, RAG for facts | They solve different problems |
Before either, try a proper prompt with a few good examples. A surprising share of "we need fine-tuning" turns out to be an under-specified prompt.
The order of work to propose
- Build the eval set — real questions with reference answers.
- Ship RAG — with citations and an "I don't know" path.
- Measure — find what still fails.
- Fine-tune only against a measured failure — for example, "answers are correct but ignore our 5-part reply template 30% of the time".
This protects the team from spending six weeks on a fine-tune and being unable to prove it helped.
A real-life example
Scenario (illustrative numbers). A pharmaceutical distributor's leadership wants a model fine-tuned on 8,000 product and regulatory documents. The engineer runs a quick test: a LoRA fine-tune on the documents, and a basic RAG setup, both scored on 150 real questions from the sales team.
The fine-tuned model answers 58% correctly and cites nothing. When a storage-temperature guideline is revised a week later, it keeps giving the old value. RAG answers 84% correctly, with a document link on every answer, and the update takes one re-index. The remaining RAG failures are mostly formatting: sales reps want a fixed "product, dosage, storage, caution" layout. A small fine-tune on 1,500 well-formatted example answers, used with RAG, fixes the layout without touching the facts.
Follow-up questions to expect
- "Doesn't fine-tuning reduce hallucination?" — Not reliably. It can make the model sound more confident in the domain while still inventing specifics; grounding and citations are the stronger control.
- "What about very long context instead of RAG?" — For a small, stable document set, putting it all in the prompt with caching can work. For thousands of changing documents with permissions, retrieval is cheaper and enforceable.
- "How much data does a fine-tune need?" — For a format or style behaviour, often a few hundred to a few thousand high-quality examples; quality matters more than count.