Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Leadership wants to fine-tune a model on the company's documents so it "knows" them. RAG or fine-tuning?


Two different kinds of learningRAG: changes what it sees• Facts arrive in the prompt• Revised guideline: re-index one file• Every answer links its source• Filtered per userFine-tune: changes what it is• Behaviour stored in weights• Revised guideline: retrain• No source to show• Best for format, tone, layout
Fine-tune for the shape of the answer and retrieve for its facts; most teams that need both are solving two problems.

What you need to know

The request "make the model know our documents" mixes up two kinds of learning.

RAG changes what the model sees

  • Facts arrive in the prompt for each request
  • Update one document: re-index it in seconds
  • Every answer can cite its source chunk
  • Filter per user: each person sees only what they may see

Fine-tuning changes what the model is

  • Behaviour is stored in the weights
  • Update a fact: run training again
  • No source to point to
  • Cannot filter or delete per user

Fine-tuning on raw documents does teach some facts, but unreliably: the model may recall a fact in one phrasing and not another, blend two versions of a policy, or still invent details. And you cannot show the user where an answer came from.

A decision rule

SituationChooseWhy
Answers must cite sources, data changes often, or access differs per userRAGFacts stay editable, traceable and filterable
Facts are right but the output shape is wrong: tone, JSON, taxonomy, lengthFine-tuneThat is a behaviour problem
A frontier model already works on a narrow, high-volume task, but costs too muchFine-tune a small model on its outputs (distillation)Same behaviour, lower cost and latency
Both a knowledge and a format problemBoth: fine-tune for format, RAG for factsThey solve different problems

Before either, try a proper prompt with a few good examples. A surprising share of "we need fine-tuning" turns out to be an under-specified prompt.

The order of work to propose

  1. Build the eval set — real questions with reference answers.
  2. Ship RAG — with citations and an "I don't know" path.
  3. Measure — find what still fails.
  4. Fine-tune only against a measured failure — for example, "answers are correct but ignore our 5-part reply template 30% of the time".

This protects the team from spending six weeks on a fine-tune and being unable to prove it helped.

A real-life example

Scenario (illustrative numbers). A pharmaceutical distributor's leadership wants a model fine-tuned on 8,000 product and regulatory documents. The engineer runs a quick test: a LoRA fine-tune on the documents, and a basic RAG setup, both scored on 150 real questions from the sales team.

The fine-tuned model answers 58% correctly and cites nothing. When a storage-temperature guideline is revised a week later, it keeps giving the old value. RAG answers 84% correctly, with a document link on every answer, and the update takes one re-index. The remaining RAG failures are mostly formatting: sales reps want a fixed "product, dosage, storage, caution" layout. A small fine-tune on 1,500 well-formatted example answers, used with RAG, fixes the layout without touching the facts.

Follow-up questions to expect

  • "Doesn't fine-tuning reduce hallucination?" — Not reliably. It can make the model sound more confident in the domain while still inventing specifics; grounding and citations are the stronger control.
  • "What about very long context instead of RAG?" — For a small, stable document set, putting it all in the prompt with caching can work. For thousands of changing documents with permissions, retrieval is cheaper and enforceable.
  • "How much data does a fine-tune need?" — For a format or style behaviour, often a few hundred to a few thousand high-quality examples; quality matters more than count.