Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

How do you decide between fine-tuning, RAG, and prompt engineering for a given problem?


Sort the failures, then fix each by typeEval set: 300real Hindi chatsBest-promptbaselinefails 35%Label eachfailure by causeWrongfacts: add RAGWrongregister:fine-tuneMissed escalations were fixed by one rule in the system prompt.
The cause of a failure picks the tool — facts go to retrieval, behaviour goes to training, unclear rules go to the prompt.

What you need to know

Each option changes a different thing, and each costs a different amount to change later.

Prompt engineeringRAGFine-tuning
What it changesThe instructions on each callThe facts the model sees on each callThe model's weights
FixesUnclear task, missing rulesMissing or changing knowledgeWrong behaviour, format, tone
Time to changeMinutesHours (re-index)Days (data, train, evaluate)
NeedsA good writer and an eval setA document store and a retrieverHundreds to thousands of labelled examples
Can cite sourcesNoYesNo

The two tests interviewers want to hear

  1. Knowledge or behaviour? If the model gets the facts wrong, it needs knowledge — use RAG. If it has the facts but writes them in the wrong way, it needs behaviour — prompt first, then fine-tune.
  2. Will the right answer change tomorrow? If yes, do not put it in weights. Retraining for every price change is slow and costly.

A process, not a guess

  1. Build an eval set — 100–300 real inputs with expected outputs, before choosing anything.
  2. Prompt baseline — the best prompt you can write, with a few examples and a clear output schema.
  3. Sort the failures — label each failure as "missing knowledge", "wrong behaviour" or "task unclear".
  4. Fix by type — RAG for knowledge, a better prompt for unclear tasks, fine-tuning for behaviour that prompting could not fix.
  5. Re-measure — every change must beat the previous best on the same eval set.

Other reasons that tip the decision

  • Cost and latency. A long prompt on a large model can cost more per month than fine-tuning a small model once.
  • Privacy and hosting. If data may not leave your servers, you will self-host an open-weight model, and a small model often needs fine-tuning to match a large one.
  • Long context does not replace RAG. Even with very long context windows, stuffing every document into every call is slow and expensive, and models still miss details buried in the middle.

A real-life example

An Indian e-commerce company builds a Hindi customer-support assistant. On 300 real chats, the prompted model fails 35% of the time. The team sorts the failures:

  • 40% of failures: wrong return windows and delivery charges. These change monthly — knowledge, so they add RAG over the policy pages.
  • 35% of failures: replies in formal textbook Hindi or in English, when customers write casual Hinglish ("mera order kab aayega?"). The model knows Hinglish but does not default to it — behaviour. After prompting fails to fix it reliably, they fine-tune on 3,000 chats that senior agents rated as good.
  • 25% of failures: not escalating angry customers. A clear rule in the system prompt fixes most of these.

The final system is a fine-tuned model, a retriever and a short prompt. Each tool fixes the failure type it is good at.

Follow-up questions to expect

  • "Can you fine-tune a model to use RAG better?" — Yes. Train on examples that include retrieved passages, some of them irrelevant, with answers that quote the right passage or say "not in the documents". This teaches the model to use context and to admit when it is missing.
  • "When would you fine-tune first?" — When the blocker is cost or latency of a large model, or when you must run a small model on-premises or on a device.
  • "What if you have no labelled data?" — Then prompting and RAG are your only options at first. Log production traffic and human corrections; that becomes the fine-tuning dataset later.