Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is fine-tuning? In what situations does it make sense to fine-tune an LLM?


What you need to know

What actually happens during fine-tuning

The model reads your example, predicts the response one token at a time, and the loss measures how far each prediction was from your target token. Backpropagation then nudges the weights so the target becomes more likely next time. After a few passes over the data, the model's default behaviour on inputs like yours has moved.

You can update every weight (full fine-tuning) or freeze the model and train a small add-on such as a LoRA adapter (parameter-efficient fine-tuning, PEFT). PEFT is the usual choice in 2026.

The main kinds

KindDataWhat it changes
Supervised fine-tuning (SFT)Prompt plus one ideal responseFormat, style, task skill
Preference tuning (DPO, RLHF)Prompt, chosen answer, rejected answerTone, helpfulness, refusals
RL with verifiable rewards (GRPO)Prompts plus a checkerReasoning on maths, code, tasks with a checkable answer
Continued pretrainingRaw domain textVocabulary and "feel" of a new domain or language

Good reasons to fine-tune

  • Behaviour is the problem, not knowledge. Strict JSON every time, a house writing style, a tricky boundary between two classes.
  • You have examples, not rules. Experts can show you 1,000 correct outputs but cannot write a prompt that captures all of them.
  • The prompt has become the cost. A 3,000-token few-shot prompt on every call is slow and expensive; a fine-tuned model can do the same with a 150-token prompt.
  • You need a smaller or private model. Data must stay on your servers, or latency must be low, so a fine-tuned 8B open-weight model replaces a large API model on one task.

Bad reasons

  • "The model doesn't know our documents." Facts that change — prices, policies, stock — belong in retrieval (RAG), where you can update and cite them.
  • No evaluation set. Without a fixed test set you cannot tell whether the fine-tune helped.
  • You have not tried a good prompt yet. A strong prompt is the baseline every fine-tune must beat.

A real-life example

An Indian fashion brand needs product descriptions for 40,000 new items each season, in its own voice: warm, specific about fabric and craft, no clichés like "elevate your style".

  • Prompt only: a large model with 8 example descriptions in the prompt. That is about 3,200 input tokens per call, or 128 million tokens a season. The brand team approves 70% of drafts.
  • Fine-tuned: a LoRA adapter on an 8B open model, trained on 2,500 descriptions the brand team had already approved. The prompt drops to about 150 tokens (6 million tokens a season), and on 300 held-out items the approval rate is 88%.

Notice what was not trained in: price, fabric and size data still come from the product catalogue in each prompt. The model learned the voice, not the facts.

Follow-up questions to expect

  • "Can fine-tuning teach the model new knowledge?" — A little, but badly: new facts are learned slowly, are hard to update, and research in 2024 found that training on facts the model did not already know tends to increase hallucination. Use retrieval for knowledge.
  • "How many examples do you need?" — A few hundred good examples can fix format and style; a classifier with many classes may need a few thousand. Train on 25%, 50% and 100% of your data — if the score is still climbing, more data will help.
  • "How do you know fine-tuning worked?" — Compare it against the base model with your best prompt, on a held-out test set built before training, plus a small general-skills check to catch forgetting.