Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

How do SFT (supervised fine-tuning) and alignment training differ in goal and data?


What you need to know

SFTPreference / alignment training
Data(prompt, ideal response)(prompt, chosen, rejected), or prompts plus a reward
Signal"Write exactly this""This is better than that"
LossCross-entropy on response tokensDPO loss, reward model + RL, or GRPO
TeachesFormat, domain skill, task accuracyTone, helpfulness, refusals, safety, reasoning
Label costHigh — an expert writes the whole answerLower — someone picks the better answer

Two sample rows

JSON
{"messages": [{"role": "user", "content": "Summarise: CT chest shows 6 mm nodule..."},              {"role": "assistant", "content": "Key finding: 6 mm nodule in right upper lobe..."}]}{"prompt": [{"role": "user", "content": "Summarise: CT chest shows 6 mm nodule..."}], "chosen": [{"role": "assistant", "content": "Key finding: 6 mm nodule... Follow-up CT advised."}], "rejected": [{"role": "assistant", "content": "Normal chest CT. No significant findings."}]}

Why SFT alone is not enough

  • SFT can only be as good as its demonstrations. It cannot express "never do this", because every token in the data is something to copy.
  • Preference training gives a negative signal. The rejected answer shows the model what to move away from — invented numbers, missed findings, rude tone.
  • For tasks with a checkable answer, RL with a reward (GRPO) lets the model find reasoning paths better than any demonstration.

Why the order matters

Preference training assumes the model already writes answers in roughly the right shape. DPO on a model that cannot yet follow instructions only teaches it to prefer one poor answer over another. So: SFT, then preference training. ORPO (section 3) combines both into one stage.

A real-life example

A medical-report summariser for a diagnostics chain:

  1. SFT on 2,000 summaries that radiologists wrote or edited. The model learns the format: key findings first, then measurements, then recommended follow-up.
  2. Preference training. The SFT model sometimes softens an abnormal finding into "no significant findings" or rounds a 6 mm nodule to "small". Radiologists compare two model summaries for 1,500 reports and pick the safer, more exact one. DPO on those pairs teaches the model to move away from these specific errors.

SFT gave the shape. Preference training removed errors that were hard to fix by adding more demonstrations.

Follow-up questions to expect

  • "Can you skip SFT?" — If you start from an instruct model that already has the right format, you can go straight to DPO. ORPO combines both stages in one.
  • "How much data does each need?" — SFT for a narrow task: a few hundred to a few thousand examples. Preference: often a few thousand pairs. Quality and agreement between labellers matter more than size.
  • "What is RLHF's role compared with DPO?" — Both are preference training. RLHF uses a reward model and online RL; DPO learns from the pairs directly.