Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
How do you measure whether a fine-tuned model is actually better (and not just overfit)?
What you need to know
Three layers of evaluation
| Layer | What it tells you | Limit |
|---|---|---|
| Validation loss / perplexity | Whether the model fits unseen data at all | Lower loss does not mean better product output |
| Task metrics on a fixed test set | Accuracy, F1, JSON validity, field-level correctness | Only as good as the test set |
| Pairwise judgement (humans or LLM judge) | Which output people prefer, base vs fine-tuned | Judges have biases; randomise order |
The right baselines
Beating your own previous checkpoint proves little. Compare with:
- the base model with your best prompt (the fine-tune must beat this to be worth it);
- the base model plus RAG if the task needs knowledge;
- a larger API model, if that is the realistic alternative.
Contamination and noise
- Contamination. Search the training set for near-duplicates of each test item (embedding similarity or n-gram overlap). Leaked items make scores look better than they are.
- Noise. With 300 test items and 80% accuracy, the 95% margin of error is about ±4.5 points (
1.96 × sqrt(0.8 × 0.2 / 300) ≈ 0.045). A 2-point gain on 300 items may be luck. Compare both models on the same items and count where they differ, or grow the test set.
Forgetting and online checks
Run a general regression suite (knowledge, maths, instruction following, safety refusals) on base and fine-tuned models. Then roll out to a small share of traffic and track product metrics: task success, human escalation rate, user corrections, latency and cost per request. Offline evals earn the right to an A/B test; they do not replace it.
A real-life example
The law firm's clause classifier scores 94% on the test set, against 86% for the prompted base model. It looks like a big win.
A contamination check finds that 22% of test clauses are near-copies of the firm's boilerplate clauses in the training set. On the clean 78%, the fine-tune scores 88% against 85%. With about 230 clean items, the margin of error is roughly ±4 points, so the team cannot yet call it a win. They label 600 fresh clauses from new contracts. On those, the fine-tune leads by 4 points overall — and by 15 points on arbitration clauses, the class the partners care most about. Per-class results were the real story.
Follow-up questions to expect
- "Why isn't validation loss enough?" — Loss measures how likely the reference text is, token by token. A model can lower loss and still get the one important field wrong, or produce valid but unhelpful answers.
- "How big should a test set be?" — Big enough that the difference you care about is larger than the noise. For a few points of accuracy, that usually means several hundred to a few thousand items.
- "What about LLM-as-a-judge?" — Useful for open-ended outputs, if you calibrate it against human labels and randomise answer order. Section 5 covers its biases.