Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

How do you measure whether a fine-tuned model is actually better (and not just overfit)?


Clause accuracy before and after removing leaked test items94%86%+888%85%+3Fine-tunedPrompted baseGapFull test setClean 78% only22% of test clauses were near-copies of training boilerplate; the margin of error on 230 items is about 4 points.
Most of the headline gain was contamination — the real gain only showed up on 600 freshly labelled clauses.

What you need to know

Three layers of evaluation

LayerWhat it tells youLimit
Validation loss / perplexityWhether the model fits unseen data at allLower loss does not mean better product output
Task metrics on a fixed test setAccuracy, F1, JSON validity, field-level correctnessOnly as good as the test set
Pairwise judgement (humans or LLM judge)Which output people prefer, base vs fine-tunedJudges have biases; randomise order

The right baselines

Beating your own previous checkpoint proves little. Compare with:

  • the base model with your best prompt (the fine-tune must beat this to be worth it);
  • the base model plus RAG if the task needs knowledge;
  • a larger API model, if that is the realistic alternative.

Contamination and noise

  • Contamination. Search the training set for near-duplicates of each test item (embedding similarity or n-gram overlap). Leaked items make scores look better than they are.
  • Noise. With 300 test items and 80% accuracy, the 95% margin of error is about ±4.5 points (1.96 × sqrt(0.8 × 0.2 / 300) ≈ 0.045). A 2-point gain on 300 items may be luck. Compare both models on the same items and count where they differ, or grow the test set.

Forgetting and online checks

Run a general regression suite (knowledge, maths, instruction following, safety refusals) on base and fine-tuned models. Then roll out to a small share of traffic and track product metrics: task success, human escalation rate, user corrections, latency and cost per request. Offline evals earn the right to an A/B test; they do not replace it.

A real-life example

The law firm's clause classifier scores 94% on the test set, against 86% for the prompted base model. It looks like a big win.

A contamination check finds that 22% of test clauses are near-copies of the firm's boilerplate clauses in the training set. On the clean 78%, the fine-tune scores 88% against 85%. With about 230 clean items, the margin of error is roughly ±4 points, so the team cannot yet call it a win. They label 600 fresh clauses from new contracts. On those, the fine-tune leads by 4 points overall — and by 15 points on arbitration clauses, the class the partners care most about. Per-class results were the real story.

Follow-up questions to expect

  • "Why isn't validation loss enough?" — Loss measures how likely the reference text is, token by token. A model can lower loss and still get the one important field wrong, or produce valid but unhelpful answers.
  • "How big should a test set be?" — Big enough that the difference you care about is larger than the noise. For a few points of accuracy, that usually means several hundred to a few thousand items.
  • "What about LLM-as-a-judge?" — Useful for open-ended outputs, if you calibrate it against human labels and randomise answer order. Section 5 covers its biases.