Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is LLM-as-a-judge evaluation, and what’s the main benefit of using it?


What you need to know

Two modes

Pointwise

  • Score one answer against a rubric (1–5, or pass/fail)
  • Good for absolute tracking over time
  • Scores drift and cluster at the top

Pairwise

  • Compare two answers: A, B or tie
  • Good for "is the fine-tune better than the base?"
  • More reliable, since relative judgments are easier

A rubric that works

Text
You are checking a summary of a medical report.Report: {report}Summary: {summary}Check each point and answer in JSON:1. unsupported_claims: list every statement in the summary not supported by the report2. missing_critical: list any abnormal finding in the report missing from the summary3. format_ok: true if the summary has Findings and Impression sectionsExplain briefly first, then give the JSON.

Concrete checks with evidence beat "rate the quality from 1 to 10". Asking for a short reason before the verdict makes grades more consistent.

Known biases and fixes

BiasWhat happensFix
PositionPrefers the first (or second) answerRun both orders; count a win only if it holds both ways
LengthPrefers longer answersRubric that penalises padding; check length in results
Self-preferencePrefers answers in its own styleUse a judge from a different model family where you can
LeniencyMost answers get 4 or 5Binary checks, or pairwise comparisons

Calibrate before you trust

Label 100–200 examples by hand, run the judge on them, and report the agreement rate (or Cohen's kappa). The 2023 MT-Bench paper found GPT-4 as a judge agreed with human experts over 80% of the time, similar to how often humans agreed with each other — but that is one setting; your task needs its own number. If agreement is 60%, the judge is a rough signal, not a metric.

A real-life example

A hospital chain uses a judge in CI for its medical-report summariser. Every new adapter is scored on 500 held-out reports within 20 minutes.

Before relying on it, they compared the judge with 150 doctor-rated summaries. On "unsupported claims" the judge agreed with doctors 88% of the time, so that check blocks releases automatically. On "clinically useful", agreement was only 61%, so doctors still review a monthly sample for that. They also found the judge missed unit errors — "5 mg" written as "5 mcg" — so they added a simple code check for units. The judge handles what needs language understanding; code handles what can be checked exactly. (Made-up numbers for illustration.)

Follow-up questions to expect

  • "Can a judge grade a model from its own family?" — It can, but self-preference bias is a risk. Use a different family or confirm with human spot-checks.
  • "Pairwise or pointwise for comparing a fine-tune with its base?" — Pairwise, in both orders, and report the win rate with ties.
  • "Can a judge be a training reward?" — Yes (RLAIF), but the policy may learn to exploit the judge's biases, so keep human checks on the outputs.