Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is LLM-as-a-judge evaluation, and what’s the main benefit of using it?
What you need to know
Two modes
Pointwise
- Score one answer against a rubric (1–5, or pass/fail)
- Good for absolute tracking over time
- Scores drift and cluster at the top
Pairwise
- Compare two answers: A, B or tie
- Good for "is the fine-tune better than the base?"
- More reliable, since relative judgments are easier
A rubric that works
You are checking a summary of a medical report.Report: {report}Summary: {summary}Check each point and answer in JSON:1. unsupported_claims: list every statement in the summary not supported by the report2. missing_critical: list any abnormal finding in the report missing from the summary3. format_ok: true if the summary has Findings and Impression sectionsExplain briefly first, then give the JSON.Concrete checks with evidence beat "rate the quality from 1 to 10". Asking for a short reason before the verdict makes grades more consistent.
Known biases and fixes
| Bias | What happens | Fix |
|---|---|---|
| Position | Prefers the first (or second) answer | Run both orders; count a win only if it holds both ways |
| Length | Prefers longer answers | Rubric that penalises padding; check length in results |
| Self-preference | Prefers answers in its own style | Use a judge from a different model family where you can |
| Leniency | Most answers get 4 or 5 | Binary checks, or pairwise comparisons |
Calibrate before you trust
Label 100–200 examples by hand, run the judge on them, and report the agreement rate (or Cohen's kappa). The 2023 MT-Bench paper found GPT-4 as a judge agreed with human experts over 80% of the time, similar to how often humans agreed with each other — but that is one setting; your task needs its own number. If agreement is 60%, the judge is a rough signal, not a metric.
A real-life example
A hospital chain uses a judge in CI for its medical-report summariser. Every new adapter is scored on 500 held-out reports within 20 minutes.
Before relying on it, they compared the judge with 150 doctor-rated summaries. On "unsupported claims" the judge agreed with doctors 88% of the time, so that check blocks releases automatically. On "clinically useful", agreement was only 61%, so doctors still review a monthly sample for that. They also found the judge missed unit errors — "5 mg" written as "5 mcg" — so they added a simple code check for units. The judge handles what needs language understanding; code handles what can be checked exactly. (Made-up numbers for illustration.)
Follow-up questions to expect
- "Can a judge grade a model from its own family?" — It can, but self-preference bias is a risk. Use a different family or confirm with human spot-checks.
- "Pairwise or pointwise for comparing a fine-tune with its base?" — Pairwise, in both orders, and report the win rate with ties.
- "Can a judge be a training reward?" — Yes (RLAIF), but the policy may learn to exploit the judge's biases, so keep human checks on the outputs.