Course Content
LLM Evaluation
6 sections · 50 lessons
What are the main limitations of AI-based evaluation?
What you need to know
The main limitations
| Limitation | What it looks like | Mitigation |
|---|---|---|
| Systematic bias | Prefers first position, longer or more formatted answers, its own family's text | Swap order, control for length, cross-family judge |
| Shared blind spots | Misses a wrong dosage or tax rule the generator also got wrong | Give the judge the source or a reference; use human experts for specialist domains |
| Poor absolute calibration | 70% of answers get 4 out of 5 | Binary checks, or pairwise comparison |
| Instability | Rerunning or rewording the rubric moves the average by points | Fixed prompt, temperature 0 where allowed, repeats |
| Non-stationarity | Scores shift when the judge model version changes | Pin the version; re-calibrate before switching |
| Cost and latency | Another model call per graded item | Sampling, cheap checks first, small judges for simple checks |
| Gameability | Generator tuned against the judge learns its quirks | Hold-out human evaluation, rotate judges |
Why bias doesn't average out
Random noise shrinks with more samples. Bias does not: if the judge adds half a point for every extra paragraph, 10,000 samples just measure that bias more precisely. A larger sample gives a narrower confidence interval around the wrong number.
The competence limit
A judge can only verify what it can check. With the source document in the prompt, a mid-sized model can confirm "15 days" appears. Without a source, asking any model whether a claim about Indian GST rules is correct means asking its memory — the same memory that may have produced the error.
A real-life example
The code-review bot's team uses a judge to rate review comments as "useful" or "not useful". After three months of prompt tuning, the judge's useful-rate climbed from 61% to 83%. But developers were resolving fewer of the bot's comments, not more.
A human review of 100 comments showed why. The tuned prompts had made comments longer, more polite, and full of hedged suggestions ("you might consider..."). The judge rewarded the style; developers found them noisy. The team added a length-controlled pairwise comparison, a human-labelled set of 200 comments marked "developer acted on it", and a rule that no prompt change ships on judge score alone. The judge's useful-rate on the human set dropped to 64% — closer to reality.
Follow-up questions to expect
- "Is an LLM judge better than no judge?" — Yes, if calibrated and used for what it is good at: narrow, checkable criteria. An uncalibrated holistic judge can be worse than none because it creates false confidence.
- "How would you detect judge drift?" — Re-run a fixed calibration set on a schedule and alert when agreement with the stored human labels drops, or when the score distribution shifts with no system change.
- "Would a panel of judges fix bias?" — It reduces the bias of any single model family, at higher cost; biases common to all LLMs, such as length preference, remain.