LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What are the main limitations of AI-based evaluation?


What you need to know

The main limitations

LimitationWhat it looks likeMitigation
Systematic biasPrefers first position, longer or more formatted answers, its own family's textSwap order, control for length, cross-family judge
Shared blind spotsMisses a wrong dosage or tax rule the generator also got wrongGive the judge the source or a reference; use human experts for specialist domains
Poor absolute calibration70% of answers get 4 out of 5Binary checks, or pairwise comparison
InstabilityRerunning or rewording the rubric moves the average by pointsFixed prompt, temperature 0 where allowed, repeats
Non-stationarityScores shift when the judge model version changesPin the version; re-calibrate before switching
Cost and latencyAnother model call per graded itemSampling, cheap checks first, small judges for simple checks
GameabilityGenerator tuned against the judge learns its quirksHold-out human evaluation, rotate judges

Why bias doesn't average out

Random noise shrinks with more samples. Bias does not: if the judge adds half a point for every extra paragraph, 10,000 samples just measure that bias more precisely. A larger sample gives a narrower confidence interval around the wrong number.

The competence limit

A judge can only verify what it can check. With the source document in the prompt, a mid-sized model can confirm "15 days" appears. Without a source, asking any model whether a claim about Indian GST rules is correct means asking its memory — the same memory that may have produced the error.

A real-life example

The code-review bot's team uses a judge to rate review comments as "useful" or "not useful". After three months of prompt tuning, the judge's useful-rate climbed from 61% to 83%. But developers were resolving fewer of the bot's comments, not more.

A human review of 100 comments showed why. The tuned prompts had made comments longer, more polite, and full of hedged suggestions ("you might consider..."). The judge rewarded the style; developers found them noisy. The team added a length-controlled pairwise comparison, a human-labelled set of 200 comments marked "developer acted on it", and a rule that no prompt change ships on judge score alone. The judge's useful-rate on the human set dropped to 64% — closer to reality.

Follow-up questions to expect

  • "Is an LLM judge better than no judge?" — Yes, if calibrated and used for what it is good at: narrow, checkable criteria. An uncalibrated holistic judge can be worse than none because it creates false confidence.
  • "How would you detect judge drift?" — Re-run a fixed calibration set on a schedule and alert when agreement with the stored human labels drops, or when the score distribution shifts with no system change.
  • "Would a panel of judges fix bias?" — It reduces the bias of any single model family, at higher cost; biases common to all LLMs, such as length preference, remain.