LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What is benchmark saturation, and why does it occur?


What you need to know

Why the ranking becomes meaningless

Two limits squeeze the useful range:

  • Noise. On 500 items at 90%, the 95% confidence interval is about ±2.6 points (1.96 × sqrt(0.9 × 0.1 / 500) ≈ 0.026). Models at 90% and 91.5% are not distinguishable.
  • Label noise ceiling. If 3% of items have wrong keys, the real maximum is about 97%. Models between 92% and 97% are fighting over a few genuinely hard items and some luck.

Why it happens

  1. Real progress. The test was built for weaker models. GSM8K was hard in 2021 and routine a few years later.
  2. Goodhart's law. "When a measure becomes a target, it stops being a good measure." Benchmarks influence data mixes, fine-tuning choices and which model gets released.
  3. Contamination. Test items and answers appear on the web and end up in training data, so part of the score is memory.
  4. Broken items. Mislabelled and ambiguous questions set a common ceiling.

How the field responds

  • Harder versions: MMLU-Pro over MMLU, GPQA's expert-written questions, Humanity's Last Exam.
  • Fresher items: live benchmarks that keep adding new problems after model cutoffs (for example LiveCodeBench for code).
  • Verified subsets and then successors: SWE-bench to SWE-bench Verified to newer sets, as each was found to have issues.
  • Private held-out test sets that are never published.

A real-life example

The code-review bot's eval suite has 150 seeded-bug PRs. After a year of improvements, every candidate model and prompt passes 96 to 98%. A new model scores 97.3% versus 96.7% — a difference of one PR, well inside noise. The suite can no longer guide decisions.

The team moves the 150 cases into a fast smoke test (it still catches big breaks), then builds a new "hard" set from the last three months of production misses: 60 PRs where human reviewers found bugs the bot missed, mostly concurrency and cross-file issues. Current systems score 38% on it. That set now separates candidates, and the team plans to refresh it every quarter.

Follow-up questions to expect

  • "Is a saturated benchmark useless?" — Not entirely; it is still a regression check — a sudden drop means something broke. It just cannot rank strong models.
  • "How do you tell saturation from contamination?" — Saturation shows as many models at the ceiling on fresh equivalents too; contamination shows as high scores on old public items but lower scores on newly written, similar items.
  • "How often should an internal eval be refreshed?" — Whenever the top systems are within noise of each other, and on a schedule (quarterly) from recent production failures.