Course Content
LLM Evaluation
6 sections · 50 lessons
What is benchmark saturation, and why does it occur?
What you need to know
Why the ranking becomes meaningless
Two limits squeeze the useful range:
- Noise. On 500 items at 90%, the 95% confidence interval is about ±2.6 points (
1.96 × sqrt(0.9 × 0.1 / 500) ≈ 0.026). Models at 90% and 91.5% are not distinguishable. - Label noise ceiling. If 3% of items have wrong keys, the real maximum is about 97%. Models between 92% and 97% are fighting over a few genuinely hard items and some luck.
Why it happens
- Real progress. The test was built for weaker models. GSM8K was hard in 2021 and routine a few years later.
- Goodhart's law. "When a measure becomes a target, it stops being a good measure." Benchmarks influence data mixes, fine-tuning choices and which model gets released.
- Contamination. Test items and answers appear on the web and end up in training data, so part of the score is memory.
- Broken items. Mislabelled and ambiguous questions set a common ceiling.
How the field responds
- Harder versions: MMLU-Pro over MMLU, GPQA's expert-written questions, Humanity's Last Exam.
- Fresher items: live benchmarks that keep adding new problems after model cutoffs (for example LiveCodeBench for code).
- Verified subsets and then successors: SWE-bench to SWE-bench Verified to newer sets, as each was found to have issues.
- Private held-out test sets that are never published.
A real-life example
The code-review bot's eval suite has 150 seeded-bug PRs. After a year of improvements, every candidate model and prompt passes 96 to 98%. A new model scores 97.3% versus 96.7% — a difference of one PR, well inside noise. The suite can no longer guide decisions.
The team moves the 150 cases into a fast smoke test (it still catches big breaks), then builds a new "hard" set from the last three months of production misses: 60 PRs where human reviewers found bugs the bot missed, mostly concurrency and cross-file issues. Current systems score 38% on it. That set now separates candidates, and the team plans to refresh it every quarter.
Follow-up questions to expect
- "Is a saturated benchmark useless?" — Not entirely; it is still a regression check — a sudden drop means something broke. It just cannot rank strong models.
- "How do you tell saturation from contamination?" — Saturation shows as many models at the ceiling on fresh equivalents too; contamination shows as high scores on old public items but lower scores on newly written, similar items.
- "How often should an internal eval be refreshed?" — Whenever the top systems are within noise of each other, and on a schedule (quarterly) from recent production failures.