Course Content
LLM Evaluation
6 sections · 50 lessons
What are major benchmark datasets, and how should their results be interpreted?
What you need to know
Families
| Capability | Examples |
|---|---|
| Knowledge and reasoning | MMLU, MMLU-Pro, GPQA (Diamond subset), Humanity's Last Exam |
| Maths | GSM8K, MATH, competition sets such as AIME |
| Code and software engineering | HumanEval, MBPP, SWE-bench Verified, SWE-bench Pro, LiveCodeBench |
| Agents and tool use | τ-bench / τ²-bench, Terminal-Bench, WebArena, GAIA, BrowseComp |
| Instruction following | IFEval |
| Long context | RULER, needle-in-a-haystack variants |
| Safety | HarmBench, BBQ, TruthfulQA |
| Human preference | Arena (formerly Chatbot Arena / LMArena) |
Four questions to ask of any number
- Could the test be in the training data? Contamination is likely for anything public and a few years old. In February 2026 OpenAI said it would stop reporting SWE-bench Verified, citing contamination and flawed tests. Prefer benchmarks with items released after the model's training cutoff, or private held-out sets.
- Is it saturated? When top models all score near the ceiling (GSM8K, original MMLU, HumanEval), differences are inside the noise.
- Same harness? Prompt template, number of few-shot examples, reasoning on or off, answer-extraction code and number of attempts can each move a score by several points. Two "MMLU" numbers from different reports are often not comparable.
- What is the uncertainty? On 500 items at 90%, the 95% interval is about ±2.6 points; a 1-point lead means nothing.
Construct validity
A benchmark measures a proxy. SWE-bench measures fixing issues in a set of popular Python repositories with existing tests — not your Java monolith or your code-review style. Once a benchmark becomes a target, labs optimise for it, and the proxy weakens.
A real-life example
The team building the code-review bot must pick a model. Three vendors publish strong coding scores. Rather than choosing the top number, the team shortlists three models whose public scores are in the same band and whose price and latency fit.
They then run their own eval: 120 past pull requests with known bugs from their own repositories, in their own languages (Kotlin and TypeScript). Results differ from the public ranking: the model with the lowest published coding score catches the most bugs in Kotlin and writes the fewest false comments. The public benchmarks saved time by removing weak candidates; the private eval made the decision.
Follow-up questions to expect
- "How do you detect contamination?" — Check whether the model can complete test items verbatim from a prefix, compare performance on items released before and after its cutoff, and look for suspiciously high scores on old items versus fresh equivalents.
- "Which benchmark matters for a RAG product?" — None directly. Use general benchmarks to shortlist, then your own retrieval and faithfulness eval on your documents.
- "Why do labs keep creating new benchmarks?" — Because old ones saturate or leak; harder, fresher sets (GPQA, Humanity's Last Exam, SWE-bench Pro, live-updated sets) restore the ability to separate models.