Course Content
LLM Evaluation
6 sections · 50 lessons
How do leaderboards aggregate benchmark scores, and what are their limitations?
What you need to know
Static benchmark aggregation
Run each model through the same benchmarks with the same harness, then average. Because benchmarks have different guessing floors and ranges, boards often normalise each score (for example, rescale so random guessing is 0) before averaging. The choice of benchmarks and normalisation decides the ranking. Independent indexes, such as Artificial Analysis, also aggregate several evaluations into one number; read their method before quoting them.
Pairwise preference ratings
On the Elo scale, a 400-point gap means 10-to-1 odds:
P(A beats B) = 1 / (1 + 10 ** (-(rating_A - rating_B) / 400))100-point gap -> 0.64 30-point gap -> 0.54A 30-point lead means the higher model wins only about 54% of head-to-head votes. With vote noise, neighbouring models' confidence intervals often overlap.
Limitations
- Averages hide the profile. A model strong at code and weak at long documents can tie with its opposite.
- Hidden weighting. Which tasks are in, and how they are normalised, is a judgement call.
- Ties presented as ranks. Positions 3 to 7 may be statistically indistinguishable.
- Style bias. Voters prefer longer, well-formatted, confident answers. Arena introduced a "style control" view in 2024 that adjusts for length and markdown; rankings shift when it is applied.
- Wrong population. Arena prompts come from its users, not from your bank's customers. Cost, latency, context length, rate limits and data residency are absent.
- Gaming. Labs can test many private variants and publish the best, which inflates top scores.
A real-life example
A product manager asks why the HR assistant does not use "the number one model on the leaderboard". The engineer shows three facts. First, the top four models' confidence intervals overlap. Second, on the style-controlled view the order changes. Third, the team's own 200-question HR eval, run on the top three candidates, gives 86%, 85% and 84% correctness — a tie — but p95 latency of 9, 3 and 4 seconds, and a cost difference of about 5 times between the most and least expensive.
They choose the second model: statistically tied on quality, fastest, and available in an Indian cloud region, which the company's data policy requires. None of those deciding factors appear on a leaderboard.
Follow-up questions to expect
- "Why use Bradley-Terry instead of raw win rate?" — Models face different opponents; Bradley-Terry accounts for opponent strength and produces comparable scores with confidence intervals.
- "What is style control?" — A regression that adds features like response length and formatting to the Bradley-Terry fit, so the model's score reflects wins not explained by style.
- "How would you build an internal leaderboard?" — Same harness for all candidates, your own task evals, confidence intervals, and cost and latency columns next to quality.