LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do leaderboards aggregate benchmark scores, and what are their limitations?


Chance the higher-rated model wins, by rating gap0.540.640.760.91012330 points100 points200 points400 pointsElo scale: P(win) = 1 / (1 + 10 to the power of minus gap over 400).
A 30-point lead means winning only 54 of 100 head-to-heads, which is why neighbouring ranks are often ties.

What you need to know

Static benchmark aggregation

Run each model through the same benchmarks with the same harness, then average. Because benchmarks have different guessing floors and ranges, boards often normalise each score (for example, rescale so random guessing is 0) before averaging. The choice of benchmarks and normalisation decides the ranking. Independent indexes, such as Artificial Analysis, also aggregate several evaluations into one number; read their method before quoting them.

Pairwise preference ratings

On the Elo scale, a 400-point gap means 10-to-1 odds:

Text
P(A beats B) = 1 / (1 + 10 ** (-(rating_A - rating_B) / 400))100-point gap -> 0.64     30-point gap -> 0.54

A 30-point lead means the higher model wins only about 54% of head-to-head votes. With vote noise, neighbouring models' confidence intervals often overlap.

Limitations

  • Averages hide the profile. A model strong at code and weak at long documents can tie with its opposite.
  • Hidden weighting. Which tasks are in, and how they are normalised, is a judgement call.
  • Ties presented as ranks. Positions 3 to 7 may be statistically indistinguishable.
  • Style bias. Voters prefer longer, well-formatted, confident answers. Arena introduced a "style control" view in 2024 that adjusts for length and markdown; rankings shift when it is applied.
  • Wrong population. Arena prompts come from its users, not from your bank's customers. Cost, latency, context length, rate limits and data residency are absent.
  • Gaming. Labs can test many private variants and publish the best, which inflates top scores.

A real-life example

A product manager asks why the HR assistant does not use "the number one model on the leaderboard". The engineer shows three facts. First, the top four models' confidence intervals overlap. Second, on the style-controlled view the order changes. Third, the team's own 200-question HR eval, run on the top three candidates, gives 86%, 85% and 84% correctness — a tie — but p95 latency of 9, 3 and 4 seconds, and a cost difference of about 5 times between the most and least expensive.

They choose the second model: statistically tied on quality, fastest, and available in an Indian cloud region, which the company's data policy requires. None of those deciding factors appear on a leaderboard.

Follow-up questions to expect

  • "Why use Bradley-Terry instead of raw win rate?" — Models face different opponents; Bradley-Terry accounts for opponent strength and produces comparable scores with confidence intervals.
  • "What is style control?" — A regression that adds features like response length and formatting to the Bradley-Terry fit, so the model's score reflects wins not explained by style.
  • "How would you build an internal leaderboard?" — Same harness for all candidates, your own task evals, confidence intervals, and cost and latency columns next to quality.