LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do you compare two prompts or models in a statistically sound way?


Same 200 items, old prompt vs new prompt15861422new: passnew: failold: passold: fail82% to 86%, but 14 fixed vs 6 broken gives McNemar p about 0.12.
Only the 20 discordant items carry evidence, and 14 against 6 is still within what chance produces.

What you need to know

How much noise is in a 200-item eval?

A pass rate is a proportion, and proportions from small samples are uncertain. With 164 of 200 passing (82%), the 95% confidence interval (Wilson method) is 76.1% to 86.7%. A new prompt with 172 of 200 (86%) has an interval of 80.5% to 90.1%. The intervals overlap heavily.

If you treat the two runs as independent, the difference of 4 points has a 95% interval of about −3.2 to +11.2 points. Zero is inside: no evidence of improvement.

Pairing helps

Both prompts ran on the same 200 items. Most items give the same result for both; only discordant items — pass for one, fail for the other — carry information. Suppose the new prompt fixed 14 items and broke 6:

Python
import numpy as nprng = np.random.default_rng(0)# 200 eval items, 1 = pass. Same items, old prompt vs new prompt.old = np.array([1]*158 + [1]*6 + [0]*14 + [0]*22)   # 164 passesnew = np.array([1]*158 + [0]*6 + [1]*14 + [0]*22)   # 172 passesdiff = new - old                                     # per-item paired differenceboot = [rng.choice(diff, size=len(diff), replace=True).mean() for _ in range(10_000)]lo, hi = np.percentile(boot, [2.5, 97.5])print(f"old={old.mean():.1%} new={new.mean():.1%} "      f"diff={diff.mean():+.1%}  95% CI [{lo:+.1%}, {hi:+.1%}]")
Text
old=82.0% new=86.0% diff=+4.0%  95% CI [+0.0%, +8.5%]

The paired bootstrap resamples items (with replacement) 10,000 times and reads the middle 95% of the resulting differences. The interval is much narrower than the unpaired one, but its lower end touches zero. An exact McNemar test on 14 versus 6 gives p ≈ 0.12. Promising, not proven. Had it been 24 fixed versus 8 broken, p ≈ 0.007.

Choosing the test

DataTest
Paired pass/failMcNemar (exact binomial on discordant pairs)
Paired scores (1 to 5, 0 to 1)Paired bootstrap, or Wilcoxon signed-rank
Pairwise preferencesSign test on wins versus losses, ties excluded
Any metricBootstrap over items

Sizing the eval

For two independent groups near 50%, detecting a 5-point difference with 80% power at the 5% level needs about 1,560 items per group; a 10-point difference needs about 385. Pairing reduces this, because only discordant items matter. If you have 50 items, say you can only detect large effects.

Other rules

  • Repeats — run each item 2 to 3 times to separate sampling noise from real change.
  • Multiple comparisons — testing 10 prompts at p less than 0.05 makes a false "winner" likely; use Holm or Benjamini-Hochberg corrections, or confirm the winner on a fresh set.
  • Practical significance — a real 1-point gain that doubles latency is not a win.

A real-life example

The bank's complaint classifier team tests a new prompt on their 400-item golden set. Accuracy goes from 86.0% to 88.5%. The product manager wants to ship. The engineer looks at the pairs: 22 complaints fixed, 12 broken. McNemar p ≈ 0.12. Worse, 5 of the 12 broken ones are fraud complaints now labelled upi.

They run both prompts on 600 more complaints from last month. Combined: 51 fixed versus 27 broken (p ≈ 0.009), but fraud recall drops from 0.88 to 0.83. The team ships only after adding two fraud examples to the new prompt and re-running, which restores fraud recall to 0.88 while keeping the overall gain.

Follow-up questions to expect

  • "Why paired rather than two-sample tests?" — Some items are hard for every prompt; pairing cancels that shared difficulty, so less data is needed to see a real difference.
  • "How do you handle non-deterministic outputs?" — Run repeats, average per item first, then bootstrap over items.
  • "What if the judge is the metric?" — The judge adds its own noise and bias; keep it fixed across both arms and calibrated, and consider the uncertainty of the judge as well.