Course Content
LLM Evaluation
6 sections · 50 lessons
How do you compare two prompts or models in a statistically sound way?
What you need to know
How much noise is in a 200-item eval?
A pass rate is a proportion, and proportions from small samples are uncertain. With 164 of 200 passing (82%), the 95% confidence interval (Wilson method) is 76.1% to 86.7%. A new prompt with 172 of 200 (86%) has an interval of 80.5% to 90.1%. The intervals overlap heavily.
If you treat the two runs as independent, the difference of 4 points has a 95% interval of about −3.2 to +11.2 points. Zero is inside: no evidence of improvement.
Pairing helps
Both prompts ran on the same 200 items. Most items give the same result for both; only discordant items — pass for one, fail for the other — carry information. Suppose the new prompt fixed 14 items and broke 6:
1import numpy as np23rng = np.random.default_rng(0)4# 200 eval items, 1 = pass. Same items, old prompt vs new prompt.5old = np.array([1]*158 + [1]*6 + [0]*14 + [0]*22) # 164 passes6new = np.array([1]*158 + [0]*6 + [1]*14 + [0]*22) # 172 passes78diff = new - old # per-item paired difference9boot = [rng.choice(diff, size=len(diff), replace=True).mean() for _ in range(10_000)]10lo, hi = np.percentile(boot, [2.5, 97.5])11print(f"old={old.mean():.1%} new={new.mean():.1%} "12 f"diff={diff.mean():+.1%} 95% CI [{lo:+.1%}, {hi:+.1%}]")old=82.0% new=86.0% diff=+4.0% 95% CI [+0.0%, +8.5%]The paired bootstrap resamples items (with replacement) 10,000 times and reads the middle 95% of the resulting differences. The interval is much narrower than the unpaired one, but its lower end touches zero. An exact McNemar test on 14 versus 6 gives p ≈ 0.12. Promising, not proven. Had it been 24 fixed versus 8 broken, p ≈ 0.007.
Choosing the test
| Data | Test |
|---|---|
| Paired pass/fail | McNemar (exact binomial on discordant pairs) |
| Paired scores (1 to 5, 0 to 1) | Paired bootstrap, or Wilcoxon signed-rank |
| Pairwise preferences | Sign test on wins versus losses, ties excluded |
| Any metric | Bootstrap over items |
Sizing the eval
For two independent groups near 50%, detecting a 5-point difference with 80% power at the 5% level needs about 1,560 items per group; a 10-point difference needs about 385. Pairing reduces this, because only discordant items matter. If you have 50 items, say you can only detect large effects.
Other rules
- Repeats — run each item 2 to 3 times to separate sampling noise from real change.
- Multiple comparisons — testing 10 prompts at p less than 0.05 makes a false "winner" likely; use Holm or Benjamini-Hochberg corrections, or confirm the winner on a fresh set.
- Practical significance — a real 1-point gain that doubles latency is not a win.
A real-life example
The bank's complaint classifier team tests a new prompt on their 400-item golden set. Accuracy goes from 86.0% to 88.5%. The product manager wants to ship. The engineer looks at the pairs: 22 complaints fixed, 12 broken. McNemar p ≈ 0.12. Worse, 5 of the 12 broken ones are fraud complaints now labelled upi.
They run both prompts on 600 more complaints from last month. Combined: 51 fixed versus 27 broken (p ≈ 0.009), but fraud recall drops from 0.88 to 0.83. The team ships only after adding two fraud examples to the new prompt and re-running, which restores fraud recall to 0.88 while keeping the overall gain.
Follow-up questions to expect
- "Why paired rather than two-sample tests?" — Some items are hard for every prompt; pairing cancels that shared difficulty, so less data is needed to see a real difference.
- "How do you handle non-deterministic outputs?" — Run repeats, average per item first, then bootstrap over items.
- "What if the judge is the metric?" — The judge adds its own noise and bias; keep it fixed across both arms and calibrated, and consider the uncertainty of the judge as well.