Course Content
LLM Evaluation
6 sections · 50 lessons
What is comparative evaluation, and how does it differ from A/B testing?
What you need to know
| Comparative evaluation | A/B test | |
|---|---|---|
| Population | Curated dataset | Live users, randomly assigned |
| Signal | Preference on outputs | Behaviour and business outcomes |
| Turnaround | Minutes to hours | Days to weeks |
| Risk to users | None | Real |
| Answers | "Which output is better?" | "Does it change what users do?" |
A/B test basics
- Choose one primary metric before starting (for example, escalation rate), plus guardrail metrics that must not get worse (latency, cost, complaints).
- Size it — compute how many users or sessions you need to detect the effect you care about.
- Randomise by user, not by request, so one person has a consistent experience.
- Check the split — if you planned 50/50 and got 52/48, something is wrong with assignment (a "sample ratio mismatch").
- Run for full weeks — to cover weekday and weekend patterns, and let novelty wear off.
- Analyse with confidence intervals and do not stop early the moment p dips below 0.05.
Worked numbers
The HR assistant tests a new retrieval setup. Control: 50,000 sessions, 12.0% escalated to a human HR agent. Treatment: 50,000 sessions, 11.4% escalated. The difference is −0.6 points, 95% interval −1.0 to −0.2 points, p ≈ 0.003. Small, but real — and at company scale it means about 300 fewer HR tickets per 50,000 sessions.
Why the two methods disagree
A variant that wins 60% of pairwise judgements can be neutral online because users don't notice, or negative because it is slower. And an online win can come from something a judge never sees, like shorter answers that users read faster.
Interleaving
For ranking-style outputs (search results, retrieved passages), interleaving mixes both variants' results in one list and sees which items users click. It needs far fewer users than a classic A/B test.
A real-life example
The e-commerce team has three new description prompts. Comparative evaluation on 300 products, with a pairwise judge in both orders plus a 50-item human check, ranks prompt C first (58% win rate against the current prompt, interval 52% to 64%). Prompts A and B are not distinguishable from the current prompt.
Only C goes to an A/B test on 20% of product pages for two weeks. Primary metric: add-to-cart rate. Guardrails: "not as described" returns and page load time. Add-to-cart shows no significant change; returns fall by a small but significant amount. The team ships C for its accuracy gain and records that the judge's preference did not translate into more sales.
Follow-up questions to expect
- "Why not A/B test every candidate?" — Online tests are slow, expose users to risk, and need large traffic; comparative evaluation filters candidates cheaply first.
- "What is peeking and why is it bad?" — Checking results repeatedly and stopping as soon as p is below 0.05 inflates false positives; fix the duration in advance or use sequential testing methods designed for it.
- "How do you A/B test an LLM feature with little traffic?" — Use interleaving where possible, longer tests, more sensitive metrics, or rely on offline evaluation plus a staged rollout with close monitoring.