LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What is comparative evaluation, and how does it differ from A/B testing?


What you need to know

Comparative evaluationA/B test
PopulationCurated datasetLive users, randomly assigned
SignalPreference on outputsBehaviour and business outcomes
TurnaroundMinutes to hoursDays to weeks
Risk to usersNoneReal
Answers"Which output is better?""Does it change what users do?"

A/B test basics

  1. Choose one primary metric before starting (for example, escalation rate), plus guardrail metrics that must not get worse (latency, cost, complaints).
  2. Size it — compute how many users or sessions you need to detect the effect you care about.
  3. Randomise by user, not by request, so one person has a consistent experience.
  4. Check the split — if you planned 50/50 and got 52/48, something is wrong with assignment (a "sample ratio mismatch").
  5. Run for full weeks — to cover weekday and weekend patterns, and let novelty wear off.
  6. Analyse with confidence intervals and do not stop early the moment p dips below 0.05.

Worked numbers

The HR assistant tests a new retrieval setup. Control: 50,000 sessions, 12.0% escalated to a human HR agent. Treatment: 50,000 sessions, 11.4% escalated. The difference is −0.6 points, 95% interval −1.0 to −0.2 points, p ≈ 0.003. Small, but real — and at company scale it means about 300 fewer HR tickets per 50,000 sessions.

Why the two methods disagree

A variant that wins 60% of pairwise judgements can be neutral online because users don't notice, or negative because it is slower. And an online win can come from something a judge never sees, like shorter answers that users read faster.

Interleaving

For ranking-style outputs (search results, retrieved passages), interleaving mixes both variants' results in one list and sees which items users click. It needs far fewer users than a classic A/B test.

A real-life example

The e-commerce team has three new description prompts. Comparative evaluation on 300 products, with a pairwise judge in both orders plus a 50-item human check, ranks prompt C first (58% win rate against the current prompt, interval 52% to 64%). Prompts A and B are not distinguishable from the current prompt.

Only C goes to an A/B test on 20% of product pages for two weeks. Primary metric: add-to-cart rate. Guardrails: "not as described" returns and page load time. Add-to-cart shows no significant change; returns fall by a small but significant amount. The team ships C for its accuracy gain and records that the judge's preference did not translate into more sales.

Follow-up questions to expect

  • "Why not A/B test every candidate?" — Online tests are slow, expose users to risk, and need large traffic; comparative evaluation filters candidates cheaply first.
  • "What is peeking and why is it bad?" — Checking results repeatedly and stopping as soon as p is below 0.05 inflates false positives; fix the duration in advance or use sequential testing methods designed for it.
  • "How do you A/B test an LLM feature with little traffic?" — Use interleaving where possible, longer tests, more sensitive metrics, or rely on offline evaluation plus a staged rollout with close monitoring.