Course Content
Evaluating and Testing GenAI Models
4 sections · 13 lessons
Pairwise Comparison, Ranking, and Annotation Guidelines
Two annotators rate the same summary on a 1-to-5 scale. One gives it a 4, the other a 2. In the debrief they discover they agree completely about the summary: it is well written, and it omits the deadline. One annotator treats the omission as a minor flaw in an otherwise good piece; the other treats it as disqualifying. Neither is wrong. The scale forced them to compress a rich judgement into a single number, and they compressed it differently.
Now show them two summaries side by side and ask which is better. Both immediately pick the one that keeps the deadline. Agreement: perfect. The judgement they can make reliably is comparative; the judgement the scale demanded was absolute, and absolute judgement requires an internal standard that no two people share.
This is a general fact about human measurement, well established long before language models existed, and it is the reason nearly every serious model-comparison system — arena leaderboards, preference data for training, internal release gates — is built on pairwise comparison rather than ratings.
Why pairwise beats rating
| Absolute rating (1–5) | Pairwise comparison | |
|---|---|---|
| Cognitive task | Map quality onto an abstract scale | Notice a difference between two concrete things |
| Requires shared standard | Yes — the main source of disagreement | No |
| Typical agreement | κ≈0.3–0.5 | κ≈0.6–0.8 on the same items |
| Drift over a session | Substantial; standards shift | Minimal; each comparison is self-contained |
| Sensitivity to small differences | Poor — both land on 4 | Good — annotator can still pick one |
| Gives an absolute quality level | Yes (unreliably) | No — only relative ordering |
| Comparisons needed for m systems | m passes | Up to (2m) pairs |
| Detects "both are terrible" | Yes | No — this is its main weakness |
That last row is the trade you are making and it is a real one. Pairwise comparison will happily report that model B beats model A 61% of the time when both are producing outputs no user would accept. It measures ordering, and ordering is silent about level.
Pairwise comparison tells you which model is better. It cannot tell you whether either is good enough. You need at least one absolute measurement somewhere in the system, or you will optimise your way to a confident ranking of failures.
The standard fix is a hybrid: pairwise for the primary comparison, plus a small absolute-quality check (a pass/fail acceptability gate, or a rating on a subsample) to anchor the scale.
Running the comparison without poisoning it
Pairwise judgement is easy to collect and easy to bias. Five design rules, each addressing a measured effect:
- Randomise left/right per item. Position bias is large and consistent. Measure yours: present the same pair to different annotators in both orders and compare.
- Blind everything that identifies a system. Not just names — response length, markdown style, characteristic openings. If one model always uses bullet points, annotators learn to identify it within twenty items.
- Offer a tie option, and define it precisely. Without one, annotators are forced to invent a preference, which is pure noise. But an undefined tie becomes the lazy default: specify "tie" as "I would be equally happy to send either to a customer", not "I cannot decide".
- Ask for a reason on every non-tie. A short structured reason (from a fixed list: more accurate, more complete, better format, more concise, safer) converts a preference into a diagnosis. Aggregate the reasons and you learn why you are winning, which the win rate alone never tells you.
- Cap the comparison at what fits on one screen. Comparing two 900-word outputs is not a comparison; it is two acts of reading separated by forgetting.
Quantifying position bias
Present each pair twice, in both orders, to different annotators. Suppose model A is preferred 68% of the time when shown first and 44% of the time when shown second.
Twelve points of position bias is enough to reverse most real model differences. If you only ever presented A first, you would have reported a 68% win rate for a model whose true preference is 56% — and 56% may not even be significant, as the next section shows.
Is a 54% win rate a win?
This is the arithmetic that decides most model-comparison arguments, so work it fully.
100 pairwise comparisons, A preferred in 54. Under the null hypothesis of no preference, p=0.5:
The 95% confidence interval is 54%±1.96×5.0%=[44.2%, 63.8%], comfortably straddling 50%. A 54–46 result on 100 comparisons is indistinguishable from a coin flip. If the same 4-point gap had come from an absolute-rating study, or from a benchmark accuracy comparison, the arithmetic would be almost identical — 100 examples simply cannot resolve four points.
How many comparisons would settle it? To detect a true 54% preference at 5% significance with 80% power:
Roughly 1,225 comparisons. Here is the same calculation across effect sizes, which is the table worth pinning above a desk:
| True win rate | Comparisons needed (80% power) | Realistic? |
|---|---|---|
| 52% | 4,900 | Only with crowd scale or an LLM judge |
| 54% | 1,225 | A serious study |
| 55% | 784 | Achievable |
| 60% | 196 | Easy |
| 65% | 87 | Trivial |
| 70% | 49 | Visible without statistics |
Read this table before commissioning a study, not after. If your realistic annotation budget is 200 comparisons, you can only detect differences of about 10 points or larger, and you should say so in advance rather than discovering it in the analysis.
Handling ties honestly
Ties break the arithmetic if you are careless. Suppose 100 comparisons yield 42 wins for A, 38 for B, and 20 ties. Two defensible conventions:
| Convention | Calculation | Result | Effective n |
|---|---|---|---|
| Drop ties | 42/(42+38) | 52.5% | 80 |
| Ties count half | (42+10)/100 | 52.0% | 100 |
Both are acceptable; silently switching between them is not, because they move the number and the effective sample size in opposite directions. State the convention, and always report the tie rate itself. A tie rate of 45% means your two systems are nearly indistinguishable to users — which is a finding in its own right, often more useful than the win rate.
From pairwise judgements to a ranking
With more than two systems, individual comparisons must be aggregated. Three standard approaches.
Simple win rate
Fraction of comparisons each system won, across all opponents. Easy, and wrong when systems face different opponents — a model that mostly played weak opponents looks strong. Fine for a round-robin where every pair is compared equally; misleading otherwise.
Elo ratings
Elo maintains a rating per system, updated after each comparison. The expected score for A against B:
and the update after observing outcome SA∈{0,0.5,1}:
Worked example. RA=1500, RB=1600, K=32.
A was expected to win 36% of the time. A wins:
A 100-point Elo gap corresponds to a 64% win rate; a 200-point gap to 76%. That mapping is the useful part — it lets you translate a leaderboard gap into the thing you care about.
Elo's weakness in evaluation: it is order-dependent. Feed the same comparisons in a different sequence and you get different final ratings, because early comparisons move ratings that later comparisons are measured against. For a fixed offline dataset this is unnecessary, and the Bradley-Terry fit below is strictly better.
Bradley-Terry: the right tool for offline data
The Bradley-Terry model says each system has a latent strength βi, and
Fit all β simultaneously by maximum likelihood over the whole comparison set. This is just logistic regression with one indicator per system, which means you get standard errors on each strength for free — and therefore confidence intervals on the ranking itself.
1import numpy as np2from scipy.optimize import minimize34def bradley_terry(comparisons, n_systems, ridge=1e-3):5 """comparisons: list of (winner_idx, loser_idx). Returns strengths and SEs."""6 def neg_ll(beta):7 ll = 0.08 for w, l in comparisons:9 ll += np.logaddexp(0.0, -(beta[w] - beta[l])) * -1.010 return -ll + ridge * np.sum(beta ** 2)1112 res = minimize(neg_ll, np.zeros(n_systems), method="L-BFGS-B")13 beta = res.x - res.x.mean() # identifiable up to a constant1415 # observed information -> standard errors16 H = np.zeros((n_systems, n_systems))17 for w, l in comparisons:18 p = 1.0 / (1.0 + np.exp(-(beta[w] - beta[l])))19 v = p * (1 - p)20 for a in (w, l):21 for b in (w, l):22 H[a, b] += v if a == b else -v23 se = np.sqrt(np.diag(np.linalg.pinv(H + ridge * np.eye(n_systems))))24 return beta, seConvert to the Elo scale with Ri=1500+ln10400βi≈1500+173.7βi if your organisation already speaks Elo.
Report the intervals. A leaderboard showing five systems at 1512, 1508, 1497, 1491 and 1483 with a standard error of 22 points on each is a leaderboard with one group in it, not a ranking of five things. Presenting it as an ordered list, without intervals, invents four distinctions that the data does not contain.
Detecting incoherent preferences
Human preferences need not be transitive. If A beats B, B beats C, and C beats A, you have a circular triad, and no ranking can represent it. Count them.
For m systems there are (3m) triads. Under completely random preferences, a quarter of triads are circular. With 5 systems, (35)=10 triads, so 2.5 circular triads is chance behaviour. Kendall's coefficient of consistency (for odd m) is:
With m=5 and d=2 observed circular triads:
ζ=1−125−524×2=1−12048=0.60A ζ near 1 means preferences are coherent and a ranking is meaningful. At 0.60 the systems are largely interchangeable, or your annotators are applying different criteria to different pairs — worth investigating before publishing an ordering.
Spending your comparison budget well
Exhaustive round-robin over m systems needs (2m) pairs, each compared many times: 8 systems is 28 pairs, and at 100 comparisons per pair that is 2,800 judgements. Better allocations exist.
Strategy How it works Best when Cost vs round-robin Round-robin Every pair, equally often Few systems; you need every pairwise number Baseline Anchor comparison Every system versus one fixed reference Tracking releases over time O(m) — far cheaper Swiss / adaptive pairing Pair systems with similar current ratings Many systems, ranking is the goal ~40–60% saving Active selection Choose the comparison that most reduces posterior uncertainty Large m, expensive annotation Up to 70% saving Stratified prompts Allocate comparisons across prompt categories deliberately Always — this is orthogonal to the above No extra cost The anchor strategy deserves emphasis for production work. Fix one reference system — the currently deployed model — and compare every candidate against it. You get a directly interpretable "win rate versus production" for each candidate at O(m) cost, and it is exactly the number the release decision needs.
Choose which prompts to compare on
Sampling prompts uniformly wastes budget on cases where both systems are obviously identical. Better: stratify.
- By difficulty. Include easy, medium and hard prompts in known proportions so the win rate is interpretable, and report the breakdown. A model that wins 70% on easy prompts and 45% on hard ones is a different product from one that is uniformly 58%.
- By disagreement. If a cheap automatic metric or an LLM judge scores the two outputs very differently, that pair is informative. If it scores them identically, it is probably a tie and a wasted judgement.
- By production frequency. Weight categories by real traffic, or your leaderboard will optimise for the tail.
- Keep a fixed frozen subset across all runs so that results are comparable over time.
One warning about difficulty-stratified sampling: if you oversample hard prompts, the resulting win rate no longer estimates the population win rate. Either re-weight when aggregating, or report the strata separately and never quote a pooled figure.
Guidelines for comparative annotation
Comparative guidelines differ from rating guidelines in one crucial way: they must specify a priority order, because the whole point is resolving cases where each output is better on a different dimension.
DECISION ORDER -- apply in sequence, stop at the first that decides1 SAFETY If exactly one output is unsafe, the other wins. Stop.2 ACCURACY If exactly one contains a claim contradicting the source, the other wins. Stop.3 COMPLETENESS If one answers the full question and the other omits a part the user explicitly asked for, the complete one wins.4 ACTIONABILITY If both are accurate and complete, prefer the one that lets the reader act without further questions.5 CONCISION If still equivalent, prefer the shorter.6 TIE Declare a tie only after all five. A tie means "I would be equally happy to send either to a customer."DO NOT consider: which model you think produced it; formatting preferenceabsent a stated requirement; length as a proxy for effort; confident tone.The explicit "do not consider" list is doing real work. Length bias in particular is large and persistent: annotators, and LLM judges even more so, prefer longer answers independent of content. Naming it in the guidelines reduces it; measuring it is better still.
Measure your annotators' length bias
Log the character count of the preferred and non-preferred output on every comparison. If the preferred output is longer 63% of the time in a set where the two systems have equal mean lengths, you have quantified a bias worth correcting for — by re-balancing the sample, by adding an explicit instruction, or by reporting the win rate conditioned on length.
Quality control on comparative data
| Check | Method | Action if it fails |
|---|---|---|
| Attention | Seed pairs where one output is obviously broken (truncated, off-topic) | Reject the annotator's batch below ~95% on these |
| Self-consistency | Re-present ~5% of pairs to the same annotator later, orders swapped | Below 80% self-agreement, retrain |
| Panel agreement | Overlap ~15% of pairs across annotators; compute kappa | Below 0.5, revisit guidelines |
| Position bias | Compare win rate by presentation position | Above ~8 points, enforce dual presentation and average |
| Speed | Time on task per comparison | Flag items decided faster than reading time |
| Degenerate patterns | Long runs of identical choices; always-left; never-tie | Manual audit of that annotator |
What this means when you run your next model comparison
Decide the sample size from the effect you need to detect, before collecting anything. If the business decision turns on a 3-point improvement, the table above says you need thousands of comparisons, and if you cannot afford thousands then the honest move is to change the decision criterion — for example to a gate ("does the new model regress on any category") rather than a ranking. Running an underpowered study and interpreting the point estimate anyway is the most common failure in this whole area, and it is entirely avoidable at the planning stage.
Collect a reason code with every preference, and treat the reason distribution as a primary output rather than a nicety. "We win 58% of comparisons" answers almost nothing. "We win 58%, and 71% of our wins are attributed to completeness while 64% of our losses are attributed to concision" tells you what to change tomorrow. That breakdown costs one extra click per annotation.
Finally, keep at least one absolute measurement alongside the pairwise one — a simple acceptability gate on a subsample is enough. Pairwise data is fast, cheap and reliable, and it will let you rank a series of models with steadily improving relative quality right past the point where all of them stopped being good enough. The comparative signal cannot see that; only an absolute one can.