Evaluating and Testing GenAI Models

Pairwise Comparison, Ranking, and Annotation Guidelines


Two annotators rate the same summary on a 1-to-5 scale. One gives it a 4, the other a 2. In the debrief they discover they agree completely about the summary: it is well written, and it omits the deadline. One annotator treats the omission as a minor flaw in an otherwise good piece; the other treats it as disqualifying. Neither is wrong. The scale forced them to compress a rich judgement into a single number, and they compressed it differently.

Now show them two summaries side by side and ask which is better. Both immediately pick the one that keeps the deadline. Agreement: perfect. The judgement they can make reliably is comparative; the judgement the scale demanded was absolute, and absolute judgement requires an internal standard that no two people share.

This is a general fact about human measurement, well established long before language models existed, and it is the reason nearly every serious model-comparison system — arena leaderboards, preference data for training, internal release gates — is built on pairwise comparison rather than ratings.

Rating on a scale versus choosing between twoAbsolute 1-to-5 rating• Each rater invents their own scale• Severity of a flaw is a private weight• Drift over a long session• Two raters who agree still score 4 and 2Pairwise A-or-B• Only the ordering has to be shared• The comparison fixes the context• Position bias, and it is measurable• Aggregate withBradley-Terry, not averages
The two annotators agreed completely about the summary and disagreed only about the scale — pairwise never asks them that question.

Why pairwise beats rating

Absolute rating (1–5)Pairwise comparison
Cognitive taskMap quality onto an abstract scaleNotice a difference between two concrete things
Requires shared standardYes — the main source of disagreementNo
Typical agreementκ≈0.3\kappa \approx 0.3–0.50.5κ≈0.6\kappa \approx 0.6–0.80.8 on the same items
Drift over a sessionSubstantial; standards shiftMinimal; each comparison is self-contained
Sensitivity to small differencesPoor — both land on 4Good — annotator can still pick one
Gives an absolute quality levelYes (unreliably)No — only relative ordering
Comparisons needed for mm systemsmm passesUp to (m2)\binom{m}{2} pairs
Detects "both are terrible"YesNo — this is its main weakness

That last row is the trade you are making and it is a real one. Pairwise comparison will happily report that model B beats model A 61% of the time when both are producing outputs no user would accept. It measures ordering, and ordering is silent about level.

Pairwise comparison tells you which model is better. It cannot tell you whether either is good enough. You need at least one absolute measurement somewhere in the system, or you will optimise your way to a confident ranking of failures.

The standard fix is a hybrid: pairwise for the primary comparison, plus a small absolute-quality check (a pass/fail acceptability gate, or a rating on a subsample) to anchor the scale.

Running the comparison without poisoning it

Pairwise judgement is easy to collect and easy to bias. Five design rules, each addressing a measured effect:

  • Randomise left/right per item. Position bias is large and consistent. Measure yours: present the same pair to different annotators in both orders and compare.
  • Blind everything that identifies a system. Not just names — response length, markdown style, characteristic openings. If one model always uses bullet points, annotators learn to identify it within twenty items.
  • Offer a tie option, and define it precisely. Without one, annotators are forced to invent a preference, which is pure noise. But an undefined tie becomes the lazy default: specify "tie" as "I would be equally happy to send either to a customer", not "I cannot decide".
  • Ask for a reason on every non-tie. A short structured reason (from a fixed list: more accurate, more complete, better format, more concise, safer) converts a preference into a diagnosis. Aggregate the reasons and you learn why you are winning, which the win rate alone never tells you.
  • Cap the comparison at what fits on one screen. Comparing two 900-word outputs is not a comparison; it is two acts of reading separated by forgetting.

Quantifying position bias

Present each pair twice, in both orders, to different annotators. Suppose model A is preferred 68% of the time when shown first and 44% of the time when shown second.

Debiased preference for A=68%+44%2=56%Position bias=68%−44%2=12 points\text{Debiased preference for A} = \frac{68\% + 44\%}{2} = 56\% \qquad \text{Position bias} = \frac{68\% - 44\%}{2} = 12 \text{ points}

Twelve points of position bias is enough to reverse most real model differences. If you only ever presented A first, you would have reported a 68% win rate for a model whose true preference is 56% — and 56% may not even be significant, as the next section shows.

Is a 54% win rate a win?

This is the arithmetic that decides most model-comparison arguments, so work it fully.

100 pairwise comparisons, A preferred in 54. Under the null hypothesis of no preference, p=0.5p = 0.5:

SE=0.5×0.5100=0.0025=0.05    (5.0 percentage points)SE = \sqrt{\frac{0.5 \times 0.5}{100}} = \sqrt{0.0025} = 0.05 \;\;(5.0 \text{ percentage points})

z=0.54−0.500.05=0.80p≈0.42z = \frac{0.54 - 0.50}{0.05} = 0.80 \qquad p \approx 0.42

The 95% confidence interval is 54%±1.96×5.0%=[44.2%, 63.8%]54\% \pm 1.96 \times 5.0\% = [44.2\%,\ 63.8\%], comfortably straddling 50%. A 54–46 result on 100 comparisons is indistinguishable from a coin flip. If the same 4-point gap had come from an absolute-rating study, or from a benchmark accuracy comparison, the arithmetic would be almost identical — 100 examples simply cannot resolve four points.

How many comparisons would settle it? To detect a true 54% preference at 5% significance with 80% power:

n=(zα/2+zβ)2 p(1−p)δ2=(1.96+0.84)2×0.250.042=7.84×0.250.0016=1.960.0016=1225n = \frac{(z_{\alpha/2} + z_{\beta})^2 \, p(1-p)}{\delta^2} = \frac{(1.96 + 0.84)^2 \times 0.25}{0.04^2} = \frac{7.84 \times 0.25}{0.0016} = \frac{1.96}{0.0016} = 1225

Roughly 1,225 comparisons. Here is the same calculation across effect sizes, which is the table worth pinning above a desk:

True win rateComparisons needed (80% power)Realistic?
52%4,900Only with crowd scale or an LLM judge
54%1,225A serious study
55%784Achievable
60%196Easy
65%87Trivial
70%49Visible without statistics

Read this table before commissioning a study, not after. If your realistic annotation budget is 200 comparisons, you can only detect differences of about 10 points or larger, and you should say so in advance rather than discovering it in the analysis.

Handling ties honestly

Ties break the arithmetic if you are careless. Suppose 100 comparisons yield 42 wins for A, 38 for B, and 20 ties. Two defensible conventions:

ConventionCalculationResultEffective nn
Drop ties42/(42+38)42 / (42+38)52.5%80
Ties count half(42+10)/100(42 + 10) / 10052.0%100

Both are acceptable; silently switching between them is not, because they move the number and the effective sample size in opposite directions. State the convention, and always report the tie rate itself. A tie rate of 45% means your two systems are nearly indistinguishable to users — which is a finding in its own right, often more useful than the win rate.

From pairwise judgements to a ranking

With more than two systems, individual comparisons must be aggregated. Three standard approaches.

Simple win rate

Fraction of comparisons each system won, across all opponents. Easy, and wrong when systems face different opponents — a model that mostly played weak opponents looks strong. Fine for a round-robin where every pair is compared equally; misleading otherwise.

Elo ratings

Elo maintains a rating per system, updated after each comparison. The expected score for A against B:

EA=11+10(RB−RA)/400E_A = \frac{1}{1 + 10^{(R_B - R_A)/400}}

and the update after observing outcome SA∈{0,0.5,1}S_A \in \{0, 0.5, 1\}:

RA′=RA+K(SA−EA)R_A' = R_A + K(S_A - E_A)

Worked example. RA=1500R_A = 1500, RB=1600R_B = 1600, K=32K = 32.

EA=11+10(1600−1500)/400=11+100.25=11+1.7783=12.7783=0.360E_A = \frac{1}{1 + 10^{(1600-1500)/400}} = \frac{1}{1 + 10^{0.25}} = \frac{1}{1 + 1.7783} = \frac{1}{2.7783} = 0.360

A was expected to win 36% of the time. A wins:

RA′=1500+32(1−0.360)=1500+32×0.640=1500+20.5=1520.5R_A' = 1500 + 32(1 - 0.360) = 1500 + 32 \times 0.640 = 1500 + 20.5 = 1520.5
RB′=1600+32(0−0.640)=1600−20.5=1579.5R_B' = 1600 + 32(0 - 0.640) = 1600 - 20.5 = 1579.5

A 100-point Elo gap corresponds to a 64% win rate; a 200-point gap to 76%. That mapping is the useful part — it lets you translate a leaderboard gap into the thing you care about.

Elo's weakness in evaluation: it is order-dependent. Feed the same comparisons in a different sequence and you get different final ratings, because early comparisons move ratings that later comparisons are measured against. For a fixed offline dataset this is unnecessary, and the Bradley-Terry fit below is strictly better.

Bradley-Terry: the right tool for offline data

The Bradley-Terry model says each system has a latent strength βi\beta_i, and

P(i beats j)=11+e−(βi−βj)P(i \text{ beats } j) = \frac{1}{1 + e^{-(\beta_i - \beta_j)}}

Fit all β\beta simultaneously by maximum likelihood over the whole comparison set. This is just logistic regression with one indicator per system, which means you get standard errors on each strength for free — and therefore confidence intervals on the ranking itself.

Python
import numpy as npfrom scipy.optimize import minimizedef bradley_terry(comparisons, n_systems, ridge=1e-3):    """comparisons: list of (winner_idx, loser_idx). Returns strengths and SEs."""    def neg_ll(beta):        ll = 0.0        for w, l in comparisons:            ll += np.logaddexp(0.0, -(beta[w] - beta[l])) * -1.0        return -ll + ridge * np.sum(beta ** 2)    res = minimize(neg_ll, np.zeros(n_systems), method="L-BFGS-B")    beta = res.x - res.x.mean()                      # identifiable up to a constant    # observed information -> standard errors    H = np.zeros((n_systems, n_systems))    for w, l in comparisons:        p = 1.0 / (1.0 + np.exp(-(beta[w] - beta[l])))        v = p * (1 - p)        for a in (w, l):            for b in (w, l):                H[a, b] += v if a == b else -v    se = np.sqrt(np.diag(np.linalg.pinv(H + ridge * np.eye(n_systems))))    return beta, se

Convert to the Elo scale with Ri=1500+400ln⁡10βi≈1500+173.7 βiR_i = 1500 + \frac{400}{\ln 10}\beta_i \approx 1500 + 173.7\,\beta_i if your organisation already speaks Elo.

Report the intervals. A leaderboard showing five systems at 1512, 1508, 1497, 1491 and 1483 with a standard error of 22 points on each is a leaderboard with one group in it, not a ranking of five things. Presenting it as an ordered list, without intervals, invents four distinctions that the data does not contain.

Detecting incoherent preferences

Human preferences need not be transitive. If A beats B, B beats C, and C beats A, you have a circular triad, and no ranking can represent it. Count them.

For mm systems there are (m3)\binom{m}{3} triads. Under completely random preferences, a quarter of triads are circular. With 5 systems, (53)=10\binom{5}{3} = 10 triads, so 2.5 circular triads is chance behaviour. Kendall's coefficient of consistency (for odd mm) is:

ζ=1−24dm3−m\zeta = 1 - \frac{24d}{m^3 - m}

With m=5m = 5 and d=2d = 2 observed circular triads:

ζ=1−24×2125−5=1−48120=0.60\zeta = 1 - \frac{24 \times 2}{125 - 5} = 1 - \frac{48}{120} = 0.60

A ζ\zeta near 1 means preferences are coherent and a ranking is meaningful. At 0.60 the systems are largely interchangeable, or your annotators are applying different criteria to different pairs — worth investigating before publishing an ordering.

Spending your comparison budget well

Exhaustive round-robin over mm systems needs (m2)\binom{m}{2} pairs, each compared many times: 8 systems is 28 pairs, and at 100 comparisons per pair that is 2,800 judgements. Better allocations exist.

StrategyHow it worksBest whenCost vs round-robin
Round-robinEvery pair, equally oftenFew systems; you need every pairwise numberBaseline
Anchor comparisonEvery system versus one fixed referenceTracking releases over timeO(m)O(m) — far cheaper
Swiss / adaptive pairingPair systems with similar current ratingsMany systems, ranking is the goal~40–60% saving
Active selectionChoose the comparison that most reduces posterior uncertaintyLarge mm, expensive annotationUp to 70% saving
Stratified promptsAllocate comparisons across prompt categories deliberatelyAlways — this is orthogonal to the aboveNo extra cost

The anchor strategy deserves emphasis for production work. Fix one reference system — the currently deployed model — and compare every candidate against it. You get a directly interpretable "win rate versus production" for each candidate at O(m)O(m) cost, and it is exactly the number the release decision needs.

Choose which prompts to compare on

Sampling prompts uniformly wastes budget on cases where both systems are obviously identical. Better: stratify.

  • By difficulty. Include easy, medium and hard prompts in known proportions so the win rate is interpretable, and report the breakdown. A model that wins 70% on easy prompts and 45% on hard ones is a different product from one that is uniformly 58%.
  • By disagreement. If a cheap automatic metric or an LLM judge scores the two outputs very differently, that pair is informative. If it scores them identically, it is probably a tie and a wasted judgement.
  • By production frequency. Weight categories by real traffic, or your leaderboard will optimise for the tail.
  • Keep a fixed frozen subset across all runs so that results are comparable over time.

One warning about difficulty-stratified sampling: if you oversample hard prompts, the resulting win rate no longer estimates the population win rate. Either re-weight when aggregating, or report the strata separately and never quote a pooled figure.

Guidelines for comparative annotation

Comparative guidelines differ from rating guidelines in one crucial way: they must specify a priority order, because the whole point is resolving cases where each output is better on a different dimension.

Text
DECISION ORDER -- apply in sequence, stop at the first that decides1  SAFETY       If exactly one output is unsafe, the other wins. Stop.2  ACCURACY     If exactly one contains a claim contradicting the source,                the other wins. Stop.3  COMPLETENESS If one answers the full question and the other omits a                part the user explicitly asked for, the complete one wins.4  ACTIONABILITY If both are accurate and complete, prefer the one that                lets the reader act without further questions.5  CONCISION    If still equivalent, prefer the shorter.6  TIE          Declare a tie only after all five. A tie means "I would be                equally happy to send either to a customer."DO NOT consider: which model you think produced it; formatting preferenceabsent a stated requirement; length as a proxy for effort; confident tone.

The explicit "do not consider" list is doing real work. Length bias in particular is large and persistent: annotators, and LLM judges even more so, prefer longer answers independent of content. Naming it in the guidelines reduces it; measuring it is better still.

Measure your annotators' length bias

Log the character count of the preferred and non-preferred output on every comparison. If the preferred output is longer 63% of the time in a set where the two systems have equal mean lengths, you have quantified a bias worth correcting for — by re-balancing the sample, by adding an explicit instruction, or by reporting the win rate conditioned on length.

Quality control on comparative data

CheckMethodAction if it fails
AttentionSeed pairs where one output is obviously broken (truncated, off-topic)Reject the annotator's batch below ~95% on these
Self-consistencyRe-present ~5% of pairs to the same annotator later, orders swappedBelow 80% self-agreement, retrain
Panel agreementOverlap ~15% of pairs across annotators; compute kappaBelow 0.5, revisit guidelines
Position biasCompare win rate by presentation positionAbove ~8 points, enforce dual presentation and average
SpeedTime on task per comparisonFlag items decided faster than reading time
Degenerate patternsLong runs of identical choices; always-left; never-tieManual audit of that annotator

What this means when you run your next model comparison

Decide the sample size from the effect you need to detect, before collecting anything. If the business decision turns on a 3-point improvement, the table above says you need thousands of comparisons, and if you cannot afford thousands then the honest move is to change the decision criterion — for example to a gate ("does the new model regress on any category") rather than a ranking. Running an underpowered study and interpreting the point estimate anyway is the most common failure in this whole area, and it is entirely avoidable at the planning stage.

Collect a reason code with every preference, and treat the reason distribution as a primary output rather than a nicety. "We win 58% of comparisons" answers almost nothing. "We win 58%, and 71% of our wins are attributed to completeness while 64% of our losses are attributed to concision" tells you what to change tomorrow. That breakdown costs one extra click per annotation.

Finally, keep at least one absolute measurement alongside the pairwise one — a simple acceptability gate on a subsample is enough. Pairwise data is fast, cheap and reliable, and it will let you rank a series of models with steadily improving relative quality right past the point where all of them stopped being good enough. The comparative signal cannot see that; only an absolute one can.