Evaluating and Testing GenAI Models

Designing Human Evaluation Pipelines


A team spends 9,000 dollars having three annotators rate 5,000 model outputs on a 1-to-5 quality scale. The results come back: model A averages 3.71, model B averages 3.64. Leadership asks whether A is better. The data scientist checks agreement between annotators and finds Fleiss' kappa of 0.33.

That kappa means the annotators were, to a substantial degree, not measuring the same thing. The 0.07 difference between the models is smaller than the disagreement between two people looking at the same output. Nine thousand dollars bought a number that cannot answer the question it was commissioned to answer, and the failure happened before the first annotation was collected — in the design.

Human evaluation is the only method that can measure the things that matter most about generative systems. It is also an instrument, and like any instrument it has a noise floor, a calibration procedure, and a set of ways to be used wrong. This lesson is about building it so the number at the end means something.

Where 9,000 dollars of annotation wentDefine the construct being measuredWrite anchors, not adjectivesQualify annotators on gold itemsCollect blind, with gold seeded inCheck kappa before reading any mean
At Fleiss' kappa 0.33 the raters are barely measuring the same thing, so 3.71 against 3.64 is a difference between two noises.

The pipeline, and where each stage fails

Text
[1] Define construct   ->  What exactly are we measuring? What does bad look like?[2] Design rubric      ->  Dimensions, scales, anchors, gates[3] Write guidelines   ->  Edge cases, worked examples, disagreement resolution[4] Pilot              ->  30-50 items, 2+ raters, MEASURE AGREEMENT[5] Fix rubric         ->  Iterate 2-4 until agreement is usable. Do not skip.[6] Recruit + train    ->  Qualification task with a pass bar[7] Sample items       ->  Stratified, sized from a power calculation[8] Annotate           ->  Blind, randomised, with gold items seeded in[9] Monitor            ->  Per-annotator agreement, drift, throughput, gold accuracy[10] Adjudicate        ->  Resolve disagreements by a stated rule[11] Analyse           ->  Effect sizes and intervals, not just means

Stages 4 and 5 are the ones teams skip under deadline pressure, and skipping them is what produced the opening failure. Piloting costs about half a day. Discovering a broken rubric after collecting 15,000 judgements costs the entire budget.

Defining the construct before the rubric

The question "rate this output 1–5 for quality" fails because quality is not one thing. Before writing any scale, answer three questions in writing:

  1. What decision will this number inform? Ship or not ship? Which of two prompts? Whether to escalate to human review? Different decisions need different measurements, and a rubric built for none of them in particular serves none of them.
  2. What does a bad output look like? Collect 50 real bad outputs. Write one sentence per output about what is wrong. Cluster the sentences. Those clusters are your dimensions — derived from evidence, not from a meeting.
  3. Which dimensions are gates and which are scores? A gate is a floor below which nothing else matters (factual error, safety violation, missing required disclosure). A score ranks the outputs that clear the gates. Mixing them into one weighted average lets an elegant, dangerous output outrank a plain, correct one.

If you cannot describe, in one sentence, the decision your rating will change, you are not ready to write a rubric.

Anchors are the rubric

A scale without anchors produces central-tendency bias: almost every rating lands on 3 or 4, variance collapses, and no comparison is possible. Every scale point needs a description specific enough that two people reading it reach the same conclusion.

ScoreWeak anchor (produces kappa ~0.3)Strong anchor (produces kappa ~0.7)
5ExcellentAnswers the full question, every claim traceable to the source, no unnecessary content, correct format. A senior agent would send this unedited.
4GoodAnswers the question correctly; one stylistic edit needed, or one non-essential detail missing. Sendable after a 10-second fix.
3AcceptableAnswers the main question but omits a secondary part, or includes one correct-but-irrelevant paragraph. Needs a minute of editing.
2PoorAddresses the topic but misses the actual question, or requires rewriting more than half the text.
1Very poorWrong question answered, or contains a claim contradicting the source, or unusable format. Faster to write from scratch.

Notice what the strong anchors have in common: they refer to observable properties and actions ("sendable after a 10-second fix", "faster to write from scratch"), not to internal impressions. An annotator can check whether a claim is traceable. They cannot check whether something is excellent.

Guidelines that actually reduce disagreement

A guideline document that only restates the rubric is decoration. The parts that move agreement are the parts about hard cases.

SectionWhat it containsWhy it matters
Task framingWho reads these outputs and what they do nextAnnotators calibrate to a real user, not an imagined one
Dimension definitionsOne paragraph each, with the question to ask yourselfPrevents dimension bleed (rating fluency under accuracy)
Anchored scalesAs above, with two real examples per pointThe single biggest lever on agreement
Edge-case rulings"If the output refuses, score X." "If it is correct but off-topic, score Y."This is where most disagreement lives
Explicit non-criteria"Do not penalise British vs American spelling." "Do not reward length."Removes systematic per-annotator biases
Worked examples8–12 fully annotated items with reasoningDoubles as the training set
Escalation ruleWhat to do when genuinely unsure; a "flag" optionStops guessing, which is pure noise

The edge-case section should grow during the pilot. Every disagreement in the pilot is a missing ruling. Write the ruling, do not blame the annotator — disagreement is nearly always a specification defect.

Give annotators a "cannot decide" option and count how often it is used. An item flagged as undecidable is information; the same item guessed on is noise dressed as data. If more than about 8% of items get flagged, the rubric does not fit the output distribution.

Measuring agreement, properly

Raw percentage agreement is the wrong statistic because it credits agreement that would happen by chance. Which chance-corrected measure you need depends on your design.

MeasureUse whenHandles
Cohen's κ\kappaExactly 2 raters, categorical labels, same raters throughoutNominal categories only
Weighted κ\kappa2 raters, ordinal scaleTreats 4-vs-5 as a smaller error than 1-vs-5
Fleiss' κ\kappa3+ raters, categorical, raters may differ per itemVariable rater panels
Krippendorff's α\alphaAny number of raters, any scale type, missing dataThe most general; use it if unsure
ICCContinuous or ordinal ratings, and you care about absolute valuesRater bias as well as inconsistency
Spearman ρ\rhoYou only care about the ranking of systemsSystematic offsets between raters

Fleiss' kappa, worked end to end

Ten outputs, three raters each, three categories (Good / Fair / Poor). The counts per item:

ItemGoodFairPoor∑jnij2\sum_j n_{ij}^2PiP_i
130091.000
221050.333
302150.333
400391.000
512050.333
630091.000
701250.333
821050.333
903091.000
1011130.000
Total121175.667

Each PiP_i is the proportion of rater pairs on item ii that agree, computed as

Pi=1n(n−1)(∑jnij2−n)=∑jnij2−36P_i = \frac{1}{n(n-1)}\left(\sum_j n_{ij}^2 - n\right) = \frac{\sum_j n_{ij}^2 - 3}{6}

Mean observed agreement:

Pˉ=5.66710=0.567\bar{P} = \frac{5.667}{10} = 0.567

Category marginals over all 10×3=3010 \times 3 = 30 ratings:

pGood=1230=0.400pFair=1130=0.367pPoor=730=0.233p_{\text{Good}} = \frac{12}{30} = 0.400 \quad p_{\text{Fair}} = \frac{11}{30} = 0.367 \quad p_{\text{Poor}} = \frac{7}{30} = 0.233

Pe=0.4002+0.3672+0.2332=0.1600+0.1344+0.0544=0.3489P_e = 0.400^2 + 0.367^2 + 0.233^2 = 0.1600 + 0.1344 + 0.0544 = 0.3489
κ=Pˉ−Pe1−Pe=0.567−0.3491−0.349=0.2180.651=0.334\kappa = \frac{\bar{P} - P_e}{1 - P_e} = \frac{0.567 - 0.349}{1 - 0.349} = \frac{0.218}{0.651} = 0.334

A kappa of 0.33. Note that raw pairwise agreement was 56.7%, which sounds like the raters mostly concurred. They did not — over a third of that agreement is what three people picking from three categories would produce by chance.

κ\kappaConventional readingWhat to do
< 0.20SlightThe rubric is broken. Do not collect data.
0.21 – 0.40FairRewrite anchors, add edge-case rulings, re-pilot.
0.41 – 0.60ModerateUsable for coarse comparisons only. Aggregate 3+ raters.
0.61 – 0.80SubstantialGood. This is the realistic target for subjective dimensions.
> 0.80Almost perfectExcellent — or the task is so easy it is not worth human time.
Python
import numpy as npfrom itertools import combinationsdef fleiss_kappa(counts: np.ndarray) -> float:    """counts[i][j] = number of raters assigning item i to category j."""    N, _ = counts.shape    n = counts[0].sum()                       # raters per item (constant)    P_i = (np.sum(counts ** 2, axis=1) - n) / (n * (n - 1))    P_bar = P_i.mean()    p_j = counts.sum(axis=0) / (N * n)    P_e = np.sum(p_j ** 2)    return float((P_bar - P_e) / (1 - P_e))def pairwise_agreement(ratings_by_rater: dict) -> float:    """Raw agreement across all rater pairs -- report ALONGSIDE kappa, never instead."""    raters = list(ratings_by_rater)    agree = total = 0    for a, b in combinations(raters, 2):        ra, rb = ratings_by_rater[a], ratings_by_rater[b]        shared = set(ra) & set(rb)        agree += sum(1 for k in shared if ra[k] == rb[k])        total += len(shared)    return agree / total if total else float("nan")

Do not average away a low kappa

There is a real temptation, on seeing kappa 0.33, to say "but with three raters the mean is more reliable". It is, somewhat — averaging reduces random error by roughly n\sqrt{n}. But low kappa is usually not random error; it is systematic disagreement about what is being measured. Averaging a rater who is scoring factuality with one who is scoring fluency produces a stable number that measures neither. Fix the construct, then average.

Annotators: recruiting, training, and keeping them honest

Match expertise to the construct

Judgement requiredWho can make itCost per hour (USD)
Is this fluent English? Is the format right?General crowd workers10 – 20
Does this summary match the source?Trained crowd workers with a good rubric18 – 30
Is this the answer a support agent would give?Actual support agents30 – 60
Is this clinically safe? Legally sound?Licensed professionals150 – 500

Using cheaper annotators than the construct requires does not give you a noisier version of the right answer. It gives you a precise measurement of a different construct — usually fluency, because that is what a non-expert can perceive.

Qualification, not just training

Every annotator completes a 20-item qualification set with known answers before touching real data, and must hit a stated bar. Then review their first 30 real items individually, giving feedback on disagreements. This is expensive and it is cheaper than discovering after 3,000 items that one annotator misread a dimension throughout.

Gold items, sized correctly

Seed known-answer items into the live stream at roughly 8–10% and track each annotator's accuracy on them. But be honest about the statistical power of that check. With 200 items per annotator and 10% gold, you have 20 gold items. If the acceptable accuracy is 90%:

SE=0.90×0.1020=0.0045=0.067SE = \sqrt{\frac{0.90 \times 0.10}{20}} = \sqrt{0.0045} = 0.067

An annotator scoring 14/20 (70%) is (0.90−0.70)/0.067=3.0(0.90 - 0.70)/0.067 = 3.0 standard errors below the bar — detectable. An annotator at 80% is only 1.5 standard errors below, well inside normal variation. Gold items reliably catch gross failure and cannot detect mild degradation, so pair them with agreement monitoring against the panel, which is far more sensitive.

Sizing the study before you spend the money

This is the calculation the opening team never did. On a 1–5 scale, suppose the within-item difference between two systems has standard deviation σd=0.9\sigma_d = 0.9 and you want to detect a mean difference of δ=0.3\delta = 0.3 at 5% significance with 80% power.

Paired design — the same annotator rates both systems' outputs for the same input:

n=(zα/2+zβ)2σd2δ2=(1.96+0.84)2×0.810.09=7.84×0.810.09=6.350.09≈71 itemsn = \frac{(z_{\alpha/2} + z_{\beta})^2 \sigma_d^2}{\delta^2} = \frac{(1.96 + 0.84)^2 \times 0.81}{0.09} = \frac{7.84 \times 0.81}{0.09} = \frac{6.35}{0.09} \approx 71 \text{ items}

Unpaired design — different items for each system, with per-rating standard deviation σ=1.0\sigma = 1.0:

nper group=2(zα/2+zβ)2σ2δ2=2×7.84×1.00.09≈175 items eachn_{\text{per group}} = \frac{2(z_{\alpha/2} + z_{\beta})^2 \sigma^2}{\delta^2} = \frac{2 \times 7.84 \times 1.0}{0.09} \approx 175 \text{ items each}

71 paired items against 350 total unpaired. Pairing is the single highest-leverage design choice available, it costs nothing, and it is routinely thrown away by harnesses that sample a fresh evaluation set per run.

Now cost it. 5,000 items, 3 raters, 90 seconds per judgement:

15,000 judgements×90s=1,350,000s=375 hours×25 USD=9,375 USD15{,}000 \text{ judgements} \times 90\text{s} = 1{,}350{,}000\text{s} = 375 \text{ hours} \times 25 \text{ USD} = 9{,}375 \text{ USD}

Against 71 paired items with 3 raters: 213 judgements, about 5.3 hours, roughly 133 dollars — plus the pilot. The opening team spent seventy times what the question required and still could not answer it, because the money went into sample size when the problem was measurement error.

More annotations cannot fix a rubric that two people read differently. Spend the first day on agreement and the second on volume, never the other way round.

The interface is part of the instrument

Interface decisions change the numbers you get. These are not cosmetic:

  • Blind by default. Model identity, and any hint of it (response length, characteristic formatting, a "System A" label that always appears left), must be hidden. Unblinded annotation measures expectations.
  • Randomise order per annotator. Both the order of items and the left/right placement of compared outputs. Position effects of 5 to 15 points are routine.
  • One dimension visible at a time, or at least scored in a fixed sequence with factuality first. Presenting all dimensions together produces halo effects, where a strong impression on one bleeds into the rest.
  • Require a free-text justification on extreme scores (1 and 5). It slows extreme ratings just enough to make them deliberate, and the text is the qualitative data you will want later.
  • Show the source document inline, not behind a link. Any verification step requiring a click will be skipped under time pressure, and your faithfulness ratings will quietly become fluency ratings.
  • Log time per item. A judgement made in four seconds did not involve reading the source. Time-on-task is the cheapest quality signal available and almost nobody records it.
  • Cap session length at about 90 minutes. Agreement measurably degrades past that, and the degradation is systematic — later items get more central ratings.

What this means when you commission a human evaluation

Refuse to start collection until you have a pilot kappa. The pilot is 40 items and two raters — half a day, a few hundred dollars — and it is the only thing standing between you and the opening scenario. If the number comes back below 0.4, the correct response is to rewrite the rubric, not to add annotators, not to average harder, and not to proceed while noting the caveat in a footnote nobody reads.

Design the comparison as paired from the start, because it changes the required sample by a factor of five and the decision is irreversible once collection begins. Store every individual rating, not just the aggregate: per-rater, per-item, with timestamps and time-on-task. Aggregates cannot be re-analysed, and every interesting question that comes up later — did agreement drift, is one annotator an outlier, are the errors concentrated in one item type — needs the raw data.

And report the evaluation's own reliability alongside its results, as a matter of routine. "Model A scored 3.71 against 3.64, paired difference +0.07 with a 95% interval of [−0.11, +0.25], Fleiss kappa 0.33" is an honest sentence that tells a reader exactly how much to believe. "Model A scored 3.71 versus 3.64" is the same study with the caveats removed, and it is how a nine-thousand-dollar non-result becomes a product decision.