Course Content
Evaluating and Testing GenAI Models
4 sections · 13 lessons
Designing Human Evaluation Pipelines
A team spends 9,000 dollars having three annotators rate 5,000 model outputs on a 1-to-5 quality scale. The results come back: model A averages 3.71, model B averages 3.64. Leadership asks whether A is better. The data scientist checks agreement between annotators and finds Fleiss' kappa of 0.33.
That kappa means the annotators were, to a substantial degree, not measuring the same thing. The 0.07 difference between the models is smaller than the disagreement between two people looking at the same output. Nine thousand dollars bought a number that cannot answer the question it was commissioned to answer, and the failure happened before the first annotation was collected — in the design.
Human evaluation is the only method that can measure the things that matter most about generative systems. It is also an instrument, and like any instrument it has a noise floor, a calibration procedure, and a set of ways to be used wrong. This lesson is about building it so the number at the end means something.
The pipeline, and where each stage fails
[1] Define construct -> What exactly are we measuring? What does bad look like?[2] Design rubric -> Dimensions, scales, anchors, gates[3] Write guidelines -> Edge cases, worked examples, disagreement resolution[4] Pilot -> 30-50 items, 2+ raters, MEASURE AGREEMENT[5] Fix rubric -> Iterate 2-4 until agreement is usable. Do not skip.[6] Recruit + train -> Qualification task with a pass bar[7] Sample items -> Stratified, sized from a power calculation[8] Annotate -> Blind, randomised, with gold items seeded in[9] Monitor -> Per-annotator agreement, drift, throughput, gold accuracy[10] Adjudicate -> Resolve disagreements by a stated rule[11] Analyse -> Effect sizes and intervals, not just meansStages 4 and 5 are the ones teams skip under deadline pressure, and skipping them is what produced the opening failure. Piloting costs about half a day. Discovering a broken rubric after collecting 15,000 judgements costs the entire budget.
Defining the construct before the rubric
The question "rate this output 1–5 for quality" fails because quality is not one thing. Before writing any scale, answer three questions in writing:
- What decision will this number inform? Ship or not ship? Which of two prompts? Whether to escalate to human review? Different decisions need different measurements, and a rubric built for none of them in particular serves none of them.
- What does a bad output look like? Collect 50 real bad outputs. Write one sentence per output about what is wrong. Cluster the sentences. Those clusters are your dimensions — derived from evidence, not from a meeting.
- Which dimensions are gates and which are scores? A gate is a floor below which nothing else matters (factual error, safety violation, missing required disclosure). A score ranks the outputs that clear the gates. Mixing them into one weighted average lets an elegant, dangerous output outrank a plain, correct one.
If you cannot describe, in one sentence, the decision your rating will change, you are not ready to write a rubric.
Anchors are the rubric
A scale without anchors produces central-tendency bias: almost every rating lands on 3 or 4, variance collapses, and no comparison is possible. Every scale point needs a description specific enough that two people reading it reach the same conclusion.
| Score | Weak anchor (produces kappa ~0.3) | Strong anchor (produces kappa ~0.7) |
|---|---|---|
| 5 | Excellent | Answers the full question, every claim traceable to the source, no unnecessary content, correct format. A senior agent would send this unedited. |
| 4 | Good | Answers the question correctly; one stylistic edit needed, or one non-essential detail missing. Sendable after a 10-second fix. |
| 3 | Acceptable | Answers the main question but omits a secondary part, or includes one correct-but-irrelevant paragraph. Needs a minute of editing. |
| 2 | Poor | Addresses the topic but misses the actual question, or requires rewriting more than half the text. |
| 1 | Very poor | Wrong question answered, or contains a claim contradicting the source, or unusable format. Faster to write from scratch. |
Notice what the strong anchors have in common: they refer to observable properties and actions ("sendable after a 10-second fix", "faster to write from scratch"), not to internal impressions. An annotator can check whether a claim is traceable. They cannot check whether something is excellent.
Guidelines that actually reduce disagreement
A guideline document that only restates the rubric is decoration. The parts that move agreement are the parts about hard cases.
| Section | What it contains | Why it matters |
|---|---|---|
| Task framing | Who reads these outputs and what they do next | Annotators calibrate to a real user, not an imagined one |
| Dimension definitions | One paragraph each, with the question to ask yourself | Prevents dimension bleed (rating fluency under accuracy) |
| Anchored scales | As above, with two real examples per point | The single biggest lever on agreement |
| Edge-case rulings | "If the output refuses, score X." "If it is correct but off-topic, score Y." | This is where most disagreement lives |
| Explicit non-criteria | "Do not penalise British vs American spelling." "Do not reward length." | Removes systematic per-annotator biases |
| Worked examples | 8–12 fully annotated items with reasoning | Doubles as the training set |
| Escalation rule | What to do when genuinely unsure; a "flag" option | Stops guessing, which is pure noise |
The edge-case section should grow during the pilot. Every disagreement in the pilot is a missing ruling. Write the ruling, do not blame the annotator — disagreement is nearly always a specification defect.
Give annotators a "cannot decide" option and count how often it is used. An item flagged as undecidable is information; the same item guessed on is noise dressed as data. If more than about 8% of items get flagged, the rubric does not fit the output distribution.
Measuring agreement, properly
Raw percentage agreement is the wrong statistic because it credits agreement that would happen by chance. Which chance-corrected measure you need depends on your design.
| Measure | Use when | Handles |
|---|---|---|
| Cohen's κ | Exactly 2 raters, categorical labels, same raters throughout | Nominal categories only |
| Weighted κ | 2 raters, ordinal scale | Treats 4-vs-5 as a smaller error than 1-vs-5 |
| Fleiss' κ | 3+ raters, categorical, raters may differ per item | Variable rater panels |
| Krippendorff's α | Any number of raters, any scale type, missing data | The most general; use it if unsure |
| ICC | Continuous or ordinal ratings, and you care about absolute values | Rater bias as well as inconsistency |
| Spearman ρ | You only care about the ranking of systems | Systematic offsets between raters |
Fleiss' kappa, worked end to end
Ten outputs, three raters each, three categories (Good / Fair / Poor). The counts per item:
| Item | Good | Fair | Poor | ∑jnij2 | Pi |
|---|---|---|---|---|---|
| 1 | 3 | 0 | 0 | 9 | 1.000 |
| 2 | 2 | 1 | 0 | 5 | 0.333 |
| 3 | 0 | 2 | 1 | 5 | 0.333 |
| 4 | 0 | 0 | 3 | 9 | 1.000 |
| 5 | 1 | 2 | 0 | 5 | 0.333 |
| 6 | 3 | 0 | 0 | 9 | 1.000 |
| 7 | 0 | 1 | 2 | 5 | 0.333 |
| 8 | 2 | 1 | 0 | 5 | 0.333 |
| 9 | 0 | 3 | 0 | 9 | 1.000 |
| 10 | 1 | 1 | 1 | 3 | 0.000 |
| Total | 12 | 11 | 7 | 5.667 |
Each Pi is the proportion of rater pairs on item i that agree, computed as
Mean observed agreement:
Category marginals over all 10×3=30 ratings:
A kappa of 0.33. Note that raw pairwise agreement was 56.7%, which sounds like the raters mostly concurred. They did not — over a third of that agreement is what three people picking from three categories would produce by chance.
| κ | Conventional reading | What to do |
|---|---|---|
| < 0.20 | Slight | The rubric is broken. Do not collect data. |
| 0.21 – 0.40 | Fair | Rewrite anchors, add edge-case rulings, re-pilot. |
| 0.41 – 0.60 | Moderate | Usable for coarse comparisons only. Aggregate 3+ raters. |
| 0.61 – 0.80 | Substantial | Good. This is the realistic target for subjective dimensions. |
| > 0.80 | Almost perfect | Excellent — or the task is so easy it is not worth human time. |
1import numpy as np2from itertools import combinations34def fleiss_kappa(counts: np.ndarray) -> float:5 """counts[i][j] = number of raters assigning item i to category j."""6 N, _ = counts.shape7 n = counts[0].sum() # raters per item (constant)8 P_i = (np.sum(counts ** 2, axis=1) - n) / (n * (n - 1))9 P_bar = P_i.mean()10 p_j = counts.sum(axis=0) / (N * n)11 P_e = np.sum(p_j ** 2)12 return float((P_bar - P_e) / (1 - P_e))1314def pairwise_agreement(ratings_by_rater: dict) -> float:15 """Raw agreement across all rater pairs -- report ALONGSIDE kappa, never instead."""16 raters = list(ratings_by_rater)17 agree = total = 018 for a, b in combinations(raters, 2):19 ra, rb = ratings_by_rater[a], ratings_by_rater[b]20 shared = set(ra) & set(rb)21 agree += sum(1 for k in shared if ra[k] == rb[k])22 total += len(shared)23 return agree / total if total else float("nan")Do not average away a low kappa
There is a real temptation, on seeing kappa 0.33, to say "but with three raters the mean is more reliable". It is, somewhat — averaging reduces random error by roughly n. But low kappa is usually not random error; it is systematic disagreement about what is being measured. Averaging a rater who is scoring factuality with one who is scoring fluency produces a stable number that measures neither. Fix the construct, then average.
Annotators: recruiting, training, and keeping them honest
Match expertise to the construct
| Judgement required | Who can make it | Cost per hour (USD) |
|---|---|---|
| Is this fluent English? Is the format right? | General crowd workers | 10 – 20 |
| Does this summary match the source? | Trained crowd workers with a good rubric | 18 – 30 |
| Is this the answer a support agent would give? | Actual support agents | 30 – 60 |
| Is this clinically safe? Legally sound? | Licensed professionals | 150 – 500 |
Using cheaper annotators than the construct requires does not give you a noisier version of the right answer. It gives you a precise measurement of a different construct — usually fluency, because that is what a non-expert can perceive.
Qualification, not just training
Every annotator completes a 20-item qualification set with known answers before touching real data, and must hit a stated bar. Then review their first 30 real items individually, giving feedback on disagreements. This is expensive and it is cheaper than discovering after 3,000 items that one annotator misread a dimension throughout.
Gold items, sized correctly
Seed known-answer items into the live stream at roughly 8–10% and track each annotator's accuracy on them. But be honest about the statistical power of that check. With 200 items per annotator and 10% gold, you have 20 gold items. If the acceptable accuracy is 90%:
An annotator scoring 14/20 (70%) is (0.90−0.70)/0.067=3.0 standard errors below the bar — detectable. An annotator at 80% is only 1.5 standard errors below, well inside normal variation. Gold items reliably catch gross failure and cannot detect mild degradation, so pair them with agreement monitoring against the panel, which is far more sensitive.
Sizing the study before you spend the money
This is the calculation the opening team never did. On a 1–5 scale, suppose the within-item difference between two systems has standard deviation σd=0.9 and you want to detect a mean difference of δ=0.3 at 5% significance with 80% power.
Paired design — the same annotator rates both systems' outputs for the same input:
n=δ2(zα/2+zβ)2σd2=0.09(1.96+0.84)2×0.81=0.097.84×0.81=0.096.35≈71 itemsUnpaired design — different items for each system, with per-rating standard deviation σ=1.0:
nper group=δ22(zα/2+zβ)2σ2=0.092×7.84×1.0≈175 items each71 paired items against 350 total unpaired. Pairing is the single highest-leverage design choice available, it costs nothing, and it is routinely thrown away by harnesses that sample a fresh evaluation set per run.
Now cost it. 5,000 items, 3 raters, 90 seconds per judgement:
15,000 judgements×90s=1,350,000s=375 hours×25 USD=9,375 USDAgainst 71 paired items with 3 raters: 213 judgements, about 5.3 hours, roughly 133 dollars — plus the pilot. The opening team spent seventy times what the question required and still could not answer it, because the money went into sample size when the problem was measurement error.
More annotations cannot fix a rubric that two people read differently. Spend the first day on agreement and the second on volume, never the other way round.
The interface is part of the instrument
Interface decisions change the numbers you get. These are not cosmetic:
- Blind by default. Model identity, and any hint of it (response length, characteristic formatting, a "System A" label that always appears left), must be hidden. Unblinded annotation measures expectations.
- Randomise order per annotator. Both the order of items and the left/right placement of compared outputs. Position effects of 5 to 15 points are routine.
- One dimension visible at a time, or at least scored in a fixed sequence with factuality first. Presenting all dimensions together produces halo effects, where a strong impression on one bleeds into the rest.
- Require a free-text justification on extreme scores (1 and 5). It slows extreme ratings just enough to make them deliberate, and the text is the qualitative data you will want later.
- Show the source document inline, not behind a link. Any verification step requiring a click will be skipped under time pressure, and your faithfulness ratings will quietly become fluency ratings.
- Log time per item. A judgement made in four seconds did not involve reading the source. Time-on-task is the cheapest quality signal available and almost nobody records it.
- Cap session length at about 90 minutes. Agreement measurably degrades past that, and the degradation is systematic — later items get more central ratings.
What this means when you commission a human evaluation
Refuse to start collection until you have a pilot kappa. The pilot is 40 items and two raters — half a day, a few hundred dollars — and it is the only thing standing between you and the opening scenario. If the number comes back below 0.4, the correct response is to rewrite the rubric, not to add annotators, not to average harder, and not to proceed while noting the caveat in a footnote nobody reads.
Design the comparison as paired from the start, because it changes the required sample by a factor of five and the decision is irreversible once collection begins. Store every individual rating, not just the aggregate: per-rater, per-item, with timestamps and time-on-task. Aggregates cannot be re-analysed, and every interesting question that comes up later — did agreement drift, is one annotator an outlier, are the errors concentrated in one item type — needs the raw data.
And report the evaluation's own reliability alongside its results, as a matter of routine. "Model A scored 3.71 against 3.64, paired difference +0.07 with a 95% interval of [−0.11, +0.25], Fleiss kappa 0.33" is an honest sentence that tells a reader exactly how much to believe. "Model A scored 3.71 versus 3.64" is the same study with the caveats removed, and it is how a nine-thousand-dollar non-result becomes a product decision.