Evaluating and Testing GenAI Models

Types of Evaluation (Automatic vs. Human)


A team has two candidate models for a customer-support summariser. They run both over a held-out set of 100 support threads, score each summary as acceptable or not, and get a clean-looking result: model A is acceptable on 70 of them, model B on 74. Four points better. B ships on Friday.

By the following Thursday the support leads are complaining that summaries have started dropping the customer's order number. Nobody can find the regression in the numbers, because the numbers never showed it. The eval said 74 versus 70 and everyone read that as "B is better", when what the eval actually said was closer to "these two models are indistinguishable at this sample size, and by the way I never looked at order numbers."

Two separate failures are stacked on top of each other here, and they are the two failures that define this whole subject. The first is statistical: a 4-point gap on 100 examples is noise. The second is construct: the thing being measured was not the thing anyone cared about. Automatic evaluation is very good at giving you the first failure at scale. Human evaluation is very good at catching the second. Neither is optional, and knowing which one you are holding at any moment is most of the skill.

Model A 70, model B 74 — on the same 100 threads60141016A acceptableA not acceptableB acceptableB not acceptable76 threads agree and carry no information; the verdict rests on 24 discordant pairs.
Scoring both models on the same examples turns a four-point gap into a paired test that 24 threads cannot settle.

Why the 4-point gap was noise

Start with the arithmetic, because it is short and it settles the argument permanently.

When you score 100 examples pass/fail and get 70 passes, you have estimated a proportion p^=0.70\hat{p} = 0.70 from a sample. That estimate has a standard error — the typical amount it would wobble if you drew a fresh sample of 100 from the same underlying distribution:

SE(p^)=p^(1−p^)n=0.70×0.30100=0.0021=0.0458SE(\hat{p}) = \sqrt{\frac{\hat{p}(1-\hat{p})}{n}} = \sqrt{\frac{0.70 \times 0.30}{100}} = \sqrt{0.0021} = 0.0458

That is 4.58 percentage points. The wobble on a single model's score is already larger than the entire gap you were excited about.

Now the gap itself. If the two models were evaluated on independent samples, the standard error of the difference is the square root of the sum of the two variances:

SE(p^B−p^A)=0.74×0.26100+0.70×0.30100=0.001924+0.0021=0.004024=0.0634SE(\hat{p}_B - \hat{p}_A) = \sqrt{\frac{0.74 \times 0.26}{100} + \frac{0.70 \times 0.30}{100}} = \sqrt{0.001924 + 0.0021} = \sqrt{0.004024} = 0.0634

6.34 percentage points. Your observed difference of 4 points is 0.63 standard errors from zero:

z=0.040.0634=0.63p≈0.53z = \frac{0.04}{0.0634} = 0.63 \qquad p \approx 0.53

A coin flip. The 95% confidence interval for the true difference runs from 0.04−1.96×0.0634=−0.0840.04 - 1.96 \times 0.0634 = -0.084 to 0.04+1.96×0.0634=+0.1640.04 + 1.96 \times 0.0634 = +0.164 — that is, anywhere from B is 8 points worse to B is 16 points better. The experiment did not answer the question it was run to answer.

A difference you cannot distinguish from zero is not a small win. It is an unanswered question, and shipping on it is shipping on a guess.

How many examples would you have needed?

To detect a true 4-point difference (70% versus 74%) at the conventional 5% significance level with 80% power, the required sample per model is:

n=(zα/2+zβ)2[pA(1−pA)+pB(1−pB)]δ2=(1.96+0.84)2×(0.21+0.1924)0.042n = \frac{(z_{\alpha/2} + z_{\beta})^2 \left[p_A(1-p_A) + p_B(1-p_B)\right]}{\delta^2} = \frac{(1.96 + 0.84)^2 \times (0.21 + 0.1924)}{0.04^2}

n=7.84×0.40240.0016=3.15480.0016≈1972n = \frac{7.84 \times 0.4024}{0.0016} = \frac{3.1548}{0.0016} \approx 1972

Roughly two thousand examples per model. That number surprises people every time. It is also the single most useful number in this lesson, because it tells you what your 100-example eval set can and cannot do: it can catch a catastrophe (a model that drops from 70% to 30%), and it cannot adjudicate a 4-point improvement. Both facts matter.

The cheap fix: evaluate both models on the same examples

The calculation above assumed independent samples. You almost never need that. If both models see the same 100 threads, the examples' intrinsic difficulty cancels out, and you should analyse only the cases where the two models disagree. This is McNemar's test.

Suppose on those 100 threads: both models pass on 60, both fail on 16, B passes where A fails on 14, and A passes where B fails on 10. The marginals reproduce the original scores (B: 60+14 = 74, A: 60+10 = 70), but now the arithmetic uses only the 24 discordant items:

z=b−cb+c=14−1024=44.899=0.82p≈0.41z = \frac{b - c}{\sqrt{b + c}} = \frac{14 - 10}{\sqrt{24}} = \frac{4}{4.899} = 0.82 \qquad p \approx 0.41

Still noise — but notice the standard error fell from 6.34 points to 24/100=4.90\sqrt{24}/100 = 4.90 points just by pairing. Pairing is free. Never run an unpaired comparison when a paired one is available.

How many paired examples would resolve a 4-point gap with the same 80% power, if roughly 24% of items are discordant? The paired version of the formula replaces the two models' variances with the discordant rate:

n=(zα/2+zβ)2×0.24δ2=7.84×0.240.0016≈1176n = \frac{(z_{\alpha/2} + z_{\beta})^2 \times 0.24}{\delta^2} = \frac{7.84 \times 0.24}{0.0016} \approx 1176

About 1,180 shared examples instead of 1,972 separate ones for each model: 40% fewer items to collect and label, and about 2,350 model outputs to score instead of 3,944. That is the return on a five-line change to your harness. (Drop the power requirement and solve z=1.96z = 1.96 alone and you get n≈577n \approx 577 — but a study of that size would miss a real 4-point gap half the time.)

What automatic evaluation actually is

Automatic evaluation means any scoring procedure that runs without a person in the loop at scoring time: a function takes the model's output (and usually a reference answer or a source document) and returns a number.

The family is broader than most people assume:

KindExamplesWhat it needsWhat it is blind to
Exact / rule matchString equality, regex, JSON schema validation, unit tests on generated codeA precise expected answerAny correct answer phrased differently
N-gram overlapBLEU, ROUGE, METEOROne or more reference textsParaphrase, meaning, factuality
Embedding similarityBERTScore, MoverScore, cosine similarityReferences plus an encoder modelTruth; fluent-but-false scores well
Model-intrinsicPerplexity, token log-likelihoodAccess to model probabilitiesTask success entirely
Learned metricsBLEURT, COMET, reward modelsHuman ratings to train onAnything outside its training distribution
LLM-as-judgeA strong model scoring against a rubricA rubric and an API budgetIts own biases: position, length, self-preference
Behavioural checksRefusal rate, latency, cost per response, toxicity classifier hitsInstrumentationWhether the answer was any good

What unites them is the property that makes them valuable: they are deterministic and free at the margin. Once written, the hundred-thousandth evaluation costs the same as the first. That is what makes them usable inside continuous integration, inside a training loop, on every pull request.

The three failure modes of automatic metrics

These recur constantly, so name them:

  • Construct mismatch. The metric measures surface form; you care about meaning. ROUGE cannot tell you whether the order number survived, unless you deliberately wrote a check for order numbers.
  • Ceiling saturation. Once every candidate model scores 0.91–0.93 on your benchmark, the metric has stopped carrying information. Differences at the top of a saturated metric are almost entirely noise plus test-set idiosyncrasy.
  • Gameability. Any metric used as a training or selection target degrades as a measure. Optimise for ROUGE and you get outputs that copy long spans from the source. Optimise for an LLM judge that likes long answers and you get padding.

Goodhart's law is not a warning about the future of your metric. It is a description of what happens the week after you put it on a dashboard.

What human evaluation actually is

Human evaluation means people read model outputs and record structured judgements — a rating on a rubric, a preference between two candidates, a label for an error type, a span highlighted as unsupported.

It is the only method that can measure things with no computable definition: is this summary useful to a support agent? Is this tone appropriate for a bereavement email? Is this explanation actually going to help a beginner? For those questions, human judgement is not an approximation to the ground truth. It is the ground truth, by definition of the construct.

But human evaluation carries its own error structure, and it is not the friendly, well-behaved kind:

Source of errorWhat it looks likeMitigation
Sampling noiseSame statistics as above, but with n=50n = 50 instead of 5,000Report confidence intervals; never quote a mean without one
Rater disagreementTwo annotators agree on 70% of items and it still means almost nothingMeasure chance-corrected agreement, not raw agreement
Rubric driftWeek 4 ratings are systematically harsher than week 1Recurring gold-standard items; re-calibration sessions
Order and anchoring biasThe first output seen sets the standard for the restRandomise presentation order per rater
Length and fluency haloLonger, more confident text is rated as more accurateScore dimensions separately; blind raters to source model
Cost and latencyTwo days and several hundred dollars per full runReserve for release gates, not for every commit

Why 70% agreement can mean nothing

This is the single most common misreading of a human eval report. Two annotators rate 100 outputs as acceptable or not. They agree on 70. That sounds respectable until you subtract the agreement you would get by chance.

Suppose the confusion matrix is: both say acceptable on 45, both say unacceptable on 25, annotator A says acceptable while B says unacceptable on 20, and the reverse on 10.

Observed agreement: Po=(45+25)/100=0.70P_o = (45 + 25)/100 = 0.70.

Annotator A calls 65 items acceptable; annotator B calls 55 acceptable. Expected agreement by chance:

Pe=(0.65×0.55)+(0.35×0.45)=0.3575+0.1575=0.515P_e = (0.65 \times 0.55) + (0.35 \times 0.45) = 0.3575 + 0.1575 = 0.515

κ=Po−Pe1−Pe=0.70−0.5151−0.515=0.1850.485=0.38\kappa = \frac{P_o - P_e}{1 - P_e} = \frac{0.70 - 0.515}{1 - 0.515} = \frac{0.185}{0.485} = 0.38

A kappa of 0.38 is conventionally described as "fair" — which in practice means your rubric is ambiguous and your annotators are partly guessing. The 70% headline hid that completely. If you take one habit from this lesson, take the habit of computing kappa before you believe an agreement number.

Choosing between them

DimensionAutomaticHuman
Marginal cost per itemEffectively zero (or fractions of a cent for LLM judges)0.30–2.00 US dollars, more for expert domains
TurnaroundSeconds to minutesHours to days
ReproducibleRule and overlap metrics: bit-for-bit. LLM judges: only approximately, even with a pinned model versionNo — same rater, different day, different answer
Measures what you care aboutOnly if you engineered it toYes, if the rubric is good
Detects novel failure modesNever — it only checks what it was told to checkYes, this is its unique strength
Scales to 100k itemsTriviallyNot without an operations team
Suitable as a training signalYes, with the Goodhart caveatOnly via a learned reward model
Statistical power at fixed budgetHigh (large nn)Low (small nn) — plan for it

The decision rule that survives contact with reality:

  • Use automatic evaluation when you need a regression gate on every commit, when you are comparing many configurations, when the task has verifiable answers (code that compiles and passes tests, extraction against a schema, arithmetic), and when you need enough nn for the statistics to work.
  • Use human evaluation when the quality dimension has no computable definition, when you are about to make an irreversible decision (a release, a vendor choice), when you are exploring what is going wrong rather than measuring how much, and when you need to validate that an automatic metric still tracks reality.
  • Use both — which in practice means always, because the third bullet is not optional. An automatic metric that has never been checked against human judgement is a number of unknown meaning.

The hybrid loop that works in production

The productive arrangement is not "run both and hope they agree". It is a loop where each corrects the other's blind spot.

Text
1. Humans label a calibration set     (150-400 items, done once, refreshed quarterly)       |       v2. Fit / select the automatic metric   -- keep only metrics that correlate with (1)       |                                  report Spearman rho and kappa, not vibes       v3. Run the automatic metric everywhere -- every commit, every prompt change, full traffic sample       |       v4. Route the tails to humans           -- lowest-scoring 5%, judge-uncertain items,       |                                  plus a random 2% audit stream       v5. New human labels feed back to (1)   -- catches metric drift and new failure modes

Step 4 is what people leave out, and leaving it out is why hybrid pipelines quietly rot. If humans only ever look at the items the metric already flagged, you will never discover failures the metric is blind to. The random audit stream is small, unglamorous, and the only thing standing between you and a silent construct failure — like a summariser that dropped order numbers for two weeks.

The random audit sample is the cheapest insurance in evaluation. Two percent of traffic, reviewed by a person who is allowed to say "this is wrong in a way we do not have a metric for."

Worked cost model

Ten thousand outputs to evaluate. Human review costs 0.60 US dollars per item; an LLM judge costs 0.004 US dollars per item.

StrategyCostCoverageStatistical power
All human10,000 × 0.60 = 6,000 USD100%Excellent, but you run it twice a year
All LLM judge10,000 × 0.004 = 40 USD100%Excellent nn, unknown validity
Hybrid: judge everything, human-review 15%40 + (1,500 × 0.60) = 940 USD100% screened, 15% verifiedBoth — and the 15% validates the judge

The hybrid costs 84% less than the all-human run (6,000 down to 940) while producing the one thing the all-LLM run cannot: evidence that the judge is measuring the right thing.

Worked example: evaluating a support-thread summariser

Concretely, here is what a defensible evaluation of the model from the opening looks like.

Define the construct first. "Good summary" is too vague to score. Break it into things a person can disagree about clearly:

  1. Faithfulness — every claim in the summary is supported by the thread. Binary per claim.
  2. Key-field retention — order number, product name, and resolution status all present when present in the thread. Deterministic, checkable by regex.
  3. Actionability — an agent reading only the summary knows what to do next. 1–5 rubric, human-scored.
  4. Length discipline — under 80 words. Deterministic.

Notice that two of the four are now automatic and exact. The order-number regression that took two weeks to find would have been caught on the first CI run, because criterion 2 is a rule, not a similarity score.

Python
import reORDER_RE = re.compile(r"\bORD-\d{6,8}\b")def key_field_retention(thread: str, summary: str) -> float:    """Fraction of critical fields present in the thread that survive into the summary."""    expected = set(ORDER_RE.findall(thread))    if not expected:        return 1.0                       # nothing to retain -> not a failure    kept = expected & set(ORDER_RE.findall(summary))    return len(kept) / len(expected)def length_ok(summary: str, limit: int = 80) -> bool:    return len(summary.split()) <= limit

Then size the experiment. Key-field retention is a proportion, so the same arithmetic applies. If model A retains 96% of order numbers and you want to detect a drop to 92%, you need — by the formula above with pA=0.96p_A = 0.96, pB=0.92p_B = 0.92:

n=7.84×(0.96×0.04+0.92×0.08)0.042=7.84×(0.0384+0.0736)0.0016=0.878080.0016≈549n = \frac{7.84 \times (0.96 \times 0.04 + 0.92 \times 0.08)}{0.04^2} = \frac{7.84 \times (0.0384 + 0.0736)}{0.0016} = \frac{0.87808}{0.0016} \approx 549

549 threads per model — very achievable for an automatic check, hopeless for a human one. That asymmetry is exactly why the split above matters: put the high-precision, high-nn requirements on the automatic side, and spend the human budget on the one dimension (actionability) that no rule can capture.

Then report honestly. The results table should look like this, and a table that omits the interval column should not be accepted:

CriterionMethodnnModel AModel BDifference (95% CI)
Key-field retentionRegex, paired80096.1%88.4%−7.7 pts [−10.4, −5.0]
Length under 80 wordsRule, paired80099.5%99.8%+0.3 pts [−0.3, +0.9]
Faithfulness (claim-level)LLM judge, validated80091.2%92.0%+0.8 pts [−1.9, +3.5]
Actionability (1–5)Human, 3 raters1203.84.1+0.3 [−0.1, +0.7]

Read that table and the decision writes itself. B is genuinely worse on key-field retention — the interval excludes zero and the effect is large. B is possibly better on actionability, but the interval crosses zero, so that is a hypothesis, not a finding. The original "74 versus 70, ship it" summarised all of this into one number that pointed the wrong way.

What to do with this on Monday

Three changes to your evaluation setup, in order of how much they buy you per hour spent.

Put an interval on every number you report. Not a footnote — the same cell, same font size. The moment a metric can only be quoted with its uncertainty attached, the conversation about whether a change is real becomes automatic rather than adversarial. For a proportion from nn items, the interval half-width is roughly 1.96p^(1−p^)/n1.96\sqrt{\hat{p}(1-\hat{p})/n}; for n=100n = 100 and p^\hat{p} near 0.5 that is about 10 points, which is a number worth memorising as a sanity check.

Pair your comparisons and evaluate on shared examples. It costs nothing and, as the McNemar arithmetic showed, it cut the required sample from about 1,972 items per model to about 1,180 shared items in a realistic case. If your harness currently samples a fresh eval set per run, that is throwing away much of your statistical power for no benefit.

Convert every failure a human finds into an automatic check. The order-number regression should have produced a regex-based test within an hour of being diagnosed, so it can never recur silently. This is the ratchet that makes hybrid evaluation compound: humans discover failure modes, automation makes them permanent tests, and human attention moves on to the next unknown. An evaluation suite built this way gets strictly more informative over time. One built by adding more general-purpose similarity metrics does not.