Course Content
Evaluating and Testing GenAI Models
4 sections · 13 lessons
Combining LLM-Based Evaluators with Human Judgment
A team replaces their human evaluation panel with an LLM judge. The judge scores 10,000 outputs overnight for about 40 dollars, where the humans took two weeks and 6,000. Agreement with the previous human panel is 83%, which sounds excellent, and the team declares the migration a success.
Three weeks later the model's win rate against the previous version has climbed from 51% to 67% on the judge's scoring, while customer satisfaction has not moved at all. Someone reads a sample of the outputs and notices they have got noticeably longer. The prompt engineer had been iterating against the judge, and the judge — like nearly all LLM judges — prefers longer answers.
Two things went wrong. The 83% agreement figure was never chance-corrected, and once it is, the judge looks much weaker than advertised. And the judge was used as an optimisation target without anything watching for the specific ways it can be gamed. Both failures are avoidable, and avoiding them is the whole content of this lesson.
What an LLM judge is and is not
An LLM judge is a strong language model prompted with a rubric, one or two candidate outputs, and usually a source document, asked to return a structured verdict. It sits between automatic metrics and humans on every axis:
| Overlap metrics | LLM judge | Human panel | |
|---|---|---|---|
| Cost per item | ~0 | 0.002 – 0.02 USD | 0.30 – 2.00 USD |
| Throughput | Millions/hour | Thousands/hour | Tens/hour |
| Handles nuance | No | Substantially | Yes |
| Gives reasons | No | Yes (of variable honesty) | Yes |
| Reproducible | Exactly | Only approximately, even with a pinned model version | No |
| Systematic biases | Known and stable | Position, length, self-preference, style | Position, fatigue, anchoring |
| Detects novel failures | Never | Only if the rubric asks | Yes |
| Safe as an optimisation target | No | No | Reasonably |
The row that matters most is the last one. A judge is a model, and optimising against a model finds its weaknesses. That is not a flaw specific to LLM judges — it is Goodhart's law — but LLM judges have unusually exploitable weaknesses and unusually high perceived credibility, which is a bad combination.
Writing a judge that produces usable output
Most bad judges are bad prompts. Four things a judge prompt must do.
1JUDGE_PROMPT = """You are evaluating a customer-support answer against the2source policy document. Follow the criteria exactly and in order.34SOURCE DOCUMENT:5{source}67CUSTOMER QUESTION:8{question}910ANSWER TO EVALUATE:11{answer}1213CRITERIA, applied in this order:141. GROUNDED Every factual claim is supported by the SOURCE DOCUMENT.15 A claim that is true but absent from the source is NOT grounded.162. COMPLETE Addresses every part of the customer's question.173. ACTIONABLE The customer knows what to do next without asking again.1819Do NOT reward: length, confident tone, formatting, politeness.2021Return JSON only:22{{23 "ungrounded_claims": ["<verbatim quote>", ...],24 "grounded": 1-5,25 "complete": 1-5,26 "actionable": 1-5,27 "justification": "<= 40 words, citing the source section for each judgement"28}}2930Scale anchors for GROUNDED:315 = every claim traceable to a specific source sentence324 = all substantive claims traceable; one immaterial detail is not333 = one peripheral claim unsupported342 = a claim central to the answer is unsupported351 = a claim contradicts the source36"""- Extract evidence before scoring. The
ungrounded_claimsfield comes first deliberately. Forcing the judge to enumerate specific spans before assigning a number substantially improves accuracy and — more usefully — makes the verdict auditable. A score of 2 with an empty claims list is self-contradictory and can be caught automatically. - Anchor the scale. The same rule that governs human rubrics governs judges. Unanchored "rate 1–5" produces the same central-tendency clustering.
- Name the non-criteria. "Do NOT reward length, tone, formatting" measurably reduces those biases. It does not eliminate them.
- Force structured output and validate it. A judge that occasionally returns prose instead of JSON will silently drop items from your aggregate, biasing it in an unknown direction. Current model APIs can enforce a JSON schema natively (the next example does); use that rather than relying on "Return JSON only" in the prompt.
Pairwise judging, done properly
For comparisons, the judge is far more reliable than for absolute scoring — the same reason humans are. But it has a strong position bias, so every pair must be judged twice with the order swapped.
1import json23VERDICT = { # the API guarantees output matching this schema4 "type": "object",5 "properties": {"winner": {"type": "string", "enum": ["1", "2", "tie"]},6 "reason": {"type": "string"}},7 "required": ["winner", "reason"],8 "additionalProperties": False,9}1011def judge_pair(client, model, question, out_a, out_b, source):12 # PAIR_PROMPT: like JUDGE_PROMPT, but shows two responses and asks which is better13 def ask(first, second):14 prompt = PAIR_PROMPT.format(source=source, question=question,15 response_1=first, response_2=second)16 r = client.messages.create(17 model=model, max_tokens=4000,18 messages=[{"role": "user", "content": prompt}],19 output_config={"format": {"type": "json_schema", "schema": VERDICT}})20 text = next(b.text for b in r.content if b.type == "text")21 return json.loads(text)["winner"] # "1", "2", or "tie"2223 fwd = ask(out_a, out_b) # A shown first24 rev = ask(out_b, out_a) # B shown first2526 a_wins = (fwd == "1") + (rev == "2")27 b_wins = (fwd == "2") + (rev == "1")28 if a_wins == b_wins: # "tie" both ways is consistent; a 1-1 split is not29 return {"winner": "tie", "consistent": fwd == rev == "tie"}30 return {"winner": "A" if a_wins > b_wins else "B",31 "consistent": a_wins == 2 or b_wins == 2}The example uses the Anthropic Python SDK; other providers have equivalent options. Two details are deliberate. The output_config schema makes the API guarantee parseable JSON with one of three allowed winners, so no verdict disappears in a parse error. And there is no temperature=0: several current models (Claude Sonnet 5 and Opus 5 among them) reject sampling parameters outright, and even where temperature 0 is accepted, hosted models are not bit-for-bit deterministic. Pin the exact model version, and measure the judge's run-to-run stability instead of assuming it.
The consistent flag is the valuable part. A pair where the judge changes its mind when the order flips is a pair the judge cannot actually decide. Track the swap-consistency rate as a headline judge-quality metric — below about 80%, the judge is contributing more noise than signal.
Validating the judge: the step that was skipped
Return to the opening. The judge agrees with humans on 83% of items. Here is the confusion matrix behind that number, over 200 items with a binary accept/reject decision:
| Human \ Judge | Accept | Reject | Total |
|---|---|---|---|
| Accept | 128 | 12 | 140 |
| Reject | 22 | 38 | 60 |
| Total | 150 | 50 | 200 |
Both are heavily skewed towards "accept", so chance agreement is high:
Kappa 0.575 — moderate, not excellent. And the errors are asymmetric in a way that matters: the judge accepted 22 outputs humans rejected, against only 12 in the other direction. Its recall on rejection is 38/60=63%. The judge misses more than a third of the outputs humans consider unacceptable, which is precisely the failure mode you cannot afford if the judge is your release gate.
Raw agreement between a judge and a human is inflated by whatever the base rate happens to be. Always report kappa, and always report the two error directions separately, because their costs are never equal.
How many human labels does validation need?
To estimate agreement to within ±5 points at 95% confidence, when agreement is around 0.85:
About 200 human-labelled items. That is a single afternoon of expert time, it is the entire cost of knowing whether your judge works, and it is the thing teams skip. Refresh it quarterly, and re-run it after any change to the judge model, the judge prompt, or the output distribution being judged.
The four biases you must measure, not assume away
| Bias | How to measure it | Typical size | Mitigation |
|---|---|---|---|
| Position | Present each pair in both orders; compare win rates | 5–20 points | Always dual-present and average; report swap consistency |
| Length | Win rate as a function of the length difference, on quality-matched pairs | 10–25 points | Explicit instruction; length-matched validation set; report the correlation |
| Self-preference | Judge outputs from its own family versus others, with human labels as ground truth | 3–10 points | Use a judge from a different family than the system under test |
| Style / confidence | Compare hedged and assertive phrasings of the same correct content | 5–15 points | Name it as a non-criterion; validate on hedged examples |
Worked example of the position measurement. A judge prefers model A 74% of the time when A is shown first, and 51% when A is shown second.
Report 62.5%, not 74%. And note that with 11.5 points of position bias, any single-order evaluation of two similar systems is dominated by the presentation order rather than by the systems.
Hybrid workflows: three that work
1. Judge screens, humans adjudicate the tails
The judge scores everything. Humans review the bottom 5% by score, everything the judge marks low-confidence, and a random 2% audit sample. Simple, effective, and the random slice is what keeps the judge honest.
2. Selective prediction with a confidence threshold
The judge answers only where it is confident and abstains elsewhere. Worked numbers on 10,000 items:
| Coverage | Accuracy on covered | Cost (USD) | Overall accuracy | |
|---|---|---|---|---|
| Judge alone | 100% | 88% | 40 | 88.0% |
| Humans alone | 100% | 100% (definitional) | 6,000 | 100% |
| Judge ≥ 0.9 confidence, rest to humans | 82% judge / 18% human | 96% / 100% | 40 + 1,080 = 1,120 | 0.82(0.96)+0.18(1.0)=96.7% |
The hybrid reaches 96.7% agreement with expert judgement for 1,120 dollars against 6,000 — an 81% cost reduction for a 3.3-point accuracy shortfall. The reason it works is that judge confidence is genuinely informative about judge accuracy: on the covered 82%, the judge is right 96% of the time, against 88% overall.
Verify that claim before relying on it. Plot accuracy against confidence on your validation set. If accuracy is flat across confidence bins, the confidence signal is worthless and this workflow degenerates into "review a random 18%".
3. Multi-judge disagreement routing
Run two or three judges. Where they agree, accept. Where they disagree, route to a human. Disagreement is a better uncertainty signal than any single judge's self-reported confidence, because it is not produced by the same process that produces the error.
1from collections import Counter23def route(item, judges, human_queue, min_agree=2):4 verdicts = [j(item) for j in judges]5 top, count = Counter(verdicts).most_common(1)[0]6 if count >= min_agree and len(set(verdicts)) == 1:7 return {"verdict": top, "route": "auto", "confidence": "unanimous"}8 if count >= min_agree:9 return {"verdict": top, "route": "auto_flagged", "confidence": "majority"}10 human_queue.put(item)11 return {"verdict": None, "route": "human", "confidence": "split"}The ensemble arithmetic, and why it disappoints
Three judges, each 80% accurate. If their errors were independent, majority voting gives:
89.6% — a 9.6-point gain, which is why ensembles sound so attractive. But three prompts against the same base model do not have independent errors. They share a tokeniser, a training distribution, and the same length and style preferences. In practice, an ensemble of three prompt variants on one model measures around 82–84%, capturing perhaps a third of the theoretical gain.
To get meaningful independence, vary what actually differs: different model families, and ideally different judging methods — one LLM judge, one NLI-based faithfulness check, one rule-based structural check. Heterogeneous ensembles earn most of the theoretical gain; homogeneous ones do not.
An ensemble only helps to the extent its members fail differently. Three prompts against one model fail the same way three times, at three times the cost.
Calibrating a judge to your humans
A judge that disagrees systematically can often be fixed without touching the model.
- Few-shot with your own disagreements. Take the items where judge and human disagreed, and put the human verdict plus its reasoning into the prompt as examples. This is the highest-yield intervention and takes an hour.
- Rewrite the criterion the judge is misreading. Inspect the disagreements by category. If the judge marks "true but not in the source" as grounded, the definition needs to say so explicitly — as the example prompt above does.
- Fit a monotone recalibration. If the judge's 4 corresponds to a human 3 consistently, learn the mapping on the validation set and apply it. This fixes offset, not ordering.
- Fine-tune a small judge on your labels. With a few thousand human-labelled items, a fine-tuned small model often beats a large general model prompted zero-shot on your specific task, at a fraction of the inference cost.
Whatever you change, re-validate on held-out human labels. Tuning a judge on the same items you evaluate it on produces a judge that agrees with those items and nothing else — the same overfitting that afflicts any model, in a place people forget to look for it.
What this means when you put a judge in your pipeline
Version the judge like production code. The judge model, the prompt, the temperature, the parsing logic and the rubric together define your metric, and changing any of them changes what every historical number meant. Stamp a judge version on every stored verdict, and when you upgrade, re-score a frozen reference set with both versions so you can quantify the shift rather than discovering it as an unexplained step in a chart.
Never let the judge be both the optimisation target and the release gate. The team in the opening improved their win rate from 51% to 67% by writing longer answers, which the judge rewarded and users did not notice. Use the judge for iteration if you like — it is fast and cheap and that is what it is for — but gate the release on something the iteration loop never touched: a held-out human evaluation, an absolute acceptability check, or a production outcome metric.
And keep a permanent human sample in the loop, sized from the arithmetic above — around 200 items per quarter is enough to detect a meaningful drift in judge-human agreement. That sample costs roughly 120 dollars and is the only mechanism by which you will ever learn that your judge has stopped measuring what it used to measure. Without it, an LLM judge does not remove human evaluation from your pipeline; it removes your ability to know whether the pipeline is working.