Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Model-as-judge done carefully
Some qualities are hard to check with code or an NLI model. Does the answer respond to what the person actually asked, or to a nearby question? Is the tone right for someone asking about bereavement leave? Does a paraphrased answer really contain the key fact? Human reviewers can judge these, but grading 250 cases takes about six hours, and PolicyPal changes several times a week.
A model-as-judge uses a language model to grade outputs. It can grade the whole suite in about three minutes for a few dollars. The PolicyPal team's first judge used a simple prompt: "Rate this answer from 1 to 10 for quality." It gave 96% of answers a 7 or higher. Human reviewers, on the same answers, rated 81% as acceptable. The judge was fast, cheap and wrong.
A judge is a measuring instrument. Like any instrument, you must design it for one job, check it against a trusted reference, and know its biases before you read its numbers.
Narrow, binary criteria with reference facts
Scores from 1 to 10 are a trap. The judge clusters around 7 and 8, the meaning of "7" drifts between runs, and there is no way to say what a 6 should have been. Replace one vague score with several yes-or-no questions, each with a precise definition, and give the judge the reference data from the eval case, so it grades against Harbourline's facts, not its own knowledge.
1# policypal/evals/judge.py2from typing import Literal34from pydantic import BaseModel, ConfigDict56class Verdict(BaseModel):7 model_config = ConfigDict(extra="forbid")8 evidence: str # quotes compared, written before the verdicts9 covers_key_facts: Literal["yes", "no"]10 addresses_question: Literal["yes", "no"]11 contradicts_reference: Literal["yes", "no"]1213JUDGE_SYSTEM = """You grade answers from an HR and IT policy assistant.14Judge only against the KEY FACTS and MUST NOT lists given, never your own knowledge.15Length and style are not quality: a short answer that covers the key facts is fully correct.16First write evidence: for each key fact, quote the words in the answer that state it, or17write "missing". Then decide:18- covers_key_facts: "yes" only if every key fact is stated or clearly implied.19- addresses_question: "yes" if the answer responds to what was asked, not a nearby question.20- contradicts_reference: "yes" if the answer states anything in MUST NOT or contradicts a key fact."""2122def judge(llm, case: dict, answer: str) -> Verdict:23 facts = "\n- ".join(case["key_facts"])24 must_not = "\n- ".join(case.get("must_not") or ["(none)"])25 user = (f"QUESTION:\n{case['input']}\n\nKEY FACTS:\n- {facts}\n\n"26 f"MUST NOT:\n- {must_not}\n\nANSWER:\n{answer}")27 reply = llm.complete(JUDGE_SYSTEM, [{"role": "user", "content": user}],28 schema=Verdict.model_json_schema(), max_tokens=800)29 return Verdict.model_validate_json(reply.text)Three details matter. The evidence field comes before the verdicts, so the judge must find the words before it decides; judges that decide first and justify afterwards agree with humans noticeably less. The prompt says outright that length is not quality, which counters a known bias. And the judge runs on a larger model than PolicyPal's answering model, from LLM(model=...) with a different configuration value, because grading is harder than answering and the volume is small: 250 cases cost about $4 on the large tier.
Calibrate against humans before you trust it
A judge's numbers mean nothing until you know how often it agrees with people. PolicyPal's calibration set is 100 answers, each graded on the same three questions by two HR reviewers working separately.
First, measure how often the humans agree with each other. That is the ceiling: no judge can be expected to agree with a human more than humans agree among themselves. Then measure the judge against the humans' consensus. The usual statistic is Cohen's kappa, which measures agreement beyond what chance alone would produce: 0 means chance, 1 means perfect.
1from sklearn.metrics import cohen_kappa_score23# calib: 100 rows, each with the case, the answer and both reviewers' labels4# big_llm = LLM(model=JUDGE_MODEL), a larger model than the one being judged5human_a = [row["a_covers"] for row in calib] # "yes" / "no"6human_b = [row["b_covers"] for row in calib]7consensus = [row["consensus_covers"] for row in calib]8judged = [judge(big_llm, row["case"], row["answer"]).covers_key_facts for row in calib]910print("human vs human:", round(cohen_kappa_score(human_a, human_b), 2))11print("judge vs consensus:", round(cohen_kappa_score(consensus, judged), 2))For covers_key_facts, the two reviewers agreed with a kappa of 0.81. The first version of the judge agreed with their consensus at 0.52. The team read all 17 disagreements. Most were one pattern: the judge accepted an answer that stated a fact approximately, such as "about two weeks" for "10 working days". They changed "stated or implied" to "stated or clearly implied, with the same numbers", added the evidence-first rule, and re-measured: 0.78.
Then comes the step teams skip. The rubric was now tuned to those 100 answers, so they checked it on 50 fresh answers graded by the same reviewers. Kappa there was 0.76, close enough to trust. If it had dropped sharply, the rubric would have been overfitted to the calibration set.
Biases and how to reduce them
| Bias | What happens | Reduction |
|---|---|---|
| Position | In A-versus-B comparisons, the judge favours whichever answer comes first | Run both orders; count a win only if it holds in both |
| Length | Longer answers are rated higher, even when padded | Say length is not quality; check verdicts against answer length |
| Self-preference | A judge rates answers from its own model family more kindly | Use a different model family, or at least check for the effect |
| Leniency | Fluent, confident, wrong answers pass | Give reference facts; require quoted evidence |
Absolute grading, where each answer is judged against reference facts on its own, is what PolicyPal uses for its suite. Pairwise grading, where the judge sees two versions' answers side by side and picks the better one, is useful when there is no reference, such as comparing tone or clarity between two prompts. It is also where position bias is strongest. In PolicyPal's test, the judge preferred whichever answer came first 61% of the time when both answers were identical in substance. Running both orders and counting only consistent wins removed the effect, at twice the cost.
PolicyPal checks the length bias directly: it compares the average length of answers the judge passes and fails, within each group. After the rubric change, there was no meaningful difference. Before it, passed answers were 30% longer.
Where a judge should not be the gate
A judge is good for qualities that are hard to express as rules, across many cases, where a small error rate is acceptable. It is the wrong tool for anything you can check exactly in code, like valid citations, correct tool arguments or a leave number, because code is free, instant and never wrong about its rule. It should not be the only gate for safety-critical behaviour either; the sensitive route and the adversarial cases are checked with exact rules and reviewed by people.
Judges also drift. A provider updates a model, and the same judge prompt grades differently. PolicyPal pins the judge's model version and re-runs 30 human-labelled cases every month. If kappa falls below 0.7, judge-based numbers are marked as untrusted until someone recalibrates.
Check your understanding
0 of 3 answered
1.A judge agrees with human consensus at kappa 0.78, and the two humans agree with each other at 0.81. How should you read this?
2.Why does the judge write evidence before the yes-or-no verdicts?
3.The team tuned the judge's rubric on 100 calibration answers. Why check it again on 50 fresh answers?