LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What does it mean to use an LLM as an evaluator?


Judge vs HR specialist on 100 answers7081012judge: passjudge: failhuman: passhuman: failRaw agreement 0.82, chance agreement 0.668, so kappa = 0.46.
82% agreement sounds good until chance is removed, and the judge still misses 10 of the 22 bad answers.

What you need to know

What a judge call looks like

Python
JUDGE_PROMPT = """You grade answers from an HR policy assistant.Context:{context}Question: {question}Answer: {answer}Is every factual claim in the Answer supported by the Context?Think step by step, then return JSON:{{"reasoning": "...", "unsupported_claims": ["..."], "verdict": "pass" | "fail"}}"""

The judge sees exactly what it needs, answers one question, and returns structured output your code can parse. Most eval tools wrap this pattern: promptfoo's llm-rubric, DeepEval's GEval, LangSmith and Langfuse LLM-as-judge evaluators, and the OpenAI Evals API's model graders.

Why it works — and when it doesn't

Checking an answer is often easier than writing it, especially when the judge is given the context or a reference answer. So a judge can reliably ask "is the 15-day figure in the context?". It is much weaker when it would need knowledge it does not have, such as checking specialist medical or legal claims without a source.

Calibration: the non-negotiable step

Raw agreement is misleading when most answers pass: a judge that says "pass" to everything agrees 78% of the time with a human who passes 78%. Kappa removes that.

Python
from sklearn.metrics import cohen_kappa_score# 100 answers graded pass (1) / fail (0) by a human expert and by the judgehuman = [1]*70 + [0]*12 + [0]*10 + [1]*8judge = [1]*70 + [0]*12 + [1]*10 + [0]*8agreement = sum(h == j for h, j in zip(human, judge)) / len(human)print(f"raw agreement = {agreement:.2f}")print(f"Cohen's kappa = {cohen_kappa_score(human, judge):.2f}")
Text
raw agreement = 0.82Cohen's kappa = 0.46

By hand: the judge passes 80% and the human 78%, so chance agreement is 0.80 × 0.78 + 0.20 × 0.22 = 0.668. Kappa = (0.82 − 0.668) / (1 − 0.668) = 0.46 — only moderate. Look at the failure class too: the judge flagged 20 failures, 12 of them real (precision 0.60), and caught 12 of the 22 real failures (recall 0.55). It misses almost half the bad answers.

A real-life example

The HR assistant team wants a faithfulness judge on production traffic. An HR specialist labels 150 answers as pass or fail. The first judge prompt, "Is this answer accurate?", gets kappa 0.46 against her labels — the numbers above — and misses most answers that add an unsupported condition such as "only for permanent employees".

The team rewrites the prompt: list each claim, check each against the context, fail if any claim is missing from the context, with three graded examples. Kappa rises to 0.78, and failure recall to 0.86. Only then do they put the judge on 5% of traffic and on the CI suite. They repeat the 150-item check every quarter and whenever the judge model changes.

Follow-up questions to expect

  • "What kappa is good enough?" — Common rough guides call 0.6 to 0.8 substantial agreement; I'd also require that the judge agrees with humans about as well as two humans agree with each other.
  • "Why ask for reasoning before the verdict?" — So the verdict is conditioned on the analysis; a verdict first and reasons after tends to produce justifications for a snap judgement.
  • "Which model should be the judge?" — A strong model, ideally from a different family than the generator to limit self-preference, pinned to a specific version.