Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Build an evaluation pipeline using LLM-as-a-judge.
What you need to know
Most LLM outputs have no single correct string to compare against: "Refunds take five working days" and "You'll get your money back within 5 business days" are both right. LLM-as-a-judge uses a model to grade outputs against a written rubric — the list of criteria and what each score means.
What makes a judge trustworthy:
- Discrete, anchored scales. "1 = contradicts the reference, 3 = partly correct, 5 = fully correct" is repeatable; "rate 0–100" is not.
- Reasoning before the score. A model generates left to right; a score written first cannot be informed by reasoning written after it.
- A reference answer and the retrieved context in the prompt, so correctness and groundedness are judged against evidence, not the judge's memory.
- Calibration. Grade a sample by hand and report the judge's agreement with humans. Without that number, the scores are only a guess.
Known biases: position bias (in A-versus-B comparisons the first answer is favoured — run both orders), verbosity bias (longer answers score higher), and self-preference (a model rates its own family's style higher — use a different family as judge where you can).
1import json, re2from collections.abc import Callable3from concurrent.futures import ThreadPoolExecutor4from dataclasses import dataclass56KEYS = ("correctness", "groundedness", "completeness")7RUBRIC = """Grade the assistant's answer. Score each criterion from 1 to 5:8- correctness: 1 contradicts the reference, 3 partly right, 5 fully consistent with it9- groundedness: 1 claims not in the context, 5 every claim supported by the context10- completeness: 1 misses the question, 5 answers every part11Write brief reasoning first, then the scores. Reply with JSON only:12{"reasoning": "...", "correctness": 1, "groundedness": 1, "completeness": 1}"""1314@dataclass15class Case:16 question: str17 context: str18 reference: str1920def parse_verdict(raw: str) -> dict | None:21 """Pull the JSON object out of the reply (tolerates Markdown code fences) and validate it."""22 match = re.search(r"\{.*\}", raw, re.S)23 try:24 data = json.loads(match.group(0)) if match else None25 except json.JSONDecodeError:26 return None27 if not isinstance(data, dict) or not all(data.get(k) in (1, 2, 3, 4, 5) for k in KEYS):28 return None29 return data3031def judge(case: Case, answer: str, judge_fn: Callable[[str], str]) -> dict | None:32 prompt = (f"{RUBRIC}\n\nQuestion: {case.question}\nContext:\n{case.context}\n"33 f"Reference answer: {case.reference}\nAssistant answer: {answer}")34 return parse_verdict(judge_fn(prompt)) or parse_verdict(judge_fn(prompt)) # one retry3536def run_eval(cases: list[Case], system_fn: Callable, judge_fn: Callable, workers: int = 8):37 with ThreadPoolExecutor(max_workers=workers) as ex:38 answers = list(ex.map(lambda c: system_fn(c.question), cases))39 verdicts = list(ex.map(lambda ca: judge(ca[0], ca[1], judge_fn), zip(cases, answers)))40 good = [v for v in verdicts if v is not None]41 report = {"n": len(cases), "unparsed": len(verdicts) - len(good)}42 for k in KEYS:43 report[k] = round(sum(v[k] for v in good) / len(good), 2) if good else None44 return report, list(zip(cases, answers, verdicts))The tricky parts:
re.search(r"\{.*\}", raw, re.S)grabs the JSON even when the judge wraps it in Markdown fences or adds a sentence before it.re.Slets.match newlines.- Validating the scores (
in (1, 2, 3, 4, 5)) rejects"4",4.5or7. A judge that returns something off the scale is treated as unparsed, not averaged. - Unparsed verdicts are counted, not averaged in as zero. A spike in
unparsedmeans the judge prompt or model changed, which you must see. ex.mapkeeps input order, so answers and verdicts line up with cases even though they run concurrently.
Complexity: 2 model calls per case (3 when a verdict needs the retry), run workers at a time, so wall-clock time is roughly (2 × cases ÷ workers) × call latency. Aggregation is O(cases × criteria).
A real-life example
Three cases, a stub system and a stub judge (one reply is fenced, one is broken):
1cases = [Case("Refund time?", "Refunds take 5 working days.", "5 working days"),2 Case("COD refund?", "COD refunds go to store credit.", "Store credit"),3 Case("Cancel fee?", "Cancelling is free before dispatch.", "Free before dispatch")]4system = {"Refund time?": "5 working days.", "COD refund?": "To your bank account.",5 "Cancel fee?": "Free before dispatch."}.get6FENCE = "`" * 3 # a Markdown code fence7replies = iter([8 FENCE + 'json\n{"reasoning": "matches", "correctness": 5, "groundedness": 5, "completeness": 5}\n' + FENCE,9 '{"reasoning": "wrong destination", "correctness": 1, "groundedness": 1, "completeness": 4}',10 "I think this is good.", # unparseable, retried...11 "Still not JSON.", # ...and unparseable again12])13report, rows = run_eval(cases, system, lambda _: next(replies), workers=1)14print(report)15# {'n': 3, 'unparsed': 1, 'correctness': 3.0, 'groundedness': 3.0, 'completeness': 4.5}| case | answer | judge reply | verdict |
|---|---|---|---|
| Refund time? | 5 working days. | fenced JSON | parsed: 5, 5, 5 |
| COD refund? | To your bank account. | JSON | parsed: 1, 1, 4 |
| Cancel fee? | Free before dispatch. | prose, then prose again | unparsed (counted) |
Averages use only the two parsed verdicts: correctness (5 + 1) / 2 = 3.0, completeness (5 + 4) / 2 = 4.5. The COD answer is complete but wrong — exactly the case a single "overall quality" score would blur.
A fintech team runs a suite like this on every prompt change, and blocks the release if average correctness drops by more than a few tenths.
Follow-up questions to expect
- "How do you know the judge is right?" — Label 50–100 cases by hand, run the judge on them, and report agreement (for example Cohen's kappa). Recalibrate whenever the judge model or rubric changes.
- "Pairwise or absolute scores?" — Pairwise ("is A or B better?") is more reliable for comparing two prompts; absolute scores are better for tracking one system over time. Randomise the order in pairwise mode.
- "Can the judge be the same model as the system?" — It works, but self-preference inflates scores. Prefer a different model family, or at least a stronger model.