Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Build an evaluation pipeline using LLM-as-a-judge.


What you need to know

Most LLM outputs have no single correct string to compare against: "Refunds take five working days" and "You'll get your money back within 5 business days" are both right. LLM-as-a-judge uses a model to grade outputs against a written rubric — the list of criteria and what each score means.

What makes a judge trustworthy:

  • Discrete, anchored scales. "1 = contradicts the reference, 3 = partly correct, 5 = fully correct" is repeatable; "rate 0–100" is not.
  • Reasoning before the score. A model generates left to right; a score written first cannot be informed by reasoning written after it.
  • A reference answer and the retrieved context in the prompt, so correctness and groundedness are judged against evidence, not the judge's memory.
  • Calibration. Grade a sample by hand and report the judge's agreement with humans. Without that number, the scores are only a guess.

Known biases: position bias (in A-versus-B comparisons the first answer is favoured — run both orders), verbosity bias (longer answers score higher), and self-preference (a model rates its own family's style higher — use a different family as judge where you can).

Python
import json, refrom collections.abc import Callablefrom concurrent.futures import ThreadPoolExecutorfrom dataclasses import dataclassKEYS = ("correctness", "groundedness", "completeness")RUBRIC = """Grade the assistant's answer. Score each criterion from 1 to 5:- correctness: 1 contradicts the reference, 3 partly right, 5 fully consistent with it- groundedness: 1 claims not in the context, 5 every claim supported by the context- completeness: 1 misses the question, 5 answers every partWrite brief reasoning first, then the scores. Reply with JSON only:{"reasoning": "...", "correctness": 1, "groundedness": 1, "completeness": 1}"""@dataclassclass Case:    question: str    context: str    reference: strdef parse_verdict(raw: str) -> dict | None:    """Pull the JSON object out of the reply (tolerates Markdown code fences) and validate it."""    match = re.search(r"\{.*\}", raw, re.S)    try:        data = json.loads(match.group(0)) if match else None    except json.JSONDecodeError:        return None    if not isinstance(data, dict) or not all(data.get(k) in (1, 2, 3, 4, 5) for k in KEYS):        return None    return datadef judge(case: Case, answer: str, judge_fn: Callable[[str], str]) -> dict | None:    prompt = (f"{RUBRIC}\n\nQuestion: {case.question}\nContext:\n{case.context}\n"              f"Reference answer: {case.reference}\nAssistant answer: {answer}")    return parse_verdict(judge_fn(prompt)) or parse_verdict(judge_fn(prompt))   # one retrydef run_eval(cases: list[Case], system_fn: Callable, judge_fn: Callable, workers: int = 8):    with ThreadPoolExecutor(max_workers=workers) as ex:        answers = list(ex.map(lambda c: system_fn(c.question), cases))        verdicts = list(ex.map(lambda ca: judge(ca[0], ca[1], judge_fn), zip(cases, answers)))    good = [v for v in verdicts if v is not None]    report = {"n": len(cases), "unparsed": len(verdicts) - len(good)}    for k in KEYS:        report[k] = round(sum(v[k] for v in good) / len(good), 2) if good else None    return report, list(zip(cases, answers, verdicts))

The tricky parts:

  • re.search(r"\{.*\}", raw, re.S) grabs the JSON even when the judge wraps it in Markdown fences or adds a sentence before it. re.S lets . match newlines.
  • Validating the scores (in (1, 2, 3, 4, 5)) rejects "4", 4.5 or 7. A judge that returns something off the scale is treated as unparsed, not averaged.
  • Unparsed verdicts are counted, not averaged in as zero. A spike in unparsed means the judge prompt or model changed, which you must see.
  • ex.map keeps input order, so answers and verdicts line up with cases even though they run concurrently.

Complexity: 2 model calls per case (3 when a verdict needs the retry), run workers at a time, so wall-clock time is roughly (2 × cases ÷ workers) × call latency. Aggregation is O(cases × criteria).

A real-life example

Three cases, a stub system and a stub judge (one reply is fenced, one is broken):

Python
cases = [Case("Refund time?", "Refunds take 5 working days.", "5 working days"),         Case("COD refund?", "COD refunds go to store credit.", "Store credit"),         Case("Cancel fee?", "Cancelling is free before dispatch.", "Free before dispatch")]system = {"Refund time?": "5 working days.", "COD refund?": "To your bank account.",          "Cancel fee?": "Free before dispatch."}.getFENCE = "`" * 3                                     # a Markdown code fencereplies = iter([    FENCE + 'json\n{"reasoning": "matches", "correctness": 5, "groundedness": 5, "completeness": 5}\n' + FENCE,    '{"reasoning": "wrong destination", "correctness": 1, "groundedness": 1, "completeness": 4}',    "I think this is good.",                       # unparseable, retried...    "Still not JSON.",                             # ...and unparseable again])report, rows = run_eval(cases, system, lambda _: next(replies), workers=1)print(report)# {'n': 3, 'unparsed': 1, 'correctness': 3.0, 'groundedness': 3.0, 'completeness': 4.5}
caseanswerjudge replyverdict
Refund time?5 working days.fenced JSONparsed: 5, 5, 5
COD refund?To your bank account.JSONparsed: 1, 1, 4
Cancel fee?Free before dispatch.prose, then prose againunparsed (counted)

Averages use only the two parsed verdicts: correctness (5 + 1) / 2 = 3.0, completeness (5 + 4) / 2 = 4.5. The COD answer is complete but wrong — exactly the case a single "overall quality" score would blur.

A fintech team runs a suite like this on every prompt change, and blocks the release if average correctness drops by more than a few tenths.

Follow-up questions to expect

  • "How do you know the judge is right?" — Label 50–100 cases by hand, run the judge on them, and report agreement (for example Cohen's kappa). Recalibrate whenever the judge model or rubric changes.
  • "Pairwise or absolute scores?" — Pairwise ("is A or B better?") is more reliable for comparing two prompts; absolute scores are better for tracking one system over time. Randomise the order in pairwise mode.
  • "Can the judge be the same model as the system?" — It works, but self-preference inflates scores. Prefer a different model family, or at least a stronger model.