Course Content
LLM Evaluation
6 sections · 50 lessons
How do you design an effective evaluation prompt for an AI judge?
What you need to know
The ingredients
- Role and task — "You check whether an HR assistant's answer is supported by the policy text."
- One criterion — "Faithfulness: every claim appears in, or follows directly from, the context." Separate prompts for relevance, tone, safety.
- A small scale with defined levels — binary pass/fail, or pass / partial / fail, each defined: "fail if any number, date or condition is not in the context".
- Examples — 2 to 4 graded examples, including one borderline case and why it was graded that way. Examples raise agreement more than any other single change.
- What to ignore — "Do not reward length, formatting or politeness."
- Output format — JSON with reasoning first, then the verdict.
- An escape hatch — "If the context is missing or unreadable, return
cannot_judge."
{"reasoning": "The answer states a 30-day return window; the context says 45 days.", "unsupported_claims": ["30-day return window"], "verdict": "fail"}G-Eval style prompts
G-Eval (2023) asks the judge to first write evaluation steps from the criteria, then score. DeepEval's GEval metric implements this: you give criteria or evaluation steps in plain language, and it produces a score between 0 and 1 with a reason. Useful to start quickly; still calibrate it.
Settings
- Pin the exact model version; judges change when providers update models.
- Use temperature 0 where the model allows it; otherwise run twice and check consistency.
- Keep inputs small: only the context the criterion needs.
Iterate like a prompt engineer
- Label — 100 to 200 real outputs graded by a domain expert.
- Run the judge — on the same items.
- Read disagreements — each one shows a missing rule or a bad example.
- Revise — the rubric or examples, not the labels.
- Re-measure — on a separate held-out slice so you don't overfit the judge prompt.
A real-life example
The code-review bot's team needs a judge for "is this review comment actionable?". Version 1 asks "Is this comment helpful? Score 1-10." Against 150 comments labelled by two senior engineers, it rates almost everything 7 or 8.
Version 2 defines actionable as: names a specific line or function, states a concrete problem, and suggests a change or a test. Pass only if all three. It includes three examples — one passing, one vague ("consider improving error handling"), and one borderline (correct problem, no suggestion — fail). Output is JSON with reasoning first. Kappa with the engineers rises from 0.21 to 0.71. The disagreements left are mostly comments on test files, so the team adds one more example for tests.
Follow-up questions to expect
- "Why binary instead of 1 to 10?" — Humans and models cannot consistently separate adjacent points on a long scale; binary decisions are more reliable and easier to calibrate.
- "Should the judge see the reference answer?" — Yes, when you have one; reference-guided judging is more accurate on factual tasks.
- "How do you stop overfitting the judge prompt to the calibration set?" — Tune on one part of the labelled data and report agreement on a held-out part.