Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Grading outputs: code checks, rubrics, model-as-judge and its biases
An eval set gives you inputs and expected outputs. A grader decides whether an actual output matches. For many LLM features this is where evaluation quietly goes wrong: a team builds a careful eval set and then grades it with a single prompt that asks another model "is this answer good? Score 1 to 10."
The TiffinGo draft has two very different kinds of content. The action, lines, amount and escalation code are structured, and have exactly one right answer. The reason is free text written for an agent, where many answers are acceptable. Each needs a different kind of grader.
The principle is simple: use the cheapest, most reliable grader that can judge each part. That means code first, a model judge only where code cannot reach, and the judge itself tested before you trust it.
Code checks: most of the grade
Anything with one right answer is graded by code. It is free, instant, and gives the same result every time.
1def grade_decision(expected: dict, draft: dict, tolerance_inr: int = 2) -> dict:2 def line_keys(lines):3 return {(line["item"], line["units"], line["rule"]) for line in lines}45 checks = {6 "action": draft["action"] == expected["action"],7 "escalation_code": draft["escalation_code"] == expected["escalation_code"],8 "lines": line_keys(draft["lines"]) == line_keys(expected["lines"]),9 "amount": abs(draft["refund_inr"] - expected["refund_inr"]) <= tolerance_inr,10 }11 checks["pass"] = all(checks.values())12 return checksA draft passes only if the action, escalation code, set of affected lines and amount are all right. The ₹2 tolerance allows for rounding differences such as 30% of ₹185, which is ₹55.50; the policy says round to nearest, and both ₹55 and ₹56 have appeared in agents' own work.
Keeping each check separate matters for the next lesson. "Failed on amount but lines were right" points to arithmetic. "Failed on lines" points to reading the complaint. A single pass/fail would hide the difference.
Rubrics for free text
For the reason, write a rubric: a short list of yes/no criteria, each about one property. Avoid 1-to-10 scales. Nobody, human or model, can reliably tell a 6 from a 7, and the scores drift from day to day.
| Criterion | Question | Graded by |
|---|---|---|
| Length | Is it at most 40 words? | Code |
| Names items | Does it name every item in the lines, and no others? | Model judge |
| Shows the arithmetic | Does each refunded line show how its amount was reached? | Model judge |
| Only known facts | Is every claim supported by the ticket, the order or the lines? | Model judge, plus a code check on numbers |
Even here, code does what it can. Length is a word count. For "only known facts", code can extract every number in the reason and check that each one appears in the order or in the draft's own lines. A reason that mentions "₹250" when no such amount exists fails without asking a model.
A model as judge
The remaining criteria need language understanding, so a second model call grades them. The judge gets the same inputs, the draft, and the rubric, and must return a verdict per criterion with the words that decided it.
You check a short reason written for a support agent. For each question, answertrue or false and quote the words from the reason that decided it.names_items: Does the reason name every item in draft_lines, and no other item?shows_math: For each refunded line, does the reason show how the amount was reached, for example "30% of 220 = 66"?only_known_facts: Is every fact in the reason supported by ticket_text, order_items or draft_lines? A claim such as "the rider was rude" with no support in the input makes this false.If you are unsure about a question, answer false.1import json2import llm34def judge_reason(ticket_text: str, order: dict, draft: dict, judge_model: str) -> dict:5 user = json.dumps({6 "ticket_text": ticket_text,7 "order_items": order["items"],8 "draft_lines": draft["lines"],9 "reason": draft["reason"],10 }, ensure_ascii=False)11 reply = llm.complete(JUDGE_SYSTEM, user, model=judge_model,12 schema=JUDGE_SCHEMA, max_tokens=1500)13 return json.loads(reply.text)JUDGE_SYSTEM is the rubric text above, and JUDGE_SCHEMA requires a verdict boolean and an evidence string for each of the three criteria. Asking for evidence does two things. It makes the judge's verdicts more careful, and it lets you audit a surprising verdict in seconds.
The judge's biases, and what to do about them
A model judge is a measuring instrument with known faults. Know them before you trust its numbers.
- Leniency. Judges tend to say yes. "If you are unsure, answer false" and binary criteria push back against it.
- Verbosity. Longer answers tend to be rated better, even when the rubric does not reward length. A hard word limit graded by code removes the incentive.
- Self-preference. A judge tends to favour text from its own model family. Where you can, use a judge from a different family than the model being graded, or at least check for the effect.
- Position. When asked to compare two outputs, judges favour one position, often the first. For pairwise comparisons, run both orders and count only consistent verdicts.
- Prompt sensitivity. Small wording changes in the rubric move scores. Version the judge prompt like any other release.
The only real defence is calibration: compare the judge with humans. Two agents graded the reasons of 60 drafts against the same rubric. The judge agreed with them 97% of the time on "names items" and 93% on "shows the arithmetic", but only 81% on "only known facts". Most disagreements were the judge accepting reasonable-sounding inferences, like "customer is upset", as supported.
At 81%, that criterion was too unreliable to gate releases on. The team tightened its wording, added the code check on numbers, and re-measured at 90%. Recalibrate whenever you change the judge prompt or the judge model.
Code checks
- Action, lines, amount, escalation code, length, numbers in the reason
- Free, instant, identical every run
- Should carry most of the pass/fail decision
Model judge
- Named items, shown arithmetic, invented facts
- Costs a call per case, varies slightly between runs
- Trusted only after measuring agreement with humans
In TiffinGo's gate, only the code checks decide whether a draft passes. Judge scores are reported beside them and trigger a review when they fall, but a reason problem on its own never blocks a release. That keeps the most reliable grader in charge of the decisions that move money.
Check your understanding
0 of 3 answered
1.Why does TiffinGo grade action, lines and amount separately instead of one pass/fail?
2.The judge agrees with humans 81% of the time on "only known facts". What should the team do?
3.You compare two releases by asking a judge "which reason is better, A or B?" What protects against position bias?