Course Content
LLM Evaluation
6 sections · 50 lessons
What is criteria ambiguity, and how does it affect evaluation results?
What you need to know
How ambiguity damages evaluation
- Low agreement — if two humans disagree on 35% of items, there is no stable "truth" for a judge to match.
- Unstable judges — small wording changes in the rubric move the score, because the judge is filling in its own meaning.
- Wrong decisions — a prompt change can look like a 5-point win on one run and a loss on the next.
- Bias, not noise — more data narrows the interval around the rubric's hidden interpretation, not around the quality you meant.
Fixing it by decomposition
Turn a vague quality into observable sub-checks:
| Vague criterion | Observable checks |
|---|---|
| Helpful | Answers the question asked; includes the requested detail; gives a next step |
| Faithful | Every number, date and condition appears in the context |
| Empathetic | Acknowledges the customer's problem in the first sentence; no blaming language |
| Concise | Under 120 words; no repeated information |
Test the rubric before using it
- Two people label 50 to 100 items independently with the rubric.
- Compute kappa for each criterion.
- Read every disagreement — it usually points at one word ("relevant", "appropriate") that each person read differently.
- Rewrite and re-test until agreement is substantial.
If humans cannot agree on a criterion, a judge model will not either.
A real-life example
The bank's reply drafter is graded on "empathetic". Two customer-service leads label 100 drafts. They agree on only 65 — Cohen's kappa 0.25, barely above chance. One lead counted any apology as empathy; the other wanted the customer's specific problem acknowledged.
The team decomposes "empathetic" into two checks: "the first sentence names the customer's specific issue (for example, the failed ₹2,400 UPI payment)" and "no sentence blames the customer". On a new set of 100 drafts, the leads agree on 88 — kappa 0.75. The judge is then calibrated against these two checks, and its scores stop swinging between runs.
Follow-up questions to expect
- "Isn't some subjectivity unavoidable?" — Yes, for tone and style. Then use pairwise comparison, several raters, and report agreement alongside the score so readers know how firm it is.
- "Kappa or raw agreement?" — Kappa, because raw agreement is inflated when one label dominates.
- "What if experts disagree because the domain is unclear?" — Write a decision rule for the case (as a labelling guideline), or split it into its own category; that disagreement is information about your product.