LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What is criteria ambiguity, and how does it affect evaluation results?


Two leads grading 100 reply draftsIs it empathetic?• Agree on 65 of 100• Cohen's kappa 0.25• One counts any apology• The other wants the issue namedTwo observable checks• Agree on 88 of 100• Cohen's kappa 0.75• First line names the specific issue• No sentence blames the customer
If two people cannot agree on a criterion, no judge model can match them; rewriting the rubric fixes what more data cannot.

What you need to know

How ambiguity damages evaluation

  • Low agreement — if two humans disagree on 35% of items, there is no stable "truth" for a judge to match.
  • Unstable judges — small wording changes in the rubric move the score, because the judge is filling in its own meaning.
  • Wrong decisions — a prompt change can look like a 5-point win on one run and a loss on the next.
  • Bias, not noise — more data narrows the interval around the rubric's hidden interpretation, not around the quality you meant.

Fixing it by decomposition

Turn a vague quality into observable sub-checks:

Vague criterionObservable checks
HelpfulAnswers the question asked; includes the requested detail; gives a next step
FaithfulEvery number, date and condition appears in the context
EmpatheticAcknowledges the customer's problem in the first sentence; no blaming language
ConciseUnder 120 words; no repeated information

Test the rubric before using it

  1. Two people label 50 to 100 items independently with the rubric.
  2. Compute kappa for each criterion.
  3. Read every disagreement — it usually points at one word ("relevant", "appropriate") that each person read differently.
  4. Rewrite and re-test until agreement is substantial.

If humans cannot agree on a criterion, a judge model will not either.

A real-life example

The bank's reply drafter is graded on "empathetic". Two customer-service leads label 100 drafts. They agree on only 65 — Cohen's kappa 0.25, barely above chance. One lead counted any apology as empathy; the other wanted the customer's specific problem acknowledged.

The team decomposes "empathetic" into two checks: "the first sentence names the customer's specific issue (for example, the failed ₹2,400 UPI payment)" and "no sentence blames the customer". On a new set of 100 drafts, the leads agree on 88 — kappa 0.75. The judge is then calibrated against these two checks, and its scores stop swinging between runs.

Follow-up questions to expect

  • "Isn't some subjectivity unavoidable?" — Yes, for tone and style. Then use pairwise comparison, several raters, and report agreement alongside the score so readers know how firm it is.
  • "Kappa or raw agreement?" — Kappa, because raw agreement is inflated when one label dominates.
  • "What if experts disagree because the domain is unclear?" — Write a decision rule for the case (as a labelling guideline), or split it into its own category; that disagreement is information about your product.