LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What methods can be used to assess the quality of LLM outputs?


Cheapest check first, humans lastRules on 100%: length, attributes, banned claimsReference metrics where a gold answer existsLLM judge on a 10% sample: two yes/no questionsHuman review of 50 a week, to check the judge
The faux-leather error slipped past the judge and was caught by humans, then moved down to a one-line rule.

What you need to know

An LLM output can be wrong in many ways at once: wrong facts, wrong format, wrong tone, unsafe content, too slow, too expensive. No single metric sees all of these, so you combine methods. The skill an interviewer is testing is knowing which method fits which failure.

The four families

FamilyWhat it checksCostWeak spot
Deterministic checksJSON parses, required fields present, label in allowed set, code passes tests, right tool calledAlmost freeCannot judge meaning or tone
Reference-based metricsOutput vs a gold answer: exact match, token F1, BLEU, ROUGE, BERTScoreCheap once labels existPenalises correct answers worded differently
LLM-as-judgeA model scores the output against a rubric, or picks the better of twoA model call per itemHas its own biases; must be calibrated
Human evaluationTrained people grade a sampleSlow and expensiveDoes not scale to every change

Why layering works

Order the checks from cheapest to most expensive. If a reply is not valid JSON, there is no point paying a judge to rate its tone. Deterministic checks usually remove a big share of failures before any model is called. The judge then handles meaning, and humans handle the small set where the judge is unsure or where you need to prove the judge is still right.

Two things to decide before choosing a method

  • Is there one correct answer? A complaint label or an extracted amount has one; a product description does not. One correct answer means exact match or F1. Many good answers means a rubric or pairwise comparison.
  • Do you have labels? Reference-based metrics need gold answers. On live traffic you usually have none, so you need reference-free methods: deterministic checks, judges, and user signals.

A real-life example

An e-commerce company generates 5,000 product descriptions a day from seller spec sheets. The team builds a layered eval:

  1. Deterministic checks on 100% — 80 to 150 words, every spec attribute (material, size, warranty) appears, no banned claims such as "100% waterproof" unless the spec says so. About 6% fail and are regenerated.
  2. LLM judge on a 10% sample — two yes/no questions: "Is every claim supported by the spec sheet?" and "Does the tone match the brand guide?"
  3. Human review of 50 a week — a copy editor grades the same two questions, and the team compares her labels with the judge's to check that the judge still agrees.

When the judge started passing descriptions that said "leather" for a "faux leather" bag, the weekly human sample caught it. The team added a deterministic rule that material words must match the spec exactly, moving that failure down to the cheapest layer.

Follow-up questions to expect

  • "Why not just use human review?" — It is the most trusted signal, but at a few minutes per item it cannot run on every pull request or every production reply. Humans are best used to label a calibration set and to review disagreements.
  • "Which layer would you build first?" — Deterministic checks and a small labelled set. They are cheap, they catch the most embarrassing failures, and every later layer is measured against them.
  • "How do you know the LLM judge is right?" — Label 100 to 200 items by hand, run the judge on the same items, and measure agreement with Cohen's kappa and the judge's precision and recall on the failure class.