Course Content
LLM Evaluation
6 sections · 50 lessons
What methods can be used to assess the quality of LLM outputs?
What you need to know
An LLM output can be wrong in many ways at once: wrong facts, wrong format, wrong tone, unsafe content, too slow, too expensive. No single metric sees all of these, so you combine methods. The skill an interviewer is testing is knowing which method fits which failure.
The four families
| Family | What it checks | Cost | Weak spot |
|---|---|---|---|
| Deterministic checks | JSON parses, required fields present, label in allowed set, code passes tests, right tool called | Almost free | Cannot judge meaning or tone |
| Reference-based metrics | Output vs a gold answer: exact match, token F1, BLEU, ROUGE, BERTScore | Cheap once labels exist | Penalises correct answers worded differently |
| LLM-as-judge | A model scores the output against a rubric, or picks the better of two | A model call per item | Has its own biases; must be calibrated |
| Human evaluation | Trained people grade a sample | Slow and expensive | Does not scale to every change |
Why layering works
Order the checks from cheapest to most expensive. If a reply is not valid JSON, there is no point paying a judge to rate its tone. Deterministic checks usually remove a big share of failures before any model is called. The judge then handles meaning, and humans handle the small set where the judge is unsure or where you need to prove the judge is still right.
Two things to decide before choosing a method
- Is there one correct answer? A complaint label or an extracted amount has one; a product description does not. One correct answer means exact match or F1. Many good answers means a rubric or pairwise comparison.
- Do you have labels? Reference-based metrics need gold answers. On live traffic you usually have none, so you need reference-free methods: deterministic checks, judges, and user signals.
A real-life example
An e-commerce company generates 5,000 product descriptions a day from seller spec sheets. The team builds a layered eval:
- Deterministic checks on 100% — 80 to 150 words, every spec attribute (material, size, warranty) appears, no banned claims such as "100% waterproof" unless the spec says so. About 6% fail and are regenerated.
- LLM judge on a 10% sample — two yes/no questions: "Is every claim supported by the spec sheet?" and "Does the tone match the brand guide?"
- Human review of 50 a week — a copy editor grades the same two questions, and the team compares her labels with the judge's to check that the judge still agrees.
When the judge started passing descriptions that said "leather" for a "faux leather" bag, the weekly human sample caught it. The team added a deterministic rule that material words must match the spec exactly, moving that failure down to the cheapest layer.
Follow-up questions to expect
- "Why not just use human review?" — It is the most trusted signal, but at a few minutes per item it cannot run on every pull request or every production reply. Humans are best used to label a calibration set and to review disagreements.
- "Which layer would you build first?" — Deterministic checks and a small labelled set. They are cheap, they catch the most embarrassing failures, and every later layer is measured against them.
- "How do you know the LLM judge is right?" — Label 100 to 200 items by hand, run the judge on the same items, and measure agreement with Cohen's kappa and the judge's precision and recall on the failure class.