Course Content
LLM Evaluation
6 sections · 50 lessons
What are the different approaches to AI-as-a-judge systems?
What you need to know
The four designs
Pointwise
- One output, one rubric, one score
- Works without a second candidate
- Scores bunch at the top of the scale
- Drifts between judge versions
Pairwise
- Two outputs, "which is better?"
- Relative judgements are easier and more consistent
- Needs order swapping for position bias
- Gives a win rate, not an absolute level
Reference-guided — the judge sees a gold answer and checks the candidate against it. Much better on facts and maths than judging from nothing. Needs labels, so mostly offline.
Decomposed — replace "rate quality 1 to 10" with several yes/no checks: "answers the question asked", "cites a policy", "no invented dates". Each check is easier to get right and easier to calibrate, and failures tell you exactly what went wrong.
Pointwise vs pairwise, with numbers
A team compares two prompts for the product-description generator on 200 products. Pointwise 1-to-5 scores: 4.1 for prompt A and 4.2 for prompt B. Almost every description got a 4, so a real difference is squeezed into 0.1 points.
Pairwise, the same judge prefers B in 116 of 200 pairs: a 58% win rate. The 95% interval is 51% to 65% (a sign test gives p ≈ 0.03), so B is probably better, though not by a lot. Pairwise made the difference visible.
Turning pairs into a ranking
With several candidates, count wins and fit a Bradley-Terry model — the method public arenas use. Comparing every pair of n candidates costs n × (n − 1) / 2 comparisons per test item; for many candidates, compare each against a fixed baseline or sample pairs.
Ensembles
Several judges vote — ideally from different model families. This reduces single-model bias at several times the cost. Useful for high-stakes decisions such as choosing a launch model.
A real-life example
The bank uses three judge designs for three jobs:
- Complaint reply drafts in production — decomposed binary checks on 5% of drafts: states the next step? mentions the ticket number? promises no refund amount? no admission of liability? Each check has its own dashboard line, so when "promises refund" jumps from 0.3% to 2%, the team knows exactly what broke.
- Choosing between two reply prompts — pairwise on 300 historical complaints, each pair judged in both orders.
- Extracted fields from complaint letters (amount, date, account) — reference-guided against gold values, or simple exact match where possible; no judge needed for a date.
Follow-up questions to expect
- "When is pointwise still the right choice?" — On live traffic, where there is no second candidate, and when criteria are binary and well defined.
- "How do you handle ties in pairwise?" — Allow a "tie" answer, and treat inconsistent verdicts across swapped orders as ties; report win, tie and loss rates.
- "Is pairwise transitive?" — Not always: judges can prefer A over B, B over C and C over A. Bradley-Terry fitting smooths this, and many intransitive triples are a sign the criteria are unclear.