LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What are the different approaches to AI-as-a-judge systems?


Same 200 products, two ways to ask the judgePointwise, 1 to 5• Prompt A averages 4.1• Prompt B averages 4.2• Almost every description gets a 4• The real difference is squeezed flatPairwise, A vs B• B preferred in 116 of 200• Win rate 58%, interval 51 to 65• Sign test p about 0.03• Difference becomes visible
Relative judgements are easier than absolute ones, so pairwise shows a gap that a crowded 1-to-5 scale hides.

What you need to know

The four designs

Pointwise

  • One output, one rubric, one score
  • Works without a second candidate
  • Scores bunch at the top of the scale
  • Drifts between judge versions

Pairwise

  • Two outputs, "which is better?"
  • Relative judgements are easier and more consistent
  • Needs order swapping for position bias
  • Gives a win rate, not an absolute level

Reference-guided — the judge sees a gold answer and checks the candidate against it. Much better on facts and maths than judging from nothing. Needs labels, so mostly offline.

Decomposed — replace "rate quality 1 to 10" with several yes/no checks: "answers the question asked", "cites a policy", "no invented dates". Each check is easier to get right and easier to calibrate, and failures tell you exactly what went wrong.

Pointwise vs pairwise, with numbers

A team compares two prompts for the product-description generator on 200 products. Pointwise 1-to-5 scores: 4.1 for prompt A and 4.2 for prompt B. Almost every description got a 4, so a real difference is squeezed into 0.1 points.

Pairwise, the same judge prefers B in 116 of 200 pairs: a 58% win rate. The 95% interval is 51% to 65% (a sign test gives p ≈ 0.03), so B is probably better, though not by a lot. Pairwise made the difference visible.

Turning pairs into a ranking

With several candidates, count wins and fit a Bradley-Terry model — the method public arenas use. Comparing every pair of n candidates costs n × (n − 1) / 2 comparisons per test item; for many candidates, compare each against a fixed baseline or sample pairs.

Ensembles

Several judges vote — ideally from different model families. This reduces single-model bias at several times the cost. Useful for high-stakes decisions such as choosing a launch model.

A real-life example

The bank uses three judge designs for three jobs:

  • Complaint reply drafts in production — decomposed binary checks on 5% of drafts: states the next step? mentions the ticket number? promises no refund amount? no admission of liability? Each check has its own dashboard line, so when "promises refund" jumps from 0.3% to 2%, the team knows exactly what broke.
  • Choosing between two reply prompts — pairwise on 300 historical complaints, each pair judged in both orders.
  • Extracted fields from complaint letters (amount, date, account) — reference-guided against gold values, or simple exact match where possible; no judge needed for a date.

Follow-up questions to expect

  • "When is pointwise still the right choice?" — On live traffic, where there is no second candidate, and when criteria are binary and well defined.
  • "How do you handle ties in pairwise?" — Allow a "tie" answer, and treat inconsistent verdicts across swapped orders as ties; report win, tie and loss rates.
  • "Is pairwise transitive?" — Not always: judges can prefer A over B, B over C and C over A. Bradley-Terry fitting smooths this, and many intransitive triples are a sign the criteria are unclear.