LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What biases can AI judges introduce, and how can they be reduced?


From a 71% win to an honest one300 pairsjudged inone order: 71%Swap order:66 pairsflip, become ties234 consistentpairs: 60% winLength-matchedpairs only: 54%New descriptions averaged 142 words against 98.
Position and length bias together turned a modest gain into an apparent landslide.

What you need to know

Position bias

In pairwise judging, the judge tends to favour one slot regardless of content. The fix is cheap and standard:

  1. Judge (A, B) — record the winner.
  2. Judge (B, A) — same pair, swapped order.
  3. Combine — count a win only if both orders pick the same answer; otherwise it is a tie.
  4. Report the flip rate — the share of pairs whose verdict changed with order. It directly measures how noisy the judge is.

Verbosity (length) bias

Longer answers look more thorough. Mitigations: say in the rubric that length is not a criterion; check whether the "winning" variant simply writes more tokens; compare win rates within length buckets; or fit a regression of win on length difference, similar to the style control used by public arenas.

Self-preference

Judges tend to rate their own family's outputs higher, partly because that text is more "likely" to them. Use a judge from a different provider than the generator, or a panel from several families.

Style and authority

Markdown headings, bold text, confident tone and official-looking citations can win over plain correct answers. Decomposed factual checks — "is the cited clause real and does it say this?" — do not care about formatting.

Other cues

Hide which system produced which output: names such as "GPT answer" or "new prompt" in the judge prompt bias it.

A real-life example

The e-commerce team compares an old and a new description prompt on 300 products with a pairwise judge. Judged once, the new prompt wins 71%. Judged in both orders, 66 of 300 pairs (22%) flip verdict when swapped — those become ties. Of the remaining 234 consistent pairs, the new prompt wins 140: 60%, with a 95% interval of about 53% to 66%.

Then they check length: new descriptions average 142 words versus 98. Within pairs where the lengths differ by less than 15 words, the new prompt's win rate is 54% — barely above even. Much of the "win" was length. The team adds a 110-word limit to the new prompt and re-runs; the length-matched win rate is 57%, and product managers accept the change knowing the real gain is modest.

Follow-up questions to expect

  • "Can you remove position bias with a better prompt?" — Prompting helps a little; swapping order is the reliable fix, because it cancels the bias instead of hoping it goes away.
  • "How do you measure self-preference?" — Have judges from two families grade the same outputs from both families and compare how each rates its own family versus the other; or compare with human labels.
  • "Do humans have these biases?" — Yes, humans also prefer longer and more confident answers; that is why human evals need blinding, randomised order and clear rubrics too.