Course Content
LLM Evaluation
6 sections · 50 lessons
What biases can AI judges introduce, and how can they be reduced?
What you need to know
Position bias
In pairwise judging, the judge tends to favour one slot regardless of content. The fix is cheap and standard:
- Judge (A, B) — record the winner.
- Judge (B, A) — same pair, swapped order.
- Combine — count a win only if both orders pick the same answer; otherwise it is a tie.
- Report the flip rate — the share of pairs whose verdict changed with order. It directly measures how noisy the judge is.
Verbosity (length) bias
Longer answers look more thorough. Mitigations: say in the rubric that length is not a criterion; check whether the "winning" variant simply writes more tokens; compare win rates within length buckets; or fit a regression of win on length difference, similar to the style control used by public arenas.
Self-preference
Judges tend to rate their own family's outputs higher, partly because that text is more "likely" to them. Use a judge from a different provider than the generator, or a panel from several families.
Style and authority
Markdown headings, bold text, confident tone and official-looking citations can win over plain correct answers. Decomposed factual checks — "is the cited clause real and does it say this?" — do not care about formatting.
Other cues
Hide which system produced which output: names such as "GPT answer" or "new prompt" in the judge prompt bias it.
A real-life example
The e-commerce team compares an old and a new description prompt on 300 products with a pairwise judge. Judged once, the new prompt wins 71%. Judged in both orders, 66 of 300 pairs (22%) flip verdict when swapped — those become ties. Of the remaining 234 consistent pairs, the new prompt wins 140: 60%, with a 95% interval of about 53% to 66%.
Then they check length: new descriptions average 142 words versus 98. Within pairs where the lengths differ by less than 15 words, the new prompt's win rate is 54% — barely above even. Much of the "win" was length. The team adds a 110-word limit to the new prompt and re-runs; the length-matched win rate is 57%, and product managers accept the change knowing the real gain is modest.
Follow-up questions to expect
- "Can you remove position bias with a better prompt?" — Prompting helps a little; swapping order is the reliable fix, because it cancels the bias instead of hoping it goes away.
- "How do you measure self-preference?" — Have judges from two families grade the same outputs from both families and compare how each rates its own family versus the other; or compare with human labels.
- "Do humans have these biases?" — Yes, humans also prefer longer and more confident answers; that is why human evals need blinding, randomised order and clear rubrics too.