Course Content
LLM Evaluation
6 sections · 50 lessons
How do you measure bias in generated outputs?
What you need to know
Techniques
- Counterfactual pairs — the most actionable method, because it isolates the variable. For decisions, compare outcome rates. For generated text, compare length, sentiment, hedging, refusals, or judge scores.
- Benchmark probes — BBQ (question answering with ambiguous and clear contexts), WinoBias/WinoGender-style pronoun templates, StereoSet. Useful as smoke tests, but narrow, mostly English and US-centric, and possibly in training data.
- Production slicing — refusal rate, escalation rate and judge scores broken down by language, region or segment.
- Blind human rating — raters score outputs without knowing which variant they are reading.
Doing the statistics right
Each pair gives one of four results. Only discordant pairs — where the two versions got different outcomes — carry information. McNemar's test asks whether discordant pairs lean one way more than chance allows. With 50 pairs almost any gap is noise; with a few hundred, a gap of a few points can be real.
Choosing the attributes
Pick attributes that matter for your users and your law: in India, name-based signals of religion or caste, gender, region, and language or dialect (Hindi, Hinglish, regional English) are all realistic.
A real-life example
The bank's complaint classifier also assigns priority: urgent complaints get a call-back within 4 hours. The team suspects that complaints written in Hinglish are marked urgent less often. They take 300 real complaints, and for each one write a formal-English version and a Hinglish version with the same facts.
Results: 41% of formal versions are marked urgent, but only 33% of Hinglish versions. Looking at discordant pairs, 38 pairs were urgent only in English and 14 only in Hinglish. An exact McNemar test gives p ≈ 0.001, so the gap is real, not noise. The team adds Hinglish examples to the prompt and a rule that amounts and "unauthorised" events force urgent; the rerun shows 25 vs 21 discordant pairs, no longer significant. They keep the 300 pairs as a permanent regression test.
Follow-up questions to expect
- "Can you use an LLM judge to measure bias?" — Only after checking it on the same axis; the judge may share the bias. Test it on pairs where humans know the right answer.
- "What about bias in open-ended generation, like job descriptions?" — Generate for swapped attributes and compare measurable features: gendered words, adjectives used, length, and blind judge ratings.
- "Why not just use BBQ?" — It measures a general model tendency in English templates, not how your system behaves on your users' inputs and your task.