LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do you measure bias in generated outputs?


300 complaint pairs: formal English vs Hinglish853814163Hinglish: urgentHinglish: not urgentEnglish: urgentEnglish: not urgentOnly the shaded discordant pairs carry information: 38 vs 14, exact McNemar p about 0.001.
The 246 pairs that agree say nothing about bias; the test lives entirely in the two off-diagonal cells.

What you need to know

Techniques

  • Counterfactual pairs — the most actionable method, because it isolates the variable. For decisions, compare outcome rates. For generated text, compare length, sentiment, hedging, refusals, or judge scores.
  • Benchmark probes — BBQ (question answering with ambiguous and clear contexts), WinoBias/WinoGender-style pronoun templates, StereoSet. Useful as smoke tests, but narrow, mostly English and US-centric, and possibly in training data.
  • Production slicing — refusal rate, escalation rate and judge scores broken down by language, region or segment.
  • Blind human rating — raters score outputs without knowing which variant they are reading.

Doing the statistics right

Each pair gives one of four results. Only discordant pairs — where the two versions got different outcomes — carry information. McNemar's test asks whether discordant pairs lean one way more than chance allows. With 50 pairs almost any gap is noise; with a few hundred, a gap of a few points can be real.

Choosing the attributes

Pick attributes that matter for your users and your law: in India, name-based signals of religion or caste, gender, region, and language or dialect (Hindi, Hinglish, regional English) are all realistic.

A real-life example

The bank's complaint classifier also assigns priority: urgent complaints get a call-back within 4 hours. The team suspects that complaints written in Hinglish are marked urgent less often. They take 300 real complaints, and for each one write a formal-English version and a Hinglish version with the same facts.

Results: 41% of formal versions are marked urgent, but only 33% of Hinglish versions. Looking at discordant pairs, 38 pairs were urgent only in English and 14 only in Hinglish. An exact McNemar test gives p ≈ 0.001, so the gap is real, not noise. The team adds Hinglish examples to the prompt and a rule that amounts and "unauthorised" events force urgent; the rerun shows 25 vs 21 discordant pairs, no longer significant. They keep the 300 pairs as a permanent regression test.

Follow-up questions to expect

  • "Can you use an LLM judge to measure bias?" — Only after checking it on the same axis; the judge may share the bias. Test it on pairs where humans know the right answer.
  • "What about bias in open-ended generation, like job descriptions?" — Generate for swapped attributes and compare measurable features: gendered words, adjectives used, length, and blind judge ratings.
  • "Why not just use BBQ?" — It measures a general model tendency in English templates, not how your system behaves on your users' inputs and your task.