AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do you ensure fairness across intersectional groups?


Triage accuracy by gender and age95%65%80%, n=1590%18-4950+WomenMenBy gender alone: women 90%, men 89% — no visible gap.
The worst group hides at the intersection, and the thinnest cell is too small to judge at all.

What you need to know

Why marginal metrics hide problems

Python
from collections import defaultdictdef slice_report(rows, keys, min_n=30):    stats = defaultdict(lambda: [0, 0])    for r in rows:        k = tuple(r[x] for x in keys)        stats[k][0] += 1        stats[k][1] += r["correct"]    for k, (n, ok) in sorted(stats.items(), key=lambda kv: kv[1][1] / kv[1][0]):        flag = "  (too few to judge)" if n < min_n else ""        print(k, n, f"{ok / n:.0%}{flag}")rows = ([{"gender": "F", "age": "50+", "correct": 1}] * 52 +        [{"gender": "F", "age": "50+", "correct": 0}] * 28 +        [{"gender": "F", "age": "18-49", "correct": 1}] * 380 +        [{"gender": "F", "age": "18-49", "correct": 0}] * 20 +        [{"gender": "M", "age": "50+", "correct": 1}] * 180 +        [{"gender": "M", "age": "50+", "correct": 0}] * 20 +        [{"gender": "M", "age": "18-49", "correct": 1}] * 12 +        [{"gender": "M", "age": "18-49", "correct": 0}] * 3)slice_report(rows, ["gender", "age"])# ('F', '50+') 80 65%# ('M', '18-49') 15 80%  (too few to judge)# ('M', '50+') 200 90%# ('F', '18-49') 400 95%

By gender alone, women score 90% and men 89% — no visible gap. The intersection shows women over 50 at 65%. The report also warns that men aged 18–49 have only 15 examples, too few for a conclusion.

The statistical problems

  • Small samples: intersections get thin fast. Report confidence intervals, and do not act on n = 15.
  • Multiple comparisons: testing 60 slices means a few will look bad by chance. Use a correction or confirm on fresh data.

Practices

  • Define slices from a harm model before looking at results.
  • Oversample thin slices in eval sets. An eval set does not need to match production traffic; it needs statistical power where harm is likely.
  • Worst-group metric as a release gate, not the average.
  • Training methods: group distributionally robust optimisation (group DRO), reweighting, targeted data collection.
  • Find unnamed slices by clustering errors; the worst group is often one nobody listed.

A real-life example

A healthcare symptom-checker is evaluated for triage accuracy. By gender: 90% and 89%. By age: 83% and 94%. It looks like a mild age issue. The intersectional report shows women over 50 at 65%: the model often labels their heart-attack symptoms (fatigue, nausea, jaw pain) as "digestive", because typical descriptions in its training data follow the classic male presentation.

The team adds 400 cases for women over 50, written and reviewed by cardiologists, adds an explicit red-flag rule for atypical cardiac symptoms in this group, and gates releases on the worst-group triage accuracy being at least 90%. After the change, the group reaches 91%.

Follow-up questions to expect

  • "How many slices can you realistically track?" — Choose the combinations your harm model says matter, usually two attributes at a time, plus error clustering to find surprises.
  • "What if you don't have the attributes?" — Collect them for a consented evaluation sample, or use a carefully reviewed labelled set; you cannot measure what you never record.
  • "Doesn't optimising the worst group lower overall accuracy?" — Sometimes slightly. Report both, and make the trade-off an explicit decision.