AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How can bias appear in AI systems, and how do you detect it?


Screening log for one software role100030030%1.0080018022.5%0.75AppliedShortlistedRateImpact ratioMenWomenA name-swap test on 200 identical CVs then showed a 6-point score drop for female names.
An impact ratio below 0.8 is the four-fifths warning sign, and the counterfactual test is what tells you the model, not the pool, caused it.

What you need to know

Where bias enters

StageExample
Historical dataPast hiring decisions favoured men, so "hired" labels encode that preference.
RepresentationA skin-condition model trained mostly on light skin.
LabellingAnnotators rate Indian English or African American English as "less professional".
ModelOptimising average loss sacrifices the small group.
DeploymentA model validated in metro cities is used in rural districts.
LLM generationStereotyped completions ("the nurse... she"), different tone or advice for different names.

Fairness metrics (they conflict)

  • Demographic parity / selection rate: does each group get the positive outcome at a similar rate? The impact ratio divides each group's rate by the highest group's rate. US hiring practice uses the "four-fifths rule": a ratio below 0.8 is a warning sign. New York City's rules for automated hiring tools require publishing impact ratios from an annual bias audit.
  • Equal opportunity: among people who deserve the positive outcome, is the true-positive rate similar?
  • Equalised odds: similar true-positive and false-positive rates.
  • Calibration within groups: a score of 0.7 means 70% for every group.

When base rates differ between groups, you mathematically cannot satisfy all of these at once. Pick the one that matches the harm and explain why.

Python
from collections import defaultdictdef impact_ratios(rows):    shown, passed = defaultdict(int), defaultdict(int)    for group, shortlisted in rows:        shown[group] += 1        passed[group] += shortlisted    rates = {g: passed[g] / shown[g] for g in shown}    best = max(rates.values())    return {g: round(r / best, 2) for g, r in rates.items()}rows = ([("men", 1)] * 300 + [("men", 0)] * 700 +        [("women", 1)] * 180 + [("women", 0)] * 620)print(impact_ratios(rows))   # {'men': 1.0, 'women': 0.75} -> below 0.8

Counterfactual testing for LLMs

Hold the input fixed and change one attribute: the candidate's name (Rahul / Priya / Mohammed / Joseph), a pronoun, a college, a dialect. Run each pair many times and measure the change in score or decision. For LLM systems this is the highest-signal test you can automate.

Proxies

Dropping the gender column does not remove gender from the data. Names, colleges, career gaps and hobbies carry it. Removing the attribute only removes your ability to measure the bias.

A real-life example

An HR team uses an LLM screening assistant to shortlist 1,800 applicants for a software role. The team exports the decisions: 30% of men and 22.5% of women are shortlisted, an impact ratio of 0.75. A counterfactual test swaps only the name on 200 identical CVs; the average score drops 6 points for female names. Digging in shows the prompt asks the model to "prefer candidates with continuous experience", which penalises maternity career breaks.

The fix: remove the continuity criterion, strip names and photos before scoring, score against a written rubric of job-related skills, and add the counterfactual test to CI with a maximum allowed score gap. Amazon scrapped an experimental recruiting model for a similar reason in 2018, after it learned to penalise CVs that mentioned "women's".

Follow-up questions to expect

  • "Which fairness metric would you choose?" — Depends on the harm. For hiring shortlists, selection-rate ratio is what auditors and regulators ask for; for a fraud flag, a false-positive-rate gap matters more because false flags hurt innocent people.
  • "What if you don't have group labels?" — Collect them for measurement only, with consent and access control, or use carefully validated proxies for auditing. Without some labels you cannot prove anything.
  • "Is an LLM less biased than a trained classifier?" — Not by default. It carries biases from web text, and they show up in tone and advice as well as decisions.