Course Content
AI Ethics and Governance
3 sections · 7 lessons
Mini-Project: Bias Audit on Fine-Tuned Model
You are the data scientist on a governance team. A loan approval model is queued for deployment and your job is to sign it off, or not. You run the obvious checks and get this:
Test set: 1,500 applicantsOverall accuracy 0.804ROC-AUC 0.88Precision, group M 0.808Precision, group F 0.800 <- a 0.8-point differencePrecision is the fraction of approvals that were correct. It is nearly identical across groups: when this model says yes, it is right about 80% of the time for everyone. Accuracy is respectable. Ranking quality is good. On this evidence you would sign.
Now disaggregate — split every number by group instead of averaging over them:
group M group Fapplicants 900 600would have repaid 450 240approved by model 468 180true positive rate 0.840 0.600false negative rate 0.160 0.400Among applicants who would have repaid, the model declines 16% of group M and 40% of group F. A qualified applicant in group F is 2.5 times as likely to be declined as an equally qualified applicant in group M. That is 96 people out of 240 in a single test set, and none of it was visible in the aggregate.
This is what a bias audit is for, and why the first rule is that every number gets computed per group and never averaged across them. What follows is a build brief: four stages, real metrics, and a description of what a finding actually looks like when you write one down.
What you are building
Four artefacts, roughly three hours of work:
- A test set with known ground truth — synthetic data into which you have deliberately injected specific, documented biases, so that you can check whether your audit detects the things you know are there.
- A disaggregated metrics table, computed by hand from confusion matrices and then cross-checked against a maintained library.
- A mechanism diagnosis — not "the model is biased" but which of the four possible causes produced each gap.
- An audit report whose findings a non-technical reader can act on.
Injecting bias on purpose is the part people skip, and it is the part that makes the exercise honest. If you audit a model with unknown bias and find nothing, you cannot tell whether the model is clean or your audit is blind. With known ground truth, a null result is a bug report about your own code.
Stage 1: a test set whose biases you already know
Build 5,000 applicants with three separate mechanisms planted in them. Public alternatives exist — Adult Income, German Credit, COMPAS — and you should run against one afterwards, but start synthetic so you know the answer.
1import numpy as np, pandas as pd23RNG = np.random.default_rng(42)45def build_dataset(n=5000):6 group = RNG.choice(["M", "F"], n, p=[0.60, 0.40]) # imbalance: about 3000 / 20007 credit = np.clip(RNG.normal(680, 80, n), 300, 850)8 income = RNG.lognormal(10.5, 0.7, n)9 debt = RNG.beta(2, 5, n)1011 # MECHANISM 2 -- proxy. postcode_band carries group membership, and it12 # survives after `group` itself is dropped from the feature set.13 postcode = RNG.normal(np.where(group == "M", 0.0, 1.1), 1.0, n)1415 # Genuine creditworthiness: what a fair model should be predicting.16 merit = (credit - 680) / 80 + 0.6 * (np.log(income) - 10.5) - 1.4 * debt1718 # MECHANISM 1 -- historical label bias. At equal merit the old process19 # approved group M more often. +0.55 on the logit is roughly20 # +10 percentage points near the middle of the distribution.21 logit = 1.1 * merit + np.where(group == "M", 0.55, 0.0) - 0.1022 approved = (RNG.random(n) < 1 / (1 + np.exp(-logit))).astype(int)2324 return pd.DataFrame({"group": group, "credit_score": credit, "income": income,25 "debt_ratio": debt, "postcode_band": postcode,26 "merit": merit, "approved": approved})2728df = build_dataset()29print(df.groupby("group")["approved"].agg(["size", "mean"]))30# size mean31# group32# F 1951 0.39056933# M 3049 0.504100Three mechanisms are now in the data, and you should be able to say out loud what each one will do to which metric:
- Label bias — the historical outcome itself favours group M. A model that fits this label perfectly will reproduce a 10-point base-rate gap and call it accuracy.
- Proxy leakage —
postcode_bandpredicts group membership. Dropping thegroupcolumn removes it from the schema, not from the model. - Representation imbalance — group F is 40% of applicants, and any intersectional cell you slice will be smaller still.
Then stratify. Split so that both groups keep their proportions in train and test, and print the cell sizes before you go any further, because a metric on a cell of 40 is noise wearing a decimal point.
1from sklearn.model_selection import train_test_split23train, test = train_test_split(df, test_size=0.30, stratify=df[["group", "approved"]],4 random_state=42)56cells = test.groupby(["group", "approved"]).size().unstack()7print(cells) # F: 357 declined / 229 approved, M: 453 / 4618print("smallest cell:", int(cells.values.min()))Finally, carve out the scenarios most likely to expose a threshold problem, and check the group mix inside each one: high income with a low credit score, low income with a high credit score, thin files (under two years of employment history), and records with a missing field. The last is not decoration — differential missingness is a proxy in disguise, because a field tends to be blank exactly for the people the old process never handled.
Stage 2: compute the numbers yourself, then check them
Write the metric code by hand once. Libraries are correct and you should use them in production, but a metric you have never derived is a metric you cannot debug when it disagrees with someone else's.
1from sklearn.metrics import confusion_matrix, roc_auc_score23def per_group(y_true, y_pred, y_score, groups):4 rows = []5 for g in sorted(set(groups)):6 m = groups == g7 # labels=[0, 1] is load-bearing: without it a group containing only8 # one class returns a 1x1 matrix and .ravel() raises on unpacking.9 tn, fp, fn, tp = confusion_matrix(y_true[m], y_pred[m], labels=[0, 1]).ravel()10 n = int(m.sum())11 rows.append({"group": g, "n": n,12 "base_rate": (tp + fn) / n,13 "selection_rate": (tp + fp) / n,14 "tpr": tp / (tp + fn) if tp + fn else np.nan,15 "fpr": fp / (fp + tn) if fp + tn else np.nan,16 "ppv": tp / (tp + fp) if tp + fp else np.nan,17 "accuracy": (tp + tn) / n,18 "auc": roc_auc_score(y_true[m], y_score[m])})19 return pd.DataFrame(rows).set_index("group")Train a logistic regression on the standardised features with group dropped, predict at the default 0.5 threshold, and you get a table like this one — the table the rest of the audit turns on. Its figures are a worked illustration, rounded so the arithmetic is easy to follow.
| Metric | Group M (n=900) | Group F (n=600) | Gap or ratio | Reading |
|---|---|---|---|---|
| Base rate (label) | 0.500 | 0.400 | 10.0 pts | The injected label bias, intact |
| Selection rate | 0.520 | 0.300 | ratio 0.58 | Fails the 0.80 four-fifths rule |
| True positive rate | 0.840 | 0.600 | 24.0 pts | Fails equal opportunity badly |
| False positive rate | 0.200 | 0.100 | 10.0 pts | Model is more cautious with group F |
| Precision (PPV) | 0.808 | 0.800 | 0.8 pts | Passes — and this is the trap |
| Accuracy | 0.820 | 0.780 | 4.0 pts | Passes a 5-point bound |
| ROC-AUC | 0.890 | 0.830 | 6.0 pts | The model ranks group F less well |
Two rows deserve a hard look. Precision passing is not reassurance — it says the model is equally right when it approves someone, which is a statement about the people it said yes to and tells you nothing about the people it said no to. And the false positive rate being lower for group F is not a kindness: paired with the low TPR it is the signature of an effectively higher decision threshold for that group.
Precision measures the quality of the approvals. True positive rate measures who was missed. An audit that reports only the first is an audit of the people the model already liked.
Now put an interval on the headline gap, because a 24-point difference on 240 positives is not the same claim as a 24-point difference on 24.
1def bootstrap_tpr_gap(y_true, y_pred, groups, a="M", b="F", n_boot=2000, seed=0):2 rng, n, gaps = np.random.default_rng(seed), len(y_true), []3 for _ in range(n_boot):4 s = rng.integers(0, n, n) # resample with replacement5 rate = {}6 for g in (a, b):7 pos = (groups[s] == g) & (y_true[s] == 1)8 rate[g] = y_pred[s][pos].mean() if pos.sum() else np.nan9 gaps.append(rate[a] - rate[b])10 lo, hi = np.nanpercentile(gaps, [2.5, 97.5])11 return float(np.nanmean(gaps)), float(lo), float(hi)1213# TPR gap, M minus F: 0.240 95% CI [0.169, 0.311]The interval excludes zero by a wide margin, so this is a finding rather than noise. Run the same function on an intersectional cell and watch what happens: group F aged 60 and over has 58 test applicants, 22 of whom would have repaid, and 9 of those were approved. The point estimate is a TPR of 0.41; the 95% interval is [0.23, 0.61], which comfortably contains both group TPRs from the table above. The honest report of that cell is "we cannot say", and that sentence is itself a finding: you do not have enough data to audit older female applicants at all.
Finally, cross-check against a maintained implementation. If the two disagree by more than rounding, you have a bug, and the usual culprit is the unpacking order of confusion_matrix(...).ravel(), which returns tn, fp, fn, tp — not the order most people write from memory.
1from fairlearn.metrics import (MetricFrame, selection_rate,2 true_positive_rate, false_positive_rate)34mf = MetricFrame(metrics={"selection_rate": selection_rate,5 "tpr": true_positive_rate, "fpr": false_positive_rate},6 y_true=y_test, y_pred=y_pred, sensitive_features=g_test)7print(mf.by_group)8print("selection-rate ratio:", mf.ratio()["selection_rate"]) # 0.577Reading the table: which gap means what
"The model is biased" is not a finding, because it does not imply an action. The four patterns below do, and each points at a different repair.
| Pattern in the numbers | What it means | What fixes it | What will not |
|---|---|---|---|
| Per-group AUC equal, TPR unequal at one threshold | Ranking is equally good; the single global cutoff falls in a different place on each group's score distribution | Threshold placement, or recalibration per group where that is lawful | More data — the model already ranks correctly |
| Per-group AUC unequal (here 0.89 vs 0.83) | The model genuinely knows less about one group; fewer rows, weaker features, or more label noise there | More and better data for that group; features that carry signal for it | Threshold tuning, which just moves errors between the two kinds |
| TPR and FPR equal, selection rates unequal | Base rates genuinely differ; the model is treating like cases alike | A decision about whether parity of outcome is the right target here, argued explicitly | Forcing equal selection rates, which now costs accuracy for real reasons |
| Precision equal, TPR unequal, base rates unequal | The classic label-bias signature: the model faithfully reproduced a biased historical target | Changing what the label measures | Any reweighting or threshold work, which rebalances the wrong measurement |
Our model shows the fourth pattern and the second at once, which is common and worth stating in the report as two separate findings with two separate owners.
A gap is a symptom. Until you know whether the model ranks the group worse or merely cuts it at a different point, you cannot tell which of the available repairs would do anything at all.
Stage 3: find the mechanism
Three checks, in this order, because each one changes how you interpret the next.
1# 1. Direct use. Fit once WITH the group column to see what it is worth.2coef = dict(zip(feature_names, model_with_group.coef_[0]))3print(f"group coefficient {coef['group_M']:+.3f} "4 f"-> odds ratio {np.exp(coef['group_M']):.2f}")5# group coefficient +0.693 -> odds ratio 2.0067# 2. Proxy leakage. Can the group be recovered from the shipped features?8from sklearn.ensemble import RandomForestClassifier9from sklearn.inspection import permutation_importance10probe = RandomForestClassifier(n_estimators=300, min_samples_leaf=20).fit(Xtr, gtr)11print("probe AUC:", roc_auc_score(gte, probe.predict_proba(Xte)[:, 1])) # 0.7812imp = permutation_importance(probe, Xte, gte, scoring="roc_auc", n_repeats=10)13# postcode_band drop in AUC = 0.190 <- the leak14# income drop in AUC = 0.021Read these together. An odds ratio of 2.00 means that with the group column present, being in group M doubles the odds of approval at identical credit score, income and debt ratio — direct use, severity high. Dropping that column removes it, but a probe AUC of 0.78 says the shipped feature set still recovers group membership, driven almost entirely by postcode_band. Below about 0.55 the features are clean; 0.78 is substantial leakage; above 0.85 the attribute is effectively still there.
The third check is the base rate, and it is the one that reframes everything: the training label approves group M at 50% and group F at 40%. Because that gap was injected at equal merit, the "correct" model — the one that minimises loss against this label — is a discriminating model. No amount of reweighting or threshold tuning reaches that, because those techniques change how you fit the target, not what the target is.
Before you fix a fairness metric, ask what the label literally records. If it records the decisions of a biased process, a better fit to it is a more faithful reproduction of the bias.
As a fourth, optional check, run SHAP on the shipped model and compare its ranking of feature influence with the probe's. A feature that both correlates with the protected attribute and carries high mean absolute SHAP value is a proxy that is actually driving decisions. A feature that correlates but has near-zero SHAP is correlated and idle — worth documenting, not worth removing.
Stage 4: write findings, not metrics
A table of numbers is not an audit. A finding is a claim with a magnitude, a mechanism, an alternative explanation that you considered and rejected, and an action. Write them in this shape:
FINDING 01 -- Unequal opportunity by group Severity: HIGHClaim Among applicants who would have repaid, the model declines 40% of group F and 16% of group M.Magnitude TPR gap 24.0 pts, 95% CI [16.9, 31.1]. At 18,000 applications a year this is roughly 690 declines of creditworthy group F applicants that would not have occurred at the group M error rate.Evidence notebook cell 14; metrics/per_group_2026-05-14.csvMechanism Historical label bias (base rates 50% / 40% at equal merit) plus proxy leakage via postcode_band (probe AUC 0.78).Rejected "Base rates differ, so the gap is expected." Rejected: per-group AUC also differs (0.89 vs 0.83), so ranking quality is unequal too, which differing base rates alone would not produce.Action Do not deploy. Change the prediction target to realised repayment rather than historical approval. Owner: model owner. Due: 30 days.Check the magnitude arithmetic, because this is the number a decision-maker will quote. Of 18,000 annual applications, 40% are group F, giving 7,200; of those, 40% would repay, giving 2,880 qualified applicants; the excess false negative rate is 40% − 16% = 24 points, so 2,880 × 0.24 ≈ 691 people. Extrapolation like this is the difference between a report that gets filed and a report that gets acted on — but state the assumption (that next year's applicant mix matches the test set) in the same breath, or someone will quote the number without it.
The recommendations section then needs owners and dates, or it is a wish list:
| When | Action | Why this and not something else |
|---|---|---|
| Immediately | Block deployment; keep the existing manual process | The gap is large, the interval excludes zero, and the harm is a denied loan |
| Within 30 days | Re-derive the label from realised repayment; drop or orthogonalise postcode_band; re-audit | Attacks the two mechanisms actually identified, rather than the metric |
| Before any redeploy | Collect enough group F data over 60 to make that cell auditable | The current cell cannot support any claim, in either direction |
| Ongoing | Monitor per-group TPR and selection rate monthly; alarm below a 0.80 ratio | Fairness measured once at launch describes a distribution that has already changed |
The executive summary is one paragraph, and the test of it is whether a person who cannot read a confusion matrix knows what to do after reading it. "The model declines creditworthy applicants in group F at 2.5 times the rate of group M, affecting an estimated 690 people a year. The cause is the historical approval label, which we recommend replacing. Do not deploy."
What a finished audit looks like
- Every metric in the report is per group. There is no aggregate accuracy figure presented as evidence of anything.
- Every group-level number carries an interval, and cells too small to support a claim say "insufficient data" rather than printing a point estimate.
- Your hand-written metrics and the library's agree to within rounding, and you have run both.
- All three injected mechanisms are detected and named: label bias, proxy leakage via
postcode_band, and the representation shortfall in the intersectional cell. - At least one intersectional slice is reported, along with the honest statement of what it cannot show.
- Each finding names a mechanism, not just a metric, and states one alternative explanation you considered and why you rejected it.
- Each recommendation has an owner and a date.
- The audit is reproducible: a fixed seed, a saved metrics file, and a notebook that runs top to bottom on a clean kernel.
Where audits go wrong
| Failure | What it looks like | Fix |
|---|---|---|
| Metric shopping | Seven metrics computed, the two that passed reported | Name the criterion in the scope section before computing anything, and publish all of them |
| Aggregate reporting | "80.4% accurate, AUC 0.88" | Per-group or it does not go in the report |
| Point estimates on tiny cells | "TPR 0.41 for women over 60" | Bootstrap every gap; say "cannot say" when the interval spans the question |
| Confusion matrix unpacked wrongly | TPR and FPR look transposed or impossible | ravel() returns tn, fp, fn, tp; cross-check against a library |
| Auditing the model the model selected | Test set built only from applicants the current system approved | Hold out a random control slice decided outside the model's influence |
| Stopping at "the model is biased" | A report with no named mechanism | Run the direct / proxy / label checks; the fix depends on which one fired |
| Levelling down | Parity achieved by approving fewer group M applicants | Report absolute per-group outcomes next to every ratio |
Submit the notebook, the generated report, the saved per-group metrics file, and a short README that says how to reproduce every number in it.
What this buys you
The habit this project installs is narrow and worth more than any single metric: never let a number be averaged over the people it is about. Aggregate accuracy hid a 24-point opportunity gap in a model that was, by two respectable measures, performing well. The disaggregation took four lines of code and turned a sign-off into a block.
The second habit is diagnostic discipline. Three teams facing the same 24-point gap will apply three different fixes — reweighting, threshold adjustment, more data — and two of them will report a green metric while the harm continues, because the mechanism here was the label and none of those touch the label. The direct-use, proxy, and base-rate checks take twenty minutes and tell you which repair can possibly work.
The third is that findings are written for people who will never open your notebook. The magnitude sentence — 690 people a year, stated with its assumption — is what makes an audit an input to a decision rather than an attachment to an email. Numbers without that sentence get acknowledged; numbers with it get acted on.