AI Ethics and Governance

Mini-Project: Bias Audit on Fine-Tuned Model


You are the data scientist on a governance team. A loan approval model is queued for deployment and your job is to sign it off, or not. You run the obvious checks and get this:

Text
Test set: 1,500 applicantsOverall accuracy            0.804ROC-AUC                     0.88Precision, group M          0.808Precision, group F          0.800     <- a 0.8-point difference

Precision is the fraction of approvals that were correct. It is nearly identical across groups: when this model says yes, it is right about 80% of the time for everyone. Accuracy is respectable. Ranking quality is good. On this evidence you would sign.

Now disaggregate — split every number by group instead of averaging over them:

Text
                      group M      group Fapplicants                900          600would have repaid         450          240approved by model         468          180true positive rate       0.840        0.600false negative rate      0.160        0.400

Among applicants who would have repaid, the model declines 16% of group M and 40% of group F. A qualified applicant in group F is 2.5 times as likely to be declined as an equally qualified applicant in group M. That is 96 people out of 240 in a single test set, and none of it was visible in the aggregate.

This is what a bias audit is for, and why the first rule is that every number gets computed per group and never averaged across them. What follows is a build brief: four stages, real metrics, and a description of what a finding actually looks like when you write one down.

The loan model you are asked to sign off35.7 percent0.820.14—16.7 percent0.610.310.47Approval rateTrue positive rateFalse positive rateSelection ratioGroup AGroup BOverall accuracy is 0.86 for both groups.
Equal accuracy across groups sits comfortably beside a 0.47 selection ratio — which gap you report decides what the audit is even about.

What you are building

Four artefacts, roughly three hours of work:

  1. A test set with known ground truth — synthetic data into which you have deliberately injected specific, documented biases, so that you can check whether your audit detects the things you know are there.
  2. A disaggregated metrics table, computed by hand from confusion matrices and then cross-checked against a maintained library.
  3. A mechanism diagnosis — not "the model is biased" but which of the four possible causes produced each gap.
  4. An audit report whose findings a non-technical reader can act on.

Injecting bias on purpose is the part people skip, and it is the part that makes the exercise honest. If you audit a model with unknown bias and find nothing, you cannot tell whether the model is clean or your audit is blind. With known ground truth, a null result is a bug report about your own code.

Stage 1: a test set whose biases you already know

Build 5,000 applicants with three separate mechanisms planted in them. Public alternatives exist — Adult Income, German Credit, COMPAS — and you should run against one afterwards, but start synthetic so you know the answer.

Python
import numpy as np, pandas as pdRNG = np.random.default_rng(42)def build_dataset(n=5000):    group  = RNG.choice(["M", "F"], n, p=[0.60, 0.40])          # imbalance: about 3000 / 2000    credit = np.clip(RNG.normal(680, 80, n), 300, 850)    income = RNG.lognormal(10.5, 0.7, n)    debt   = RNG.beta(2, 5, n)    # MECHANISM 2 -- proxy. postcode_band carries group membership, and it    # survives after `group` itself is dropped from the feature set.    postcode = RNG.normal(np.where(group == "M", 0.0, 1.1), 1.0, n)    # Genuine creditworthiness: what a fair model should be predicting.    merit = (credit - 680) / 80 + 0.6 * (np.log(income) - 10.5) - 1.4 * debt    # MECHANISM 1 -- historical label bias. At equal merit the old process    # approved group M more often. +0.55 on the logit is roughly    # +10 percentage points near the middle of the distribution.    logit = 1.1 * merit + np.where(group == "M", 0.55, 0.0) - 0.10    approved = (RNG.random(n) < 1 / (1 + np.exp(-logit))).astype(int)    return pd.DataFrame({"group": group, "credit_score": credit, "income": income,                         "debt_ratio": debt, "postcode_band": postcode,                         "merit": merit, "approved": approved})df = build_dataset()print(df.groupby("group")["approved"].agg(["size", "mean"]))#         size      mean# group# F       1951  0.390569# M       3049  0.504100

Three mechanisms are now in the data, and you should be able to say out loud what each one will do to which metric:

  • Label bias — the historical outcome itself favours group M. A model that fits this label perfectly will reproduce a 10-point base-rate gap and call it accuracy.
  • Proxy leakage — postcode_band predicts group membership. Dropping the group column removes it from the schema, not from the model.
  • Representation imbalance — group F is 40% of applicants, and any intersectional cell you slice will be smaller still.

Then stratify. Split so that both groups keep their proportions in train and test, and print the cell sizes before you go any further, because a metric on a cell of 40 is noise wearing a decimal point.

Python
from sklearn.model_selection import train_test_splittrain, test = train_test_split(df, test_size=0.30, stratify=df[["group", "approved"]],                               random_state=42)cells = test.groupby(["group", "approved"]).size().unstack()print(cells)                       # F: 357 declined / 229 approved,  M: 453 / 461print("smallest cell:", int(cells.values.min()))

Finally, carve out the scenarios most likely to expose a threshold problem, and check the group mix inside each one: high income with a low credit score, low income with a high credit score, thin files (under two years of employment history), and records with a missing field. The last is not decoration — differential missingness is a proxy in disguise, because a field tends to be blank exactly for the people the old process never handled.

Stage 2: compute the numbers yourself, then check them

Write the metric code by hand once. Libraries are correct and you should use them in production, but a metric you have never derived is a metric you cannot debug when it disagrees with someone else's.

Python
from sklearn.metrics import confusion_matrix, roc_auc_scoredef per_group(y_true, y_pred, y_score, groups):    rows = []    for g in sorted(set(groups)):        m = groups == g        # labels=[0, 1] is load-bearing: without it a group containing only        # one class returns a 1x1 matrix and .ravel() raises on unpacking.        tn, fp, fn, tp = confusion_matrix(y_true[m], y_pred[m], labels=[0, 1]).ravel()        n = int(m.sum())        rows.append({"group": g, "n": n,                     "base_rate":      (tp + fn) / n,                     "selection_rate": (tp + fp) / n,                     "tpr":            tp / (tp + fn) if tp + fn else np.nan,                     "fpr":            fp / (fp + tn) if fp + tn else np.nan,                     "ppv":            tp / (tp + fp) if tp + fp else np.nan,                     "accuracy":       (tp + tn) / n,                     "auc":            roc_auc_score(y_true[m], y_score[m])})    return pd.DataFrame(rows).set_index("group")

Train a logistic regression on the standardised features with group dropped, predict at the default 0.5 threshold, and you get a table like this one — the table the rest of the audit turns on. Its figures are a worked illustration, rounded so the arithmetic is easy to follow.

MetricGroup M (n=900)Group F (n=600)Gap or ratioReading
Base rate (label)0.5000.40010.0 ptsThe injected label bias, intact
Selection rate0.5200.300ratio 0.58Fails the 0.80 four-fifths rule
True positive rate0.8400.60024.0 ptsFails equal opportunity badly
False positive rate0.2000.10010.0 ptsModel is more cautious with group F
Precision (PPV)0.8080.8000.8 ptsPasses — and this is the trap
Accuracy0.8200.7804.0 ptsPasses a 5-point bound
ROC-AUC0.8900.8306.0 ptsThe model ranks group F less well

Two rows deserve a hard look. Precision passing is not reassurance — it says the model is equally right when it approves someone, which is a statement about the people it said yes to and tells you nothing about the people it said no to. And the false positive rate being lower for group F is not a kindness: paired with the low TPR it is the signature of an effectively higher decision threshold for that group.

Precision measures the quality of the approvals. True positive rate measures who was missed. An audit that reports only the first is an audit of the people the model already liked.

Now put an interval on the headline gap, because a 24-point difference on 240 positives is not the same claim as a 24-point difference on 24.

Python
def bootstrap_tpr_gap(y_true, y_pred, groups, a="M", b="F", n_boot=2000, seed=0):    rng, n, gaps = np.random.default_rng(seed), len(y_true), []    for _ in range(n_boot):        s = rng.integers(0, n, n)                        # resample with replacement        rate = {}        for g in (a, b):            pos = (groups[s] == g) & (y_true[s] == 1)            rate[g] = y_pred[s][pos].mean() if pos.sum() else np.nan        gaps.append(rate[a] - rate[b])    lo, hi = np.nanpercentile(gaps, [2.5, 97.5])    return float(np.nanmean(gaps)), float(lo), float(hi)# TPR gap, M minus F:  0.240   95% CI [0.169, 0.311]

The interval excludes zero by a wide margin, so this is a finding rather than noise. Run the same function on an intersectional cell and watch what happens: group F aged 60 and over has 58 test applicants, 22 of whom would have repaid, and 9 of those were approved. The point estimate is a TPR of 0.41; the 95% interval is [0.23, 0.61], which comfortably contains both group TPRs from the table above. The honest report of that cell is "we cannot say", and that sentence is itself a finding: you do not have enough data to audit older female applicants at all.

Finally, cross-check against a maintained implementation. If the two disagree by more than rounding, you have a bug, and the usual culprit is the unpacking order of confusion_matrix(...).ravel(), which returns tn, fp, fn, tp — not the order most people write from memory.

Python
from fairlearn.metrics import (MetricFrame, selection_rate,                               true_positive_rate, false_positive_rate)mf = MetricFrame(metrics={"selection_rate": selection_rate,                          "tpr": true_positive_rate, "fpr": false_positive_rate},                 y_true=y_test, y_pred=y_pred, sensitive_features=g_test)print(mf.by_group)print("selection-rate ratio:", mf.ratio()["selection_rate"])   # 0.577

Reading the table: which gap means what

"The model is biased" is not a finding, because it does not imply an action. The four patterns below do, and each points at a different repair.

Pattern in the numbersWhat it meansWhat fixes itWhat will not
Per-group AUC equal, TPR unequal at one thresholdRanking is equally good; the single global cutoff falls in a different place on each group's score distributionThreshold placement, or recalibration per group where that is lawfulMore data — the model already ranks correctly
Per-group AUC unequal (here 0.89 vs 0.83)The model genuinely knows less about one group; fewer rows, weaker features, or more label noise thereMore and better data for that group; features that carry signal for itThreshold tuning, which just moves errors between the two kinds
TPR and FPR equal, selection rates unequalBase rates genuinely differ; the model is treating like cases alikeA decision about whether parity of outcome is the right target here, argued explicitlyForcing equal selection rates, which now costs accuracy for real reasons
Precision equal, TPR unequal, base rates unequalThe classic label-bias signature: the model faithfully reproduced a biased historical targetChanging what the label measuresAny reweighting or threshold work, which rebalances the wrong measurement

Our model shows the fourth pattern and the second at once, which is common and worth stating in the report as two separate findings with two separate owners.

A gap is a symptom. Until you know whether the model ranks the group worse or merely cuts it at a different point, you cannot tell which of the available repairs would do anything at all.

Stage 3: find the mechanism

Three checks, in this order, because each one changes how you interpret the next.

Python
# 1. Direct use. Fit once WITH the group column to see what it is worth.coef = dict(zip(feature_names, model_with_group.coef_[0]))print(f"group coefficient {coef['group_M']:+.3f} "      f"-> odds ratio {np.exp(coef['group_M']):.2f}")# group coefficient +0.693 -> odds ratio 2.00# 2. Proxy leakage. Can the group be recovered from the shipped features?from sklearn.ensemble import RandomForestClassifierfrom sklearn.inspection import permutation_importanceprobe = RandomForestClassifier(n_estimators=300, min_samples_leaf=20).fit(Xtr, gtr)print("probe AUC:", roc_auc_score(gte, probe.predict_proba(Xte)[:, 1]))   # 0.78imp = permutation_importance(probe, Xte, gte, scoring="roc_auc", n_repeats=10)# postcode_band   drop in AUC = 0.190     <- the leak# income          drop in AUC = 0.021

Read these together. An odds ratio of 2.00 means that with the group column present, being in group M doubles the odds of approval at identical credit score, income and debt ratio — direct use, severity high. Dropping that column removes it, but a probe AUC of 0.78 says the shipped feature set still recovers group membership, driven almost entirely by postcode_band. Below about 0.55 the features are clean; 0.78 is substantial leakage; above 0.85 the attribute is effectively still there.

The third check is the base rate, and it is the one that reframes everything: the training label approves group M at 50% and group F at 40%. Because that gap was injected at equal merit, the "correct" model — the one that minimises loss against this label — is a discriminating model. No amount of reweighting or threshold tuning reaches that, because those techniques change how you fit the target, not what the target is.

Before you fix a fairness metric, ask what the label literally records. If it records the decisions of a biased process, a better fit to it is a more faithful reproduction of the bias.

As a fourth, optional check, run SHAP on the shipped model and compare its ranking of feature influence with the probe's. A feature that both correlates with the protected attribute and carries high mean absolute SHAP value is a proxy that is actually driving decisions. A feature that correlates but has near-zero SHAP is correlated and idle — worth documenting, not worth removing.

Stage 4: write findings, not metrics

A table of numbers is not an audit. A finding is a claim with a magnitude, a mechanism, an alternative explanation that you considered and rejected, and an action. Write them in this shape:

Text
FINDING 01 -- Unequal opportunity by group                        Severity: HIGHClaim     Among applicants who would have repaid, the model declines 40% of          group F and 16% of group M.Magnitude TPR gap 24.0 pts, 95% CI [16.9, 31.1]. At 18,000 applications a year          this is roughly 690 declines of creditworthy group F applicants that          would not have occurred at the group M error rate.Evidence  notebook cell 14; metrics/per_group_2026-05-14.csvMechanism Historical label bias (base rates 50% / 40% at equal merit) plus          proxy leakage via postcode_band (probe AUC 0.78).Rejected  "Base rates differ, so the gap is expected." Rejected: per-group AUC          also differs (0.89 vs 0.83), so ranking quality is unequal too, which          differing base rates alone would not produce.Action    Do not deploy. Change the prediction target to realised repayment          rather than historical approval. Owner: model owner. Due: 30 days.

Check the magnitude arithmetic, because this is the number a decision-maker will quote. Of 18,000 annual applications, 40% are group F, giving 7,200; of those, 40% would repay, giving 2,880 qualified applicants; the excess false negative rate is 40% − 16% = 24 points, so 2,880 × 0.24 ≈ 691 people. Extrapolation like this is the difference between a report that gets filed and a report that gets acted on — but state the assumption (that next year's applicant mix matches the test set) in the same breath, or someone will quote the number without it.

The recommendations section then needs owners and dates, or it is a wish list:

WhenActionWhy this and not something else
ImmediatelyBlock deployment; keep the existing manual processThe gap is large, the interval excludes zero, and the harm is a denied loan
Within 30 daysRe-derive the label from realised repayment; drop or orthogonalise postcode_band; re-auditAttacks the two mechanisms actually identified, rather than the metric
Before any redeployCollect enough group F data over 60 to make that cell auditableThe current cell cannot support any claim, in either direction
OngoingMonitor per-group TPR and selection rate monthly; alarm below a 0.80 ratioFairness measured once at launch describes a distribution that has already changed

The executive summary is one paragraph, and the test of it is whether a person who cannot read a confusion matrix knows what to do after reading it. "The model declines creditworthy applicants in group F at 2.5 times the rate of group M, affecting an estimated 690 people a year. The cause is the historical approval label, which we recommend replacing. Do not deploy."

What a finished audit looks like

  1. Every metric in the report is per group. There is no aggregate accuracy figure presented as evidence of anything.
  2. Every group-level number carries an interval, and cells too small to support a claim say "insufficient data" rather than printing a point estimate.
  3. Your hand-written metrics and the library's agree to within rounding, and you have run both.
  4. All three injected mechanisms are detected and named: label bias, proxy leakage via postcode_band, and the representation shortfall in the intersectional cell.
  5. At least one intersectional slice is reported, along with the honest statement of what it cannot show.
  6. Each finding names a mechanism, not just a metric, and states one alternative explanation you considered and why you rejected it.
  7. Each recommendation has an owner and a date.
  8. The audit is reproducible: a fixed seed, a saved metrics file, and a notebook that runs top to bottom on a clean kernel.

Where audits go wrong

FailureWhat it looks likeFix
Metric shoppingSeven metrics computed, the two that passed reportedName the criterion in the scope section before computing anything, and publish all of them
Aggregate reporting"80.4% accurate, AUC 0.88"Per-group or it does not go in the report
Point estimates on tiny cells"TPR 0.41 for women over 60"Bootstrap every gap; say "cannot say" when the interval spans the question
Confusion matrix unpacked wronglyTPR and FPR look transposed or impossibleravel() returns tn, fp, fn, tp; cross-check against a library
Auditing the model the model selectedTest set built only from applicants the current system approvedHold out a random control slice decided outside the model's influence
Stopping at "the model is biased"A report with no named mechanismRun the direct / proxy / label checks; the fix depends on which one fired
Levelling downParity achieved by approving fewer group M applicantsReport absolute per-group outcomes next to every ratio

Submit the notebook, the generated report, the saved per-group metrics file, and a short README that says how to reproduce every number in it.

What this buys you

The habit this project installs is narrow and worth more than any single metric: never let a number be averaged over the people it is about. Aggregate accuracy hid a 24-point opportunity gap in a model that was, by two respectable measures, performing well. The disaggregation took four lines of code and turned a sign-off into a block.

The second habit is diagnostic discipline. Three teams facing the same 24-point gap will apply three different fixes — reweighting, threshold adjustment, more data — and two of them will report a green metric while the harm continues, because the mechanism here was the label and none of those touch the label. The direct-use, proxy, and base-rate checks take twenty minutes and tell you which repair can possibly work.

The third is that findings are written for people who will never open your notebook. The magnitude sentence — 690 people a year, stated with its assumption — is what makes an audit an input to a decision rather than an attachment to an email. Numbers without that sentence get acknowledged; numbers with it get acted on.