Course Content
AI Ethics and Governance
3 sections · 7 lessons
Mitigation and Auditing Strategies
A lending team runs a fairness check and finds their approval model selects Group B applicants at 16.7% and Group A applicants at 35.7%. The ratio is 0.47, well under 0.8 — the "four-fifths" rule of thumb for adverse impact that US employment regulators have used since the 1978 Uniform Guidelines, and that lending fairness teams commonly borrow. They reweight the training data, retrain, and the ratio comes back at 0.98. Everyone signs off.
Seven months later the ratio has drifted back to 0.61. Nothing was changed. The reweighting had rebalanced the labels, but the feature set still contained a postcode field that reconstructed group membership with an AUC of 0.89, and the retraining pipeline had been quietly fed six months of new outcomes generated by the model's own decisions — decisions that had already reshaped who got a loan and therefore who had a repayment record.
This is the standard shape of failed mitigation. A technique is applied at one point in the pipeline, the metric it targets moves, and the underlying mechanism is untouched. Mitigation only works if you have already identified which mechanism you are fighting.
Find the mechanism before choosing the tool
There are three places you can intervene: before training (the data), during training (the objective), and after training (the outputs). Each fixes a different class of problem and each fails at the others.
| Intervention point | You change | Works when | Fails when | Needs at inference |
|---|---|---|---|---|
| Pre-processing | Training data: weights, samples, feature values | The imbalance lives in the data — skewed labels, under-representation | The label itself measures the wrong thing; proxies remain | Nothing |
| In-processing | Loss function or training procedure | You control training and want an explicit accuracy/fairness dial | You are using a third-party or foundation model | Nothing |
| Post-processing | Thresholds or scores applied to predictions | The model is fixed, opaque, or vendor-supplied | Group-specific thresholds are illegal or unobservable at decision time | The protected attribute — often a blocker |
That last column decides more real projects than anything else. Post-processing usually requires knowing each individual's group membership when the decision is made. In hiring and lending, asking for it at that moment is often unlawful, so the elegant threshold-optimiser cannot be deployed.
Step zero: find the proxies
Every technique below is undermined by proxy leakage, so this comes first. Take your feature matrix with the protected attribute removed and try to predict the attribute from it.
1from sklearn.ensemble import RandomForestClassifier2from sklearn.inspection import permutation_importance3from sklearn.model_selection import train_test_split4from sklearn.metrics import roc_auc_score56Xtr, Xte, atr, ate = train_test_split(X_no_attr, a, test_size=0.3, stratify=a)7probe = RandomForestClassifier(n_estimators=300, min_samples_leaf=20).fit(Xtr, atr)89auc = roc_auc_score(ate, probe.predict_proba(Xte)[:, 1])10print(f"protected attribute recoverable: AUC = {auc:.3f}")1112# which features carry it?13imp = permutation_importance(probe, Xte, ate, scoring="roc_auc", n_repeats=10)14for i in imp.importances_mean.argsort()[::-1][:8]:15 print(f"{X_no_attr.columns[i]:<28} drop in AUC = {imp.importances_mean[i]:.3f}")Interpretation: AUC near 0.50 means the features are clean. 0.65 is meaningful leakage. Above 0.85 the attribute is effectively still in the dataset and any claim of "we don't use race" is cosmetic.
What you do next is a judgement, not a computation. Some proxies are pure redundancy and should go. Others carry genuine predictive signal that is only partly group-related — postcode may really predict repayment through commuting cost or local employment. The usable middle option is orthogonalisation: regress the suspect feature on the protected attribute, keep the residual, discard the fitted part.
1import numpy as np2from sklearn.linear_model import LinearRegression34# keep only the part of `income_proxy` that is NOT explained by group membership5resid = income_proxy - LinearRegression().fit(a.values.reshape(-1, 1),6 income_proxy).predict(a.values.reshape(-1, 1))7X["income_proxy_resid"] = residThis is honest but partial: it removes the linear component only, and it silently discards predictive power that may have been legitimate. Document the trade rather than hiding it.
Every mitigation technique assumes you know which door the bias came through. Applied to the wrong door, they move the metric and leave the harm in place.
Pre-processing: change the data
Reweighing, worked in full
Reweighing (Kamiran and Calders, 2012) assigns each training example a weight so that, in the weighted dataset, group and label are statistically independent. No row is deleted and no value is altered — only the weight the loss function gives it.
The weight for a group–label cell is the count you would expect under independence, divided by the count you observe:
Take 1,000 historical hiring records. Group A: 700 people, 250 hired. Group B: 300 people, 50 hired.
| Cell | Observed count | Expected if independent | Weight |
|---|---|---|---|
| A, hired | 250 | 1000 × 0.7 × 0.3 = 210 | 210 / 250 = 0.84 |
| A, not hired | 450 | 1000 × 0.7 × 0.7 = 490 | 490 / 450 = 1.089 |
| B, hired | 50 | 1000 × 0.3 × 0.3 = 90 | 90 / 50 = 1.80 |
| B, not hired | 250 | 1000 × 0.3 × 0.7 = 210 | 210 / 250 = 0.84 |
Check the result. Weighted positives in A: 250 × 0.84 = 210, out of weighted total 210 + 490 = 700, so 30%. Weighted positives in B: 50 × 1.80 = 90, out of 90 + 210 = 300, so 30%. Selection rate parity in the training distribution, exactly.
Before: 250/700 = 35.7% versus 50/300 = 16.7%, ratio 0.47. After: 30% versus 30%, ratio 1.00.
The 1.80 weight is the number to worry about. Fifty examples now carry the influence of ninety. Every quirk of those fifty people — a rare degree subject, one unusual employer — is amplified by 80%, which raises variance and can cause the model to overfit to a handful of individuals. Cross-validate per group, and if any weight exceeds about 3, treat it as a signal that you do not have enough data rather than a fix.
Resampling
Same goal, blunter instrument. Oversample the under-represented cell (duplicate or synthesise) or undersample the over-represented one.
| Method | What it does | Main risk |
|---|---|---|
| Random oversampling | Duplicates minority-cell rows | Exact duplicates encourage memorisation; no new information |
| SMOTE | Creates synthetic rows by interpolating between near neighbours | Interpolating between two real people can produce a combination that cannot exist; poor with categorical features |
| Random undersampling | Drops majority-cell rows | Throws away real data; reduces accuracy for everyone |
| Reweighing | Leaves data intact, changes loss weights | High weights inflate variance |
Prefer reweighing when your learner supports sample weights, which almost all of them do. Reserve SMOTE for cases where a subgroup is so small that the model has literally nothing to fit — and then be explicit that you are fabricating people.
What pre-processing cannot fix
If the label is the wrong measurement, rebalancing it just balances the wrong measurement. Obermeyer et al. (2019) found a widely deployed healthcare algorithm ranking patients by predicted cost as a stand-in for medical need; because less is spent on Black patients with equivalent illness, the score understated their need. Reweighing that dataset would have produced parity in predicted cost while leaving sicker patients unenrolled. The fix was to change the prediction target to actual health conditions, which raised the share of Black patients auto-enrolled from 17.7% to 46.5%. No amount of sample weighting reaches that.
In-processing: change the objective
Constrained optimisation
Instead of minimising loss alone, minimise loss subject to a fairness constraint. Fairlearn's ExponentiatedGradient implements the Agarwal et al. (2018) reduction: it repeatedly re-fits your ordinary classifier on re-weighted data, driving the constraint violation down, and returns a randomised mixture of those classifiers.
1from fairlearn.reductions import ExponentiatedGradient, DemographicParity, EqualizedOdds2from fairlearn.metrics import MetricFrame, selection_rate, true_positive_rate3from sklearn.ensemble import GradientBoostingClassifier45base = GradientBoostingClassifier(n_estimators=200, max_depth=3)67mitigated = ExponentiatedGradient(8 estimator=base,9 constraints=EqualizedOdds(difference_bound=0.05), # allow 5 points of slack10)11mitigated.fit(X_train, y_train, sensitive_features=a_train)1213frame = MetricFrame(14 metrics={"selection_rate": selection_rate, "tpr": true_positive_rate},15 y_true=y_test, y_pred=mitigated.predict(X_test),16 sensitive_features=a_test,17)18print(frame.by_group)19print("TPR gap:", frame.difference(method="between_groups")["tpr"])Two things to know before you use it. First, difference_bound is the dial: setting it to 0 demands exact parity and will cost you the most accuracy, while 0.05 usually costs very little. Second, the output is a randomised classifier — two identical applicants can receive different decisions. That is defensible in a research paper and very hard to defend to a regulator or a rejected applicant, so check whether you need predict to be deterministic before you ship it.
Adversarial debiasing
Train two networks at once (Zhang, Lemoine and Mitchell, 2018). The predictor maps features to the outcome. The adversary takes the predictor's output and tries to recover the protected attribute. The predictor is trained to minimise its own loss and maximise the adversary's, with the gradient component that would help the adversary projected out.
features --> [ predictor ] --> y_hat --> [ adversary ] --> a_hat ^ | | | minimise prediction loss maximise adversary loss AND remove the gradient direction that helps the adversaryConverged state: y_hat is as accurate as possible subject to carryingno information the adversary can use to recover `a`.The appeal is that it attacks proxies directly: it does not care which feature leaked group information, only that the output stops carrying it. The cost is that you are training a minimax game, with all the instability that implies — oscillation, adversary collapse, and sensitivity to the learning-rate ratio between the two networks. Budget for tuning, and monitor the adversary's accuracy during training: if it drops to chance immediately, it is probably too weak to be teaching the predictor anything.
Post-processing: change the decision rule
Leave the model alone; move the thresholds. Suppose your score is well ranked but distributed lower for Group B. At a single cutoff of 0.50:
| Rule | Group A TPR | Group B TPR | Group A FPR | Group B FPR | Group B PPV |
|---|---|---|---|---|---|
| Single threshold 0.50 | 0.80 | 0.55 | 0.12 | 0.09 | 0.71 |
| A at 0.50, B at 0.38 | 0.80 | 0.80 | 0.12 | 0.22 | 0.58 |
Equal opportunity is achieved: both groups' qualified members are now found at 80%. The price is explicit and worth stating out loud — Group B's false positive rate rises from 9% to 22%, and the meaning of a positive decision degrades from 71% correct to 58% correct for that group. Whether that is an improvement depends entirely on whether a false positive harms the individual or merely costs the institution.
1from fairlearn.postprocessing import ThresholdOptimizer23opt = ThresholdOptimizer(4 estimator=trained_model,5 constraints="true_positive_rate_parity", # or "equalized_odds", "demographic_parity"6 objective="accuracy_score",7 prefit=True,8)9opt.fit(X_train, y_train, sensitive_features=a_train)10y_pred = opt.predict(X_test, sensitive_features=a_test) # attribute REQUIRED at inferenceNote that last line. The protected attribute must be available when the decision is made. In US employment testing this approach also runs into a hard legal wall: section 106 of the Civil Rights Act of 1991 prohibits adjusting scores or using different cutoffs on the basis of race, colour, religion, sex or national origin. A technique that is optimal in a notebook can be unlawful in production.
The fairness–accuracy trade-off, honestly stated
Two claims are both true and people usually believe only one.
Sometimes the trade-off is real. If you constrain the classifier away from the loss-minimising solution, overall accuracy falls. On typical tabular benchmarks, going from unconstrained to a 5-point demographic-parity bound costs roughly 1–3 accuracy points; demanding exact parity costs considerably more.
Sometimes it is an illusion caused by a bad label or bad data. Fixing the Optum cost proxy improved both fairness and the algorithm's ability to identify sick patients, because it started predicting the right thing. Adding data for an under-represented group improves that group's accuracy without hurting anyone's.
| Situation | Is the trade-off real? | Right response |
|---|---|---|
| Subgroup has 200 training examples | No | Collect more data for that subgroup |
| Label is a proxy that is worse for one group | No | Change the target variable |
| Base rates genuinely differ and prediction is imperfect | Yes, and it is a theorem | Choose which metric to equalise, with justification |
| Constraint set to exact parity | Yes, and self-inflicted | Use a slack bound; exact parity is rarely required |
There is one failure mode worth naming: levelling down. You can always achieve perfect demographic parity by making the model worse for the advantaged group — reject more of them until the rates match. The metric goes green and nobody is better off. Always report absolute per-group outcomes alongside the ratio, so that a "fix" that helps nobody is visible.
If a fairness intervention improves the ratio without improving anyone's absolute outcome, you have not reduced harm — you have redistributed it downwards.
The audit: five steps that produce evidence
An audit is not a metric printout. It is a document that someone outside the team can read and use to decide whether to deploy.
1. Scope
Write down: the decision the system makes, who is affected, which groups will be examined and why, which fairness criterion is being prioritised and on what grounds, what the harm of a false positive is, and what the harm of a false negative is. Both harm sentences must name a person, not a metric. This step is where the fairness criterion is chosen — before any number is computed, so the choice cannot be reverse-engineered from whichever metric happened to look best.
2. Data audit
- Group sizes in train, validation and test, in absolute counts. A group with 60 test rows cannot support a claim about its error rate.
- Base rate of the positive label per group — this determines which fairness metrics are even simultaneously achievable.
- Missingness per group. Differential missingness is a common hidden proxy: the field is blank for people the old process never processed.
- Proxy leakage AUC, plus the top leaking features.
- Provenance: who collected it, when, under what process, and what changed since. This is the datasheet (Gebru et al., 2018).
3. Model and fairness audit
Per group, report: n, base rate, selection rate, accuracy, TPR, FPR, PPV, and a confidence interval on each. Then intersections — the Gender Shades result was invisible at the level of "gender" or "skin tone" alone and only appeared at darker-skinned-and-female. Then slice by everything else that matters: age band, region, device, language, and time period.
1from fairlearn.metrics import MetricFrame, false_positive_rate, true_positive_rate2from sklearn.metrics import accuracy_score, precision_score3import pandas as pd45intersection = a_test.astype(str) + " | " + age_band_test.astype(str)67mf = MetricFrame(8 metrics={"accuracy": accuracy_score, "tpr": true_positive_rate,9 "fpr": false_positive_rate, "ppv": precision_score},10 y_true=y_test, y_pred=y_pred, sensitive_features=intersection,11)12report = mf.by_group.copy()13report["n"] = pd.Series(intersection).value_counts()14print(report.sort_values("n")) # smallest cells first — that is where it breaks4. Documentation
A model card (Mitchell et al., 2019) states intended use, out-of-scope uses, training data, per-group evaluation results, ethical considerations, and known limitations. The out-of-scope section is the one that does real work: COMPAS was built to assess supervision needs and used for detention decisions, and an explicit "this model must not be used to inform custodial decisions" line is the artefact that lets someone object later.
5. Continuous monitoring
Fairness is not a property you establish once. Populations shift, upstream data pipelines change, and models retrained on their own outputs drift.
| Monitor | Cadence | Alert when |
|---|---|---|
| Per-group selection rate | Daily | Ratio between any two groups falls below 0.80 |
| Per-group TPR / FPR | Weekly, as labels arrive | Gap widens by more than 5 points versus the audited baseline |
| Group composition of inputs | Weekly | Any group's share shifts by more than 20% relative |
| Feature drift on top proxies | Weekly | Distribution shift beyond the audited range |
| Appeal and override rates by group | Monthly | One group appeals or is overridden far more often — a sign the model is wrong about them |
The last row is the most under-used signal in production. Humans overriding the model disproportionately for one group is direct field evidence of a disparity that your offline test set did not capture.
Where mitigation goes wrong
| Failure mode | What it looks like | Fix |
|---|---|---|
| Metric shopping | Compute all fairness metrics, report the one that passes | Commit to the criterion in step 1, in writing, before measuring |
| Fixing the symptom | Reweighing applied to a mislabelled target | Diagnose the entry point first; label problems need label fixes |
| Levelling down | Parity achieved by degrading the advantaged group | Report absolute per-group outcomes, not only ratios |
| Train-time only | Audited once at launch, never again | Automated per-group monitoring with alert thresholds |
| Aggregate-only reporting | "94% accurate" with no group breakdown | Ban aggregate-only metrics from review documents |
| Underpowered subgroups | A 40-row group reported to two decimal places | Publish confidence intervals; state when n is too small to conclude |
| Feedback contamination | Retraining on outcomes the model itself produced | Hold out a random control slice from model influence |
What this looks like in a real project
Order matters more than technique. Run the proxy probe first, because its result tells you whether pre-processing has any chance of working. Fix the label second, because a wrong target invalidates everything downstream and is often the cheapest fix available. Only then reach for reweighing or a constraint, and pick the intervention point by asking a practical question rather than an academic one: do you control training (in-processing available?), and will you know the person's group at decision time (post-processing available?). Those two answers eliminate most of the menu.
Set a slack bound rather than exact parity, because exact parity buys almost nothing over a 5-point bound and costs a great deal of accuracy. Report absolute outcomes next to every ratio so levelling down cannot hide. Write the model card before launch rather than after, because the out-of-scope-uses section is the only thing that will still be there when someone repurposes your model in two years.
And keep the control slice. A small randomly selected fraction of decisions made outside the model's influence is the only clean data you will ever have for measuring what the model is actually doing to the world, as opposed to what it is doing to its own training set.