AI Ethics and Governance

Mitigation and Auditing Strategies


A lending team runs a fairness check and finds their approval model selects Group B applicants at 16.7% and Group A applicants at 35.7%. The ratio is 0.47, well under 0.8 — the "four-fifths" rule of thumb for adverse impact that US employment regulators have used since the 1978 Uniform Guidelines, and that lending fairness teams commonly borrow. They reweight the training data, retrain, and the ratio comes back at 0.98. Everyone signs off.

Seven months later the ratio has drifted back to 0.61. Nothing was changed. The reweighting had rebalanced the labels, but the feature set still contained a postcode field that reconstructed group membership with an AUC of 0.89, and the retraining pipeline had been quietly fed six months of new outcomes generated by the model's own decisions — decisions that had already reshaped who got a loan and therefore who had a repayment record.

This is the standard shape of failed mitigation. A technique is applied at one point in the pipeline, the metric it targets moves, and the underlying mechanism is untouched. Mitigation only works if you have already identified which mechanism you are fighting.

Three places you can intervene on a 0.47 ratioMeasure: 16.7 against 35.7 percent, ratio 0.47Pre-processing — reweigh or resample the dataIn-processing — constrain the objectivePost-processing — per-group decision thresholdsAudit — evidence a third party can re-run
Each layer fixes a different mechanism, so choosing the tool before you have found the mechanism is how a fairness fix moves the number without moving the harm.

Find the mechanism before choosing the tool

There are three places you can intervene: before training (the data), during training (the objective), and after training (the outputs). Each fixes a different class of problem and each fails at the others.

Intervention pointYou changeWorks whenFails whenNeeds at inference
Pre-processingTraining data: weights, samples, feature valuesThe imbalance lives in the data — skewed labels, under-representationThe label itself measures the wrong thing; proxies remainNothing
In-processingLoss function or training procedureYou control training and want an explicit accuracy/fairness dialYou are using a third-party or foundation modelNothing
Post-processingThresholds or scores applied to predictionsThe model is fixed, opaque, or vendor-suppliedGroup-specific thresholds are illegal or unobservable at decision timeThe protected attribute — often a blocker

That last column decides more real projects than anything else. Post-processing usually requires knowing each individual's group membership when the decision is made. In hiring and lending, asking for it at that moment is often unlawful, so the elegant threshold-optimiser cannot be deployed.

Step zero: find the proxies

Every technique below is undermined by proxy leakage, so this comes first. Take your feature matrix with the protected attribute removed and try to predict the attribute from it.

Python
from sklearn.ensemble import RandomForestClassifierfrom sklearn.inspection import permutation_importancefrom sklearn.model_selection import train_test_splitfrom sklearn.metrics import roc_auc_scoreXtr, Xte, atr, ate = train_test_split(X_no_attr, a, test_size=0.3, stratify=a)probe = RandomForestClassifier(n_estimators=300, min_samples_leaf=20).fit(Xtr, atr)auc = roc_auc_score(ate, probe.predict_proba(Xte)[:, 1])print(f"protected attribute recoverable: AUC = {auc:.3f}")# which features carry it?imp = permutation_importance(probe, Xte, ate, scoring="roc_auc", n_repeats=10)for i in imp.importances_mean.argsort()[::-1][:8]:    print(f"{X_no_attr.columns[i]:<28} drop in AUC = {imp.importances_mean[i]:.3f}")

Interpretation: AUC near 0.50 means the features are clean. 0.65 is meaningful leakage. Above 0.85 the attribute is effectively still in the dataset and any claim of "we don't use race" is cosmetic.

What you do next is a judgement, not a computation. Some proxies are pure redundancy and should go. Others carry genuine predictive signal that is only partly group-related — postcode may really predict repayment through commuting cost or local employment. The usable middle option is orthogonalisation: regress the suspect feature on the protected attribute, keep the residual, discard the fitted part.

Python
import numpy as npfrom sklearn.linear_model import LinearRegression# keep only the part of `income_proxy` that is NOT explained by group membershipresid = income_proxy - LinearRegression().fit(a.values.reshape(-1, 1),                                              income_proxy).predict(a.values.reshape(-1, 1))X["income_proxy_resid"] = resid

This is honest but partial: it removes the linear component only, and it silently discards predictive power that may have been legitimate. Document the trade rather than hiding it.

Every mitigation technique assumes you know which door the bias came through. Applied to the wrong door, they move the metric and leave the harm in place.

Pre-processing: change the data

Reweighing, worked in full

Reweighing (Kamiran and Calders, 2012) assigns each training example a weight so that, in the weighted dataset, group and label are statistically independent. No row is deleted and no value is altered — only the weight the loss function gives it.

The weight for a group–label cell is the count you would expect under independence, divided by the count you observe:

w(g,y)  =  N⋅P(g)⋅P(y)N(g,y)w(g, y) \;=\; \frac{N \cdot P(g) \cdot P(y)}{N(g, y)}

Take 1,000 historical hiring records. Group A: 700 people, 250 hired. Group B: 300 people, 50 hired.

CellObserved countExpected if independentWeight
A, hired2501000 × 0.7 × 0.3 = 210210 / 250 = 0.84
A, not hired4501000 × 0.7 × 0.7 = 490490 / 450 = 1.089
B, hired501000 × 0.3 × 0.3 = 9090 / 50 = 1.80
B, not hired2501000 × 0.3 × 0.7 = 210210 / 250 = 0.84

Check the result. Weighted positives in A: 250 × 0.84 = 210, out of weighted total 210 + 490 = 700, so 30%. Weighted positives in B: 50 × 1.80 = 90, out of 90 + 210 = 300, so 30%. Selection rate parity in the training distribution, exactly.

Before: 250/700 = 35.7% versus 50/300 = 16.7%, ratio 0.47. After: 30% versus 30%, ratio 1.00.

The 1.80 weight is the number to worry about. Fifty examples now carry the influence of ninety. Every quirk of those fifty people — a rare degree subject, one unusual employer — is amplified by 80%, which raises variance and can cause the model to overfit to a handful of individuals. Cross-validate per group, and if any weight exceeds about 3, treat it as a signal that you do not have enough data rather than a fix.

Resampling

Same goal, blunter instrument. Oversample the under-represented cell (duplicate or synthesise) or undersample the over-represented one.

MethodWhat it doesMain risk
Random oversamplingDuplicates minority-cell rowsExact duplicates encourage memorisation; no new information
SMOTECreates synthetic rows by interpolating between near neighboursInterpolating between two real people can produce a combination that cannot exist; poor with categorical features
Random undersamplingDrops majority-cell rowsThrows away real data; reduces accuracy for everyone
ReweighingLeaves data intact, changes loss weightsHigh weights inflate variance

Prefer reweighing when your learner supports sample weights, which almost all of them do. Reserve SMOTE for cases where a subgroup is so small that the model has literally nothing to fit — and then be explicit that you are fabricating people.

What pre-processing cannot fix

If the label is the wrong measurement, rebalancing it just balances the wrong measurement. Obermeyer et al. (2019) found a widely deployed healthcare algorithm ranking patients by predicted cost as a stand-in for medical need; because less is spent on Black patients with equivalent illness, the score understated their need. Reweighing that dataset would have produced parity in predicted cost while leaving sicker patients unenrolled. The fix was to change the prediction target to actual health conditions, which raised the share of Black patients auto-enrolled from 17.7% to 46.5%. No amount of sample weighting reaches that.

In-processing: change the objective

Constrained optimisation

Instead of minimising loss alone, minimise loss subject to a fairness constraint. Fairlearn's ExponentiatedGradient implements the Agarwal et al. (2018) reduction: it repeatedly re-fits your ordinary classifier on re-weighted data, driving the constraint violation down, and returns a randomised mixture of those classifiers.

Python
from fairlearn.reductions import ExponentiatedGradient, DemographicParity, EqualizedOddsfrom fairlearn.metrics import MetricFrame, selection_rate, true_positive_ratefrom sklearn.ensemble import GradientBoostingClassifierbase = GradientBoostingClassifier(n_estimators=200, max_depth=3)mitigated = ExponentiatedGradient(    estimator=base,    constraints=EqualizedOdds(difference_bound=0.05),   # allow 5 points of slack)mitigated.fit(X_train, y_train, sensitive_features=a_train)frame = MetricFrame(    metrics={"selection_rate": selection_rate, "tpr": true_positive_rate},    y_true=y_test, y_pred=mitigated.predict(X_test),    sensitive_features=a_test,)print(frame.by_group)print("TPR gap:", frame.difference(method="between_groups")["tpr"])

Two things to know before you use it. First, difference_bound is the dial: setting it to 0 demands exact parity and will cost you the most accuracy, while 0.05 usually costs very little. Second, the output is a randomised classifier — two identical applicants can receive different decisions. That is defensible in a research paper and very hard to defend to a regulator or a rejected applicant, so check whether you need predict to be deterministic before you ship it.

Adversarial debiasing

Train two networks at once (Zhang, Lemoine and Mitchell, 2018). The predictor maps features to the outcome. The adversary takes the predictor's output and tries to recover the protected attribute. The predictor is trained to minimise its own loss and maximise the adversary's, with the gradient component that would help the adversary projected out.

Text
features --> [ predictor ] --> y_hat --> [ adversary ] --> a_hat                  ^                                    |                  |                                    |     minimise prediction loss              maximise adversary loss     AND remove the gradient direction that helps the adversaryConverged state: y_hat is as accurate as possible subject to carryingno information the adversary can use to recover `a`.

The appeal is that it attacks proxies directly: it does not care which feature leaked group information, only that the output stops carrying it. The cost is that you are training a minimax game, with all the instability that implies — oscillation, adversary collapse, and sensitivity to the learning-rate ratio between the two networks. Budget for tuning, and monitor the adversary's accuracy during training: if it drops to chance immediately, it is probably too weak to be teaching the predictor anything.

Post-processing: change the decision rule

Leave the model alone; move the thresholds. Suppose your score is well ranked but distributed lower for Group B. At a single cutoff of 0.50:

RuleGroup A TPRGroup B TPRGroup A FPRGroup B FPRGroup B PPV
Single threshold 0.500.800.550.120.090.71
A at 0.50, B at 0.380.800.800.120.220.58

Equal opportunity is achieved: both groups' qualified members are now found at 80%. The price is explicit and worth stating out loud — Group B's false positive rate rises from 9% to 22%, and the meaning of a positive decision degrades from 71% correct to 58% correct for that group. Whether that is an improvement depends entirely on whether a false positive harms the individual or merely costs the institution.

Python
from fairlearn.postprocessing import ThresholdOptimizeropt = ThresholdOptimizer(    estimator=trained_model,    constraints="true_positive_rate_parity",   # or "equalized_odds", "demographic_parity"    objective="accuracy_score",    prefit=True,)opt.fit(X_train, y_train, sensitive_features=a_train)y_pred = opt.predict(X_test, sensitive_features=a_test)   # attribute REQUIRED at inference

Note that last line. The protected attribute must be available when the decision is made. In US employment testing this approach also runs into a hard legal wall: section 106 of the Civil Rights Act of 1991 prohibits adjusting scores or using different cutoffs on the basis of race, colour, religion, sex or national origin. A technique that is optimal in a notebook can be unlawful in production.

The fairness–accuracy trade-off, honestly stated

Two claims are both true and people usually believe only one.

Sometimes the trade-off is real. If you constrain the classifier away from the loss-minimising solution, overall accuracy falls. On typical tabular benchmarks, going from unconstrained to a 5-point demographic-parity bound costs roughly 1–3 accuracy points; demanding exact parity costs considerably more.

Sometimes it is an illusion caused by a bad label or bad data. Fixing the Optum cost proxy improved both fairness and the algorithm's ability to identify sick patients, because it started predicting the right thing. Adding data for an under-represented group improves that group's accuracy without hurting anyone's.

SituationIs the trade-off real?Right response
Subgroup has 200 training examplesNoCollect more data for that subgroup
Label is a proxy that is worse for one groupNoChange the target variable
Base rates genuinely differ and prediction is imperfectYes, and it is a theoremChoose which metric to equalise, with justification
Constraint set to exact parityYes, and self-inflictedUse a slack bound; exact parity is rarely required

There is one failure mode worth naming: levelling down. You can always achieve perfect demographic parity by making the model worse for the advantaged group — reject more of them until the rates match. The metric goes green and nobody is better off. Always report absolute per-group outcomes alongside the ratio, so that a "fix" that helps nobody is visible.

If a fairness intervention improves the ratio without improving anyone's absolute outcome, you have not reduced harm — you have redistributed it downwards.

The audit: five steps that produce evidence

An audit is not a metric printout. It is a document that someone outside the team can read and use to decide whether to deploy.

1. Scope

Write down: the decision the system makes, who is affected, which groups will be examined and why, which fairness criterion is being prioritised and on what grounds, what the harm of a false positive is, and what the harm of a false negative is. Both harm sentences must name a person, not a metric. This step is where the fairness criterion is chosen — before any number is computed, so the choice cannot be reverse-engineered from whichever metric happened to look best.

2. Data audit

  • Group sizes in train, validation and test, in absolute counts. A group with 60 test rows cannot support a claim about its error rate.
  • Base rate of the positive label per group — this determines which fairness metrics are even simultaneously achievable.
  • Missingness per group. Differential missingness is a common hidden proxy: the field is blank for people the old process never processed.
  • Proxy leakage AUC, plus the top leaking features.
  • Provenance: who collected it, when, under what process, and what changed since. This is the datasheet (Gebru et al., 2018).

3. Model and fairness audit

Per group, report: n, base rate, selection rate, accuracy, TPR, FPR, PPV, and a confidence interval on each. Then intersections — the Gender Shades result was invisible at the level of "gender" or "skin tone" alone and only appeared at darker-skinned-and-female. Then slice by everything else that matters: age band, region, device, language, and time period.

Python
from fairlearn.metrics import MetricFrame, false_positive_rate, true_positive_ratefrom sklearn.metrics import accuracy_score, precision_scoreimport pandas as pdintersection = a_test.astype(str) + " | " + age_band_test.astype(str)mf = MetricFrame(    metrics={"accuracy": accuracy_score, "tpr": true_positive_rate,             "fpr": false_positive_rate, "ppv": precision_score},    y_true=y_test, y_pred=y_pred, sensitive_features=intersection,)report = mf.by_group.copy()report["n"] = pd.Series(intersection).value_counts()print(report.sort_values("n"))          # smallest cells first — that is where it breaks

4. Documentation

A model card (Mitchell et al., 2019) states intended use, out-of-scope uses, training data, per-group evaluation results, ethical considerations, and known limitations. The out-of-scope section is the one that does real work: COMPAS was built to assess supervision needs and used for detention decisions, and an explicit "this model must not be used to inform custodial decisions" line is the artefact that lets someone object later.

5. Continuous monitoring

Fairness is not a property you establish once. Populations shift, upstream data pipelines change, and models retrained on their own outputs drift.

MonitorCadenceAlert when
Per-group selection rateDailyRatio between any two groups falls below 0.80
Per-group TPR / FPRWeekly, as labels arriveGap widens by more than 5 points versus the audited baseline
Group composition of inputsWeeklyAny group's share shifts by more than 20% relative
Feature drift on top proxiesWeeklyDistribution shift beyond the audited range
Appeal and override rates by groupMonthlyOne group appeals or is overridden far more often — a sign the model is wrong about them

The last row is the most under-used signal in production. Humans overriding the model disproportionately for one group is direct field evidence of a disparity that your offline test set did not capture.

Where mitigation goes wrong

Failure modeWhat it looks likeFix
Metric shoppingCompute all fairness metrics, report the one that passesCommit to the criterion in step 1, in writing, before measuring
Fixing the symptomReweighing applied to a mislabelled targetDiagnose the entry point first; label problems need label fixes
Levelling downParity achieved by degrading the advantaged groupReport absolute per-group outcomes, not only ratios
Train-time onlyAudited once at launch, never againAutomated per-group monitoring with alert thresholds
Aggregate-only reporting"94% accurate" with no group breakdownBan aggregate-only metrics from review documents
Underpowered subgroupsA 40-row group reported to two decimal placesPublish confidence intervals; state when n is too small to conclude
Feedback contaminationRetraining on outcomes the model itself producedHold out a random control slice from model influence

What this looks like in a real project

Order matters more than technique. Run the proxy probe first, because its result tells you whether pre-processing has any chance of working. Fix the label second, because a wrong target invalidates everything downstream and is often the cheapest fix available. Only then reach for reweighing or a constraint, and pick the intervention point by asking a practical question rather than an academic one: do you control training (in-processing available?), and will you know the person's group at decision time (post-processing available?). Those two answers eliminate most of the menu.

Set a slack bound rather than exact parity, because exact parity buys almost nothing over a 5-point bound and costs a great deal of accuracy. Report absolute outcomes next to every ratio so levelling down cannot hide. Write the model card before launch rather than after, because the out-of-scope-uses section is the only thing that will still be there when someone repurposes your model in two years.

And keep the control slice. A small randomly selected fraction of decisions made outside the model's influence is the only clean data you will ever have for measuring what the model is actually doing to the world, as opposed to what it is doing to its own training set.