AI Ethics and Governance

Types of Bias in AI


Between 2014 and 2017 Amazon built an internal tool to score CVs one to five stars, the way it scores products. The training data was ten years of CVs the company had received, labelled by who actually got hired. Nobody wrote a rule about gender. Gender was not a column in the data.

By 2015 the engineers noticed the model was systematically downgrading CVs that contained the word women's — as in "women's chess club captain" — and penalising graduates of two all-women's colleges. Reuters reported the story in 2018; the project had been scrapped. Amazon had removed the protected attribute and the model found it anyway, because ten years of tech hiring at a company with a male-dominated engineering organisation had left a fingerprint on the vocabulary of successful CVs.

This is the shape of almost every real AI bias failure. Not a malicious rule. Not a variable someone forgot to delete. A model doing exactly what it was asked — reproduce the pattern in this data — on data that encoded a pattern nobody wanted reproduced.

To do anything about that, you need to know precisely where the bias entered, and you need to be able to measure it in numbers rather than adjectives.

Deleting the column does not delete the attributeDrop genderfrom the CV dataProxies survive:clubs, colleges, verbsModelreconstructsthe attributeLabelsencode whowas hired beforeSame downgrade,no rule writtenAmazon's tool penalised the token "women's" without anyone naming gender anywhere.
Bias enters through the data and the label, so a model with no protected column can still learn the protected decision.

Two completely different things are called "bias"

The word does double duty, and conflating the two senses produces conversations where nobody is wrong and nobody agrees.

Statistical biasSocietal / harm bias
MeansSystematic error: the model's average prediction is off the true value in a consistent directionThe system's behaviour disadvantages a group of people in a way that is unjust
Measured byExpected prediction minus truth; the "bias" in bias–variance decompositionDifferences in error rates, selection rates, or outcomes across groups
Fixed byA more flexible model, better features, more trainingChanging the data, the objective, the threshold, or the decision to deploy at all
Zero isAlways desirableNot a single well-defined target — see the impossibility result below

A model can have near-zero statistical bias and still be socially catastrophic. A recidivism model that perfectly predicts re-arrest is statistically unbiased with respect to re-arrest, and still encodes decades of unequal policing, because re-arrest is not the same event as re-offending. The model is accurate about the wrong thing.

A model is only ever unbiased with respect to the label you gave it. If the label is a biased measurement of what you actually care about, perfect accuracy reproduces the bias exactly.

The four doors bias walks through

People talk about "biased algorithms" as if the bias lives in the model weights. It almost never originates there. There are four distinct entry points, and each one needs a different fix, so distinguishing them is the whole job.

Door 1 — Historical data: the world was already unequal

The label records what happened, and what happened was shaped by human decisions that were themselves unequal. Amazon's label was "did we hire this person", and past hiring was skewed. A loan default model's label is "did this loan default", but you only observe defaults among people who were approved — the applicants that past loan officers rejected have no outcome at all. That is selection bias in the label, and it means the training set is a censored view of reality.

A related variant is measurement bias: your label is a stand-in for the thing you care about, and the stand-in is differentially wrong across groups. Arrest is a proxy for crime. Diagnosis codes are a proxy for illness. Both are filtered through who gets stopped and who can afford to see a doctor.

Door 2 — Representation: some people are barely in the data

If a group is 3% of your training set, the loss function barely notices when the model gets that group wrong. Reducing error on the other 97% pays better, so gradient descent optimises the majority.

The clearest documented case is Gender Shades (Buolamwini and Gebru, 2018), which audited commercial gender-classification APIs from Microsoft, IBM and Face++. Error rates on lighter-skinned men were at worst 0.8%. Error rates on darker-skinned women reached 34.7% — roughly a 43-fold difference within the same product. The paper also checked the benchmarks everyone had been reporting on: one widely used face benchmark was 79.6% lighter-skinned, another 86.2%. So the industry's own scoreboard could not have detected the problem. That is evaluation bias — an unrepresentative test set makes a broken model look fine.

Representation failures have a sibling: aggregation bias, where the groups are all present in adequate numbers but a single model is forced to serve populations that genuinely differ. HbA1c, the standard blood marker for diabetes, has documented differences in its relationship to blood glucose across ethnic groups; a single fixed threshold applied to everyone under-diagnoses some populations. One model, one threshold, several different underlying relationships — the average fits nobody well.

Door 3 — Proxy variables: the attribute you deleted is still there

Remove race from the feature set and the model can still use postcode, surname, the name of a secondary school, which shops a card is used at, the phone's operating system, or the phrase "women's chess club". In a segregated housing market, postcode is not correlated with race by accident — it is a near-recording of it.

This deserves its own section below, because it is where most well-intentioned mitigation efforts fail.

Door 4 — How the output is used

A model can be well-calibrated, evenly representative, proxy-clean, and still cause harm because of what humans do with the number.

COMPAS, the risk tool at the centre of the ProPublica investigation, was designed by its vendor to identify needs for supervision and treatment. Courts used it to inform detention and sentencing decisions. Same score, entirely different consequence. That is deployment bias — a mismatch between the context a model was validated in and the context it is used in.

A second form is the feedback loop. Lum and Isaac's 2016 analysis of predictive policing applied a PredPol-style algorithm to Oakland drug-crime data. Because the training data was recorded arrests, and arrests concentrate where officers are sent, the model directed patrols back to the same neighbourhoods, generating more arrests there, which strengthened the model's belief. Public-health estimates of actual drug use in Oakland were far more evenly distributed across the city than the arrest data. The model was not predicting crime; it was predicting policing, and then causing it.

Algorithmic bias proper — introduced by the model or objective rather than the data — is the smallest door. An objective maximising overall accuracy will sacrifice a 5% minority; a ranking loss that rewards clicks rewards outrage.

TypeDoorSignature symptomDetected by
Historical / label bias1Model is accurate but reproduces a known past injusticeAsking what the label actually measures, and who produced it
Selection bias1No outcomes exist for people the old process rejectedComparing applicant population to training population
Measurement bias1Proxy label diverges from the true target differently per groupFinding a second, independent measure of the true target
Representation bias2Error rate far worse on a small subgroupPer-group metrics, never aggregate accuracy
Aggregation bias2One model underperforms every group it servesFitting separate models and comparing
Evaluation bias2Benchmark performance excellent, field performance poorAuditing the demographic composition of the test set
Proxy leakage3Removing the protected attribute changes nothingTraining a model to predict the attribute from remaining features
Deployment bias4Used for a decision it was never validated forComparing documented intended use to observed use
Feedback loop4Predictions get "more accurate" over time on their own dataHolding out a region or cohort from the model's influence

Why deleting the protected attribute does not delete its influence

The intuition "if the model can't see race, it can't discriminate by race" is called fairness through unawareness. It fails whenever any combination of the remaining features can reconstruct the attribute — which, in real datasets, is almost always.

Work it out. Suppose you drop ethnicity from a lending model but keep postcode. Take three postcodes in a segregated city:

PostcodeShare of residents in Group AHistorical approval rate
PC-188%41%
PC-252%63%
PC-39%78%

The model never sees group membership. It sees that applicants from PC-1 were approved 41% of the time historically and learns to score them down. An applicant from PC-1 has an 88% chance of being in Group A. The model has not "guessed" anyone's ethnicity — it has simply learned a variable that carries most of the same information. Statistically, if you know the postcode you can predict group membership with high confidence, so any function of postcode is also, partly, a function of group.

Worse, unawareness makes things harder to fix. Without the attribute you cannot compute per-group error rates, so the discrimination becomes unmeasurable while remaining fully operative. This is why regulators increasingly permit collecting protected attributes strictly for the purpose of testing for discrimination.

You need the protected attribute to detect discrimination. Deleting it does not remove the bias; it removes your ability to see it.

The proxy leakage test

There is a direct empirical check. Take your feature matrix with the protected attribute removed, and try to predict the protected attribute from it. If a simple model succeeds, your features contain the attribute.

Python
from sklearn.ensemble import RandomForestClassifierfrom sklearn.model_selection import cross_val_score# X_no_attr: every feature the real model will use, protected attribute dropped# a: the protected attribute (held aside purely for auditing)auc = cross_val_score(    RandomForestClassifier(n_estimators=300, min_samples_leaf=20),    X_no_attr, a, cv=5, scoring="roc_auc").mean()print(f"Attribute recoverable from features: AUC = {auc:.3f}")# AUC 0.50  -> features carry no information about the attribute# AUC 0.65  -> meaningful leakage# AUC 0.85+ -> the attribute is effectively still in your dataset

Then find which features leak, by permuting each one and seeing how much the AUC drops. In credit and hiring datasets the usual culprits are postcode, first name, school or university, employer name, and any text field at all.

Knowing a feature is a proxy does not automatically mean you delete it. Postcode may genuinely predict repayment for reasons unrelated to group. The decision is whether the predictive power that survives after group information is stripped out is worth keeping — and that is a judgement, not a computation.

Three documented cases, read as mechanism

COMPAS (ProPublica, 2016)

ProPublica obtained risk scores for roughly 7,000 people arrested in Broward County, Florida in 2013–2014 and checked who was re-arrested within two years. Among defendants who did not re-offend, 44.9% of Black defendants had been labelled high risk versus 23.5% of white defendants. Among those who did re-offend, 47.7% of white defendants had been labelled low risk versus 28.0% of Black defendants. The vendor, Northpointe, replied that the tool was equally calibrated: a score of 7 meant roughly the same probability of re-offending regardless of race, which it did.

Both parties were reporting correct arithmetic on the same data. The next section explains why they had to disagree.

Optum's healthcare risk score (Obermeyer et al., Science, 2019)

A commercial algorithm used across US health systems to enrol patients in extra care management ranked patients by predicted future healthcare cost, on the reasonable-sounding theory that expensive patients are sick patients. Cost is a proxy for need — and a group-dependent one, because less money is spent on Black patients with the same conditions, for reasons ranging from access to distrust to under-treatment.

The consequence: at any given risk score, Black patients were measurably sicker than white patients with the same score. The researchers estimated that fixing the label — predicting actual health conditions rather than cost — would raise the share of Black patients auto-enrolled in the extra-care programme from 17.7% to 46.5%. By industry estimates, tools of this kind are applied to about 200 million people in the US each year. This one contained no race variable. It was a pure Door 1 measurement-bias failure, fixable by changing the label rather than the model.

Dutch childcare benefits (toeslagenaffaire)

The Dutch tax authority ran a risk model to flag childcare-benefit fraud. Having a second nationality was among the factors that raised risk. Around 26,000 families were wrongly accused, ordered to repay tens of thousands of euros, driven into debt, and in some cases had children removed. The Dutch data protection authority found the nationality processing unlawful, and the government resigned in January 2021. This is Door 3 and Door 4 together: a proxy in the features, and an output used as near-automatic enforcement with no meaningful appeal.

CasePrimary doorWhat the fix actually was
Amazon CV scorer1 (historical label) + 3 (text proxies)None available; project cancelled
Gender Shades2 (representation + evaluation)Rebalance training and benchmark sets; report per-subgroup error
Optum risk score1 (measurement — wrong label)Change the prediction target from cost to health conditions
Predictive policing4 (feedback loop)Stop training on arrests; no purely technical fix exists
Dutch benefits3 + 4Remove nationality; restore human appeal; compensation scheme

Measuring fairness: three metrics, defined precisely

Everything below comes out of a per-group confusion matrix. Split your test set by group, and for each group count:

Text
                      Actually positive    Actually negativePredicted positive          TP                   FPPredicted negative          FN                   TNSelection rate = (TP + FP) / N        share flaggedTPR (recall)   = TP / (TP + FN)       of those who are positive, share caughtFPR            = FP / (FP + TN)       of those who are negative, share wrongly flaggedPPV (precision)= TP / (TP + FP)       of those flagged, share truly positiveprevalence p   = (TP + FN) / N        base rate of the outcome in this group
Fairness criterionRequires equal across groupsPlain EnglishRight when
Demographic paritySelection rateThe same share of each group gets the positive decisionThe label itself is suspect, or the goal is redistributive (outreach, shortlisting)
Equal opportunityTPRAmong people who genuinely qualify, the same share are foundMissing a qualified person is the main harm (hiring, screening for treatment)
Equalised oddsTPR and FPRBoth error types fall equally on each groupFalse positives and false negatives are both costly (bail, fraud enforcement)
Predictive parity / calibrationPPV (or score-conditional risk)A given score means the same thing whoever you areA human acts on the score's face value (clinicians, judges, underwriters)

The impossibility result, with the arithmetic

Take a screening tool used on two groups of 1,000 people each. Group A has a base rate of 50% (500 true positives); Group B has a base rate of 25% (250). Base rates differ — which in the real world is the norm, and is itself usually a consequence of historical inequity.

First, tune the tool to be calibrated: make PPV equal at 0.60 in both groups, and equal opportunity holds too, TPR = 0.60 in both.

Group A (p = 0.50)Actually +Actually −Group B (p = 0.25)Actually +Actually −
Flagged300200Flagged150100
Not flagged200300Not flagged100650
  • Group A: PPV = 300/500 = 0.60; TPR = 300/500 = 0.60; FPR = 200/500 = 0.40; selection rate = 500/1000 = 50%
  • Group B: PPV = 150/250 = 0.60; TPR = 150/250 = 0.60; FPR = 100/750 = 0.133; selection rate = 250/1000 = 25%

Calibration holds. Equal opportunity holds. But an innocent person in Group A is three times as likely to be wrongly flagged (40% versus 13.3%), so equalised odds fails badly — and demographic parity fails too, 50% versus 25%.

Now force demographic parity instead: flag exactly 375 people in each group, by taking the top 375 scores within each group.

Group A, 375 flaggedActually +Actually −Group B, 375 flaggedActually +Actually −
Flagged255120Flagged180195
Not flagged245380Not flagged70555
  • Group A: selection rate 37.5%; PPV = 255/375 = 0.68; TPR = 255/500 = 0.51; FPR = 120/500 = 0.24
  • Group B: selection rate 37.5%; PPV = 180/375 = 0.48; TPR = 180/250 = 0.72; FPR = 195/750 = 0.26

Demographic parity now holds exactly, and FPR is nearly equal (24% versus 26%). But PPV has split apart: a "flagged" label is right 68% of the time for Group A and only 48% for Group B. A caseworker who trusts the flag equally is now wrong far more often about Group B. Meanwhile Group B's qualified members are found at 72% versus Group A's 51% — equal opportunity has broken in the other direction.

Why this is a theorem, not bad luck

The four quantities are algebraically locked together. For any group with prevalence pp:

FPR  =  TPR⋅p1−p⋅1−PPVPPV\text{FPR} \;=\; \text{TPR}\cdot\frac{p}{1-p}\cdot\frac{1-\text{PPV}}{\text{PPV}}

Check it against the first table. Group A: 0.60×0.50.5×0.40.6=0.400.60 \times \frac{0.5}{0.5} \times \frac{0.4}{0.6} = 0.40. Group B: 0.60×0.250.75×0.40.6=0.1330.60 \times \frac{0.25}{0.75} \times \frac{0.4}{0.6} = 0.133. Both match.

Now read the formula as a constraint. Fix TPR equal across groups and fix PPV equal across groups; the only remaining term is p/(1−p)p/(1-p). If the two groups have different prevalence, that term differs, so FPR must differ. There is no model clever enough to escape it. Chouldechova (2016) proved exactly this for predictive parity versus equalised error rates; Kleinberg, Mullainathan and Raghavan (2016) proved the parallel result for calibration versus balance in both classes. The only escapes are perfect prediction (every error rate zero) or equal base rates — neither of which you get.

Demographic parity, equal opportunity and predictive parity cannot generally hold at once. Choosing between them is not an engineering decision with a correct answer; it is a decision about who bears which kind of error.

That last point is the one that changes practice. If a false positive means a person spends months in pre-trial detention, and a false negative means a small increase in risk to the public, then the people who bear the false positives are not the same people who bear the false negatives — and they should have a say in the trade. A team of engineers picking a fairness metric in a sprint planning meeting is quietly making a distributive-justice decision on behalf of people who were not in the room.

Where people get this wrong

BeliefWhy it fails
"We removed the protected attribute, so it's fair."Proxies reconstruct it, and you have now lost the ability to measure the disparity.
"Our accuracy is 94% across the board."Aggregate accuracy hides subgroup collapse. Gender Shades sat behind vendor-reported accuracies above 90%.
"Just balance the dataset 50/50."Fixes Door 2 only. It does nothing about a mislabelled target (Optum) or a feedback loop (predictive policing).
"The data is objective; the model just reflects reality."The data reflects a measurement process run by people. Arrests are not crimes; spend is not sickness.
"We'll pick the fairest metric."There isn't one. Picking demographic parity over predictive parity is a choice about which harm you accept.
"A human reviews every decision, so it's fine."Automation bias is well documented: reviewers overwhelmingly confirm the machine's suggestion, especially under time pressure.
"Small disparity, so it doesn't matter."A 2-point disparity applied to 40 million decisions a year is a large number of people. Scale converts small rates into mass effects.

What to do on Monday

The practical consequence of all this is a short set of habits that are cheap to adopt and catch most of what goes wrong.

Interrogate the label before the model. Write down, in one sentence, what you actually want to predict and what you are actually predicting. If those sentences differ — need versus cost, crime versus arrest, competence versus past hiring decisions — you have found your biggest problem before writing a line of training code, and no fairness algorithm downstream will fix it.

Never report an aggregate metric alone. Every evaluation table gets a group column. Include group sizes: a 4% subgroup with 60 test examples gives you a confidence interval so wide the metric is nearly meaningless, and that itself is a finding.

Run the proxy leakage test and record the AUC. It takes ten minutes and turns "we removed race" into a number you can defend.

Choose your fairness criterion in writing, with a reason, before you tune. Write the sentence: "We prioritise equal false-negative rates because the cost of missing a qualified applicant falls on the applicant, who has no recourse, while the cost of a false positive falls on us in interview time." That sentence is auditable. "We used the fairness library" is not.

Get the affected group into the decision. The impossibility result means someone is choosing which error a real person absorbs. If that someone is only the build team, the ethics of the system is being decided by whoever happened to be staffed on it.