AI Ethics and Governance

Building Ethical AI Checklists


Here is a checklist item signed off on thousands of real projects:

Text
□ Have we considered bias in this system?     [x] Yes  -- reviewed 14 March, J. Okafor

The tick was honest. Someone did sit down and consider bias: they split the CV-screening model's accuracy by every demographic field in the applicant database, found nothing above a two-point gap, and moved on.

Six months after launch, an external audit found graduates of women's colleges landing in the bottom quartile 2.3 times as often as graduates of comparable co-educational institutions. The model had never seen a gender field — the company had deliberately not collected one. It had seen the institution name, the sports under "interests", and the phrase "women in engineering society" in the activities free text.

Nothing in that review was dishonest and nothing in it was useful. "Have we considered bias?" has no answer that fails: a team that considered nothing would have ticked the same box, because considering is not an artefact. Here is the question that would have caught it — not a more strongly worded version of the same sentiment:

For each protected attribute we do not collect, which features in the model predict it, and how well? Report the AUC.

That one has a failing answer you can recognise when you hear it: "we don't collect gender, so the model can't discriminate on gender." That sentence is the sound of a control not working.

What follows is a working set of such questions, ordered by where each problem is still cheap to fix.

A signature line versus a controlSignature line• "Have we considered bias?" Yes or no• Signed once, at the end• No artefact attached to the answer• Has no failing stateControl• Which subgroups, at what sample size?• Blocks the gate until answered• Names the file that proves it• Can fail and stop the launch
A checklist changes decisions only where an item can come back "no" and something downstream is not allowed to proceed.

What separates a control from a signature line

A checklist question earns its place if it does three things. It names an artefact — a number, a file path, a person's name, something existing outside the reviewer's head. It has a recognisable failing answer you can write down in advance; if you cannot, the question is decoration. It has an owner and a date: not "the team", a person, and when they last produced the artefact.

Most published AI ethics checklists fail all three: they are lists of values, not controls. The translation is mechanical:

Principle-shaped questionWhy it cannot failMechanism-shaped replacementWhat it catches
Is the system fair?Every team believes its system isWhich fairness metric did we commit to before measuring, who signed that, and when?Metric shopping — computing six, reporting the one that passed

A checklist question that cannot be answered wrongly is not a control. It is a signature line, and its only function is to distribute blame after the fact.

Before any data is touched

These questions can still stop a project, which is exactly why they get skipped — by the time anyone convenes a review, the work has a budget and a launch date. A choice corrected here costs an afternoon; the same correction after an audit costs a withdrawal.

QuestionWhat it catchesWhat a bad answer sounds like
State the decision as one sentence with a verb and a person: "this system ___ a ___".Systems later defended as "only a recommendation" while nobody ever overrides them"It surfaces insights to support the hiring team's judgement."
Describe the harm of a wrong answer in each direction, in two sentences, each naming a person rather than a metric.Asymmetry blindness — treating a false positive and a false negative as the same kind of error"Precision drops, so we'd see more churn in the funnel."
What is the non-model baseline, and what does it score?Projects where twelve explicit rules perform within a point, fully auditable, with no proxy risk"There isn't really a baseline; this is a new capability."
Who is affected without ever choosing to interact with this?Indirect stakeholders — the rejected applicant, the employee whose role is reshapedA list of paying customers and internal users.

One more belongs here: what single number will be reported as success, and what harm is invisible to it? "Throughput, up 40%" rewards the exact behaviour you are worried about.

The second row is the one people underestimate. Force both harm sentences to name a person and the asymmetry becomes obvious: a false positive costs the bank a write-off, a false negative costs a person a house. Every threshold decision downstream depends on having said that in writing, once.

Data: the cheapest place to catch anything

Can this test set support the claim you are about to make?

"Is the data representative?" is unanswerable; the version with an answer is arithmetic. Say you intend to report that true positive rates differ by no more than 10 percentage points across groups. The smallest gap a test set can reliably detect, at 5% significance and 80% power, is

Δmin⁡=(zα/2+zβ)p(1−p)(1nA+1nB)\Delta_{\min} = (z_{\alpha/2} + z_{\beta})\sqrt{p(1-p)\left(\tfrac{1}{n_A} + \tfrac{1}{n_B}\right)}

with zα/2=1.96z_{\alpha/2} = 1.96, zβ=0.84z_\beta = 0.84 and pp the rate you expect. Run it on your actual cell counts before you run any model:

Python
import numpy as npfrom scipy.stats import normdef minimum_detectable_gap(n_ref, n_group, p=0.80, alpha=0.05, power=0.80):    """Smallest true rate difference this test set could reliably detect."""    z = norm.ppf(1 - alpha / 2) + norm.ppf(power)        # 1.960 + 0.842    se = np.sqrt(p * (1 - p) * (1 / n_ref + 1 / n_group))    return z * seCLAIM_TOLERANCE = 0.10          # the gap we intend to claim is not exceededfor name, n in [("Group A", 4000), ("Group B", 1000), ("Group B, age 60+", 58)]:    gap = minimum_detectable_gap(4000, n)    verdict = "CANNOT SUPPORT THE CLAIM" if gap > CLAIM_TOLERANCE else "ok"    print(f"{name:<18} n={n:<5} smallest detectable gap = {gap:.3f}  {verdict}")
Text
Group A            n=4000  smallest detectable gap = 0.025  okGroup B            n=1000  smallest detectable gap = 0.040  okGroup B, age 60+   n=58    smallest detectable gap = 0.148  CANNOT SUPPORT THE CLAIM

The third row is the interesting one. With 58 rows a real 14-point gap would fail to reach significance, so the audit reports "no significant disparity" — technically true and completely misleading. The honest output is not a pass but "we cannot say": collect more data for that cell, or stop publishing a number for it.

Can the model reconstruct what you refused to collect?

Drop the protected attribute from the feature matrix, train a small classifier to predict it from what remains, and read off the AUC. Around 0.50 means the features carry nothing; 0.65 is meaningful leakage; above 0.85 the attribute is effectively still in the dataset, and "we don't use it" describes the schema, not the behaviour. The failing answer is the confident one — "we removed the sensitive columns" — because removal is a schema operation and prediction is what the model does.

Is the label measuring what you think it measures?

In a hiring model the label is almost never "was a good employee". It is "was hired by the previous process", or "was still employed after 12 months" — measurements of the old system's preferences and of who could afford to stay. Rebalancing a dataset whose label is the wrong measurement gives you a balanced wrong measurement. So: write down in one sentence what the label literally records, without using the word it is named after. If good_hire records "received an offer from a panel of four managers", say that. Half the time, writing the sentence changes the project.

Ask, in the same breath, for the missingness rate of every feature broken down by group. Differential missingness is a common hidden proxy: a field is blank precisely for the people the old process never processed, and a tree model happily learns that "blank" predicts the outcome. The bad answer is "we impute the median" without a per-group table — imputing over a group with 40% missingness quietly assigns that whole group the majority's typical value.

Development: commit to the metric before you see the numbers

Every fairness metric passes for some model. Here is one classifier on two groups — 4,000 people in Group A, 1,000 in Group B, base rates 40% and 25%.

GroupTPFNFPTNn
A1,2004002402,1604,000
B150100756751,000

Every standard metric follows from those cells:

MetricGroup AGroup BGap or ratioVerdict at the usual threshold
Selection rate (1,440/4,000 vs 225/1,000)0.3600.225ratio 0.63Fails the 0.80 four-fifths rule
True positive rate (1,200/1,600 vs 150/250)0.7500.60015.0 ptsFails a 5-point bound
False positive rate (240/2,400 vs 75/750)0.1000.1000.0 ptsPasses, exactly
Precision (1,200/1,440 vs 150/225)0.8330.66716.7 ptsFails predictive parity
Accuracy (3,360/4,000 vs 825/1,000)0.8400.8251.5 ptsPasses

A team reporting "false positive rates are identical to three decimal places, and accuracy differs by 1.5 points" is telling the literal truth about a model that misses 40% of qualified Group B applicants against 25% of Group A's. No fabrication is required. The only defence is procedural: the metric is chosen and recorded before the numbers exist, with the reason, and the review reports all of them anyway.

Choose the metric before you see the numbers, and write down who chose it and why. Otherwise the metric you report is simply the one that passed.

Two more questions with teeth here. "Show me the explanation a rejected person receives" — not the tooling, the actual sentence. A force plot is an engineering artefact; "declined mainly because of a debt-to-income ratio of 0.61 and two missed payments in 24 months" is an explanation, and if the model cannot produce the second, the stakes may justify a simpler model. And "what does it do with a missing or out-of-range field?" Feed it a negative income and a 200-year employment history: systems that degrade gracefully refuse, systems that fail dangerously return a confident score built on an imputed value.

Testing: does the evidence come from the world or from the model?

One question dominates here and is rarely asked: was this test data influenced by the model you are testing? If your evaluation set is made of applicants the current system already screened, you are measuring how well the model reproduces itself. The fix is a small randomly selected control slice held outside the model's influence, and it is the only clean data you will ever have.

Alongside it: report groups by intersection, not one attribute at a time — a model can show a 2-point gap by sex and a 3-point gap by skin tone while the darker-skinned-women cell is off by 30 points. Sort cells by size ascending and read the small ones first. Check the metrics on realistic input — typos, truncated fields, wrong units — not the cleaned test set. And for models trained on personal records, check whether training data can be recovered; memorisation is a privacy failure that never appears in an accuracy metric.

Deployment readiness: accountability you can test

QuestionWhat it catchesWhat a bad answer sounds like
Name the individual accountable for this system's outcomes.Diffusion — a committee that meets quarterly and owns nothing"The AI governance board."
How is it turned off, who is authorised, how long does it take, when was that last exercised end to end?Kill switches that exist as a config flag nobody has ever flipped"We'd push a change to disable the endpoint." (Never tested. Needs a release. Four hours.)
During the pilot, how many people appealed, and how many appeals were upheld?Appeal routes that technically exist and are structurally unusable"There's a contact address on the rejection notice."

Two more: if a person corrects a field about themselves, does the decision get re-run — correction rights that update a database but not the outcome are not rights. And what is the fallback while the system is off, since a system nothing else can replace cannot really be switched off. Note too that a zero appeal rate across 10,000 pilot decisions is not evidence everyone was satisfied. It is evidence nobody could find the route.

After launch: the questions nobody schedules

A pre-launch review catches only what was visible before launch. Much real harm appears afterwards, through drift — the deployed population diverging from the training population — and through feedback, where the model's decisions shape the data it is next trained on.

The feedback effect is worth the arithmetic. Of every 5,000 applicants above, 4,000 are Group A and 1,000 Group B, so Group B is 20% of the pool. Repayment outcomes exist only for approved applicants, so the next training set gets 1,440 labelled rows from Group A and 225 from Group B — Group B's share of the labelled data being

2251,440+225=2251,665=13.5%\frac{225}{1{,}440 + 225} = \frac{225}{1{,}665} = 13.5\%

One retraining cycle has cut Group B's representation from 20% to 13.5% — a relative drop of a third — with nobody changing a line of code. Fewer rows means a worse fit, which means lower scores, fewer approvals, and fewer rows again. The system is manufacturing the evidence for its own conclusions.

Anything you measure only before launch, you have measured on the one distribution the system was never going to face.

SignalCadenceAlarm whenWhat it catches
Selection rate per groupDailyAny pairwise ratio drops below 0.80Drift; pipeline changes
TPR and FPR per groupWeekly, as outcomes arriveGap widens by over 5 points versus the certified baselineDegradation invisible in aggregate accuracy
Group mix of the next training setEvery retrainDiverges from the applicant pool by over 10% relativeThe feedback loop, before it compounds
Human override rate per groupMonthlyOne group overridden far more oftenField evidence the model is wrong about a group

Re-run the proxy probe quarterly too; new data sources bring new proxies. But the override row is the most under-used signal in production — a human overturning the model disproportionately for one group is direct field evidence of a disparity your test set did not contain. Two questions round this out: when was the incident plan last rehearsed (a plan never run is a document, not a capability), and what changed in the applicable regulation, since a system certified against last year's rules is not certified.

Impact assessment: turning "harm" into rows

Impact assessment answers something narrower than "is this ethical?": if this system behaves as designed, and if it fails, who is harmed, how badly, and what evidence would settle it? Severity must be anchored to consequence — Critical for irreversible harm to health, liberty or livelihood; High for a material opportunity lost and hard to recover, such as credit or employment; Medium for harm an appeal can undo; Low for friction with no lasting effect. Without those anchors, severity becomes a measure of how embarrassing something would be for the organisation. Each row must then be falsifiable: "mitigation" names a mechanism, "evidence" names something you can go and look at.

HarmSeverityEvidence that would settle itMitigation
Qualified applicants declined at different rates by groupHighPer-group TPR with confidence intervals, on the control sliceConstrained retraining with a 5-point bound; monthly monitoring
Decision cannot be explained to the person it affectsHighTwenty sampled rejection notices, read by someone outside the teamReason codes generated at scoring time, stored with the decision
Sensitive attributes recoverable from retained featuresMediumProxy probe AUC on the live feature setDrop or orthogonalise the top leaking features; document the trade

Store the whole assessment as a record whose fields are pointers to evidence, and gate deployment on the presence of that evidence rather than on somebody's verdict:

Python
REQUIRED_EVIDENCE = {    "problem":        ["decision_sentence", "non_model_baseline_score"],    "data":           ["subgroup_counts_path", "proxy_probe_auc",                       "missingness_by_group_path"],    "fairness":       ["metric_committed_on", "metric_chosen_by",                       "metric_frame_path", "smallest_reported_cell_n"],    "explainability": ["sample_rejection_notice_path"],    "privacy":        ["dpia_path", "lawful_basis", "retention_days"],    "accountability": ["accountable_person", "kill_switch_last_tested_on",                       "fallback_process_doc"],    "monitoring":     ["alert_config_path", "control_slice_pct", "review_due_on"],}class BlockedDeployment(Exception):    passdef gate(record: dict) -> str:    """Fail on ABSENT EVIDENCE, never on a self-reported verdict."""    missing = {area: [k for k in keys if not record.get(area, {}).get(k)]               for area, keys in REQUIRED_EVIDENCE.items()}    missing = {area: keys for area, keys in missing.items() if keys}    if missing:        raise BlockedDeployment(missing)    return "clear to deploy"

What matters is what the gate checks. A review object whose methods return {"status": "passed"} is the signature-line problem rendered in code: it records that someone asserted a verdict. This version cannot be satisfied by assertion, only by producing a file, a number, or a name. It also composes with the paperwork you owe anyway — a regulatory risk assessment and a data protection impact assessment want most of these same fields, and should be one document.

When the checklist has no right answer

Some conflicts are not oversights. Two stakeholders want genuinely incompatible things, and no checklist item resolves that — only an explicit design conversation does. Value-sensitive design, developed by Batya Friedman and colleagues, is that conversation held deliberately.

It begins by separating direct stakeholders, who use the system, from indirect stakeholders, affected without using it. In lending the loan officer is direct and the applicant indirect — exactly backwards from where the harm falls, which is why indirect stakeholders get designed for last, if at all. Then you map the tensions, writing down the mechanism rather than the sentiment:

TensionThe mechanism underneath itResolved badlyA defensible resolution
Accuracy vs. demographic parityThe loss-minimising classifier is not the parity-satisfying one when base rates differDemand exact parity, then abandon it when accuracy dropsA slack bound justified in advance, with per-group outcomes reported beside the ratio
Privacy vs. measurable fairnessYou cannot measure a gap on an attribute you refuse to collectCollect nothing, then claim fairness cannot be assessedCollect it under a separate lawful basis, hold it segregated, use it only for aggregate audit — never as a feature
Automation speed vs. contestabilityAn automated decision has no natural point for a person to interveneAdd a review queue nobody has time to workAutomate reversible decisions; route irreversible ones to a human with a time budget

The third step is iteration with real stakeholders, and here is the failure mode that gives the field its bad name. Ethics-washing is rarely fabricated results. It is a genuine consultation, honestly conducted, held after the design is locked: everyone is sincere, nothing they say can change anything, and the exercise produces legitimacy instead of information.

If no possible outcome of the consultation would have changed the design, it was not a review. Ask what would have to be said in that room to delay the launch — if there is no answer, you already know.

Running this so that it changes decisions

When a question does fail, do not settle it in a corridor. State the fact and the conflict separately — "the TPR gap by age is 15 points" is the fact, "equalising it costs 2 points of accuracy" is the conflict, and most arguments happen because people dispute different halves. Write down at least three options, one of which is not shipping; if "do not ship" is never on the list, this is not a decision procedure. Check what the law allows: group-specific thresholds are unlawful in several jurisdictions for employment and credit however well they perform. Then decide, record the reasoning, and set a review date — when the question returns in eighteen months with different people in the room, that record is the difference between a decision and a rumour.

Checklists themselves fail in recognisable ways. Two are obvious once named: answers arriving as paragraphs of prose mean the review records opinions, not artefacts; and a review scheduled after the launch date has already been decided. The rest:

SymptomWhat it meansWhat to do
A 100% pass rate across two years of reviewsThe questions have no failing answersFor each item, write the bad answer; delete any item where you cannot
The reviewer built the thingNobody is positioned to say noRotate reviewers across teams; the reviewer's job is to ask for the artefact
It has never stopped anythingIt is a signature-collection processThe single best indicator; treat it as urgent

Put the questions where the work happens, not in a policy document. Data questions belong in the dataset card that must exist before a training run; development questions in the pull request template; deployment questions in the release gate, as code that refuses to proceed when evidence is missing. A question someone has to remember to ask gets asked when there is time, which is never. Cap each gate at five to nine items and retire anything an automated test now covers — a list grown to 120 items, one per past incident, gets completed in ten minutes with a ruler.

Then accept the cost. The first time one of these questions delays a launch by three weeks, someone senior will ask whether the process has become too heavy. The answer is that a control which never fires is not a control, and three weeks is the whole price of finding out before a regulator does.