Course Content
AI Ethics and Governance
3 sections · 7 lessons
Building Ethical AI Checklists
Here is a checklist item signed off on thousands of real projects:
□ Have we considered bias in this system? [x] Yes -- reviewed 14 March, J. OkaforThe tick was honest. Someone did sit down and consider bias: they split the CV-screening model's accuracy by every demographic field in the applicant database, found nothing above a two-point gap, and moved on.
Six months after launch, an external audit found graduates of women's colleges landing in the bottom quartile 2.3 times as often as graduates of comparable co-educational institutions. The model had never seen a gender field — the company had deliberately not collected one. It had seen the institution name, the sports under "interests", and the phrase "women in engineering society" in the activities free text.
Nothing in that review was dishonest and nothing in it was useful. "Have we considered bias?" has no answer that fails: a team that considered nothing would have ticked the same box, because considering is not an artefact. Here is the question that would have caught it — not a more strongly worded version of the same sentiment:
For each protected attribute we do not collect, which features in the model predict it, and how well? Report the AUC.
That one has a failing answer you can recognise when you hear it: "we don't collect gender, so the model can't discriminate on gender." That sentence is the sound of a control not working.
What follows is a working set of such questions, ordered by where each problem is still cheap to fix.
What separates a control from a signature line
A checklist question earns its place if it does three things. It names an artefact — a number, a file path, a person's name, something existing outside the reviewer's head. It has a recognisable failing answer you can write down in advance; if you cannot, the question is decoration. It has an owner and a date: not "the team", a person, and when they last produced the artefact.
Most published AI ethics checklists fail all three: they are lists of values, not controls. The translation is mechanical:
| Principle-shaped question | Why it cannot fail | Mechanism-shaped replacement | What it catches |
|---|---|---|---|
| Is the system fair? | Every team believes its system is | Which fairness metric did we commit to before measuring, who signed that, and when? | Metric shopping — computing six, reporting the one that passed |
A checklist question that cannot be answered wrongly is not a control. It is a signature line, and its only function is to distribute blame after the fact.
Before any data is touched
These questions can still stop a project, which is exactly why they get skipped — by the time anyone convenes a review, the work has a budget and a launch date. A choice corrected here costs an afternoon; the same correction after an audit costs a withdrawal.
| Question | What it catches | What a bad answer sounds like |
|---|---|---|
| State the decision as one sentence with a verb and a person: "this system ___ a ___". | Systems later defended as "only a recommendation" while nobody ever overrides them | "It surfaces insights to support the hiring team's judgement." |
| Describe the harm of a wrong answer in each direction, in two sentences, each naming a person rather than a metric. | Asymmetry blindness — treating a false positive and a false negative as the same kind of error | "Precision drops, so we'd see more churn in the funnel." |
| What is the non-model baseline, and what does it score? | Projects where twelve explicit rules perform within a point, fully auditable, with no proxy risk | "There isn't really a baseline; this is a new capability." |
| Who is affected without ever choosing to interact with this? | Indirect stakeholders — the rejected applicant, the employee whose role is reshaped | A list of paying customers and internal users. |
One more belongs here: what single number will be reported as success, and what harm is invisible to it? "Throughput, up 40%" rewards the exact behaviour you are worried about.
The second row is the one people underestimate. Force both harm sentences to name a person and the asymmetry becomes obvious: a false positive costs the bank a write-off, a false negative costs a person a house. Every threshold decision downstream depends on having said that in writing, once.
Data: the cheapest place to catch anything
Can this test set support the claim you are about to make?
"Is the data representative?" is unanswerable; the version with an answer is arithmetic. Say you intend to report that true positive rates differ by no more than 10 percentage points across groups. The smallest gap a test set can reliably detect, at 5% significance and 80% power, is
with zα/2=1.96, zβ=0.84 and p the rate you expect. Run it on your actual cell counts before you run any model:
1import numpy as np2from scipy.stats import norm34def minimum_detectable_gap(n_ref, n_group, p=0.80, alpha=0.05, power=0.80):5 """Smallest true rate difference this test set could reliably detect."""6 z = norm.ppf(1 - alpha / 2) + norm.ppf(power) # 1.960 + 0.8427 se = np.sqrt(p * (1 - p) * (1 / n_ref + 1 / n_group))8 return z * se910CLAIM_TOLERANCE = 0.10 # the gap we intend to claim is not exceeded1112for name, n in [("Group A", 4000), ("Group B", 1000), ("Group B, age 60+", 58)]:13 gap = minimum_detectable_gap(4000, n)14 verdict = "CANNOT SUPPORT THE CLAIM" if gap > CLAIM_TOLERANCE else "ok"15 print(f"{name:<18} n={n:<5} smallest detectable gap = {gap:.3f} {verdict}")Group A n=4000 smallest detectable gap = 0.025 okGroup B n=1000 smallest detectable gap = 0.040 okGroup B, age 60+ n=58 smallest detectable gap = 0.148 CANNOT SUPPORT THE CLAIMThe third row is the interesting one. With 58 rows a real 14-point gap would fail to reach significance, so the audit reports "no significant disparity" — technically true and completely misleading. The honest output is not a pass but "we cannot say": collect more data for that cell, or stop publishing a number for it.
Can the model reconstruct what you refused to collect?
Drop the protected attribute from the feature matrix, train a small classifier to predict it from what remains, and read off the AUC. Around 0.50 means the features carry nothing; 0.65 is meaningful leakage; above 0.85 the attribute is effectively still in the dataset, and "we don't use it" describes the schema, not the behaviour. The failing answer is the confident one — "we removed the sensitive columns" — because removal is a schema operation and prediction is what the model does.
Is the label measuring what you think it measures?
In a hiring model the label is almost never "was a good employee". It is "was hired by the previous process", or "was still employed after 12 months" — measurements of the old system's preferences and of who could afford to stay. Rebalancing a dataset whose label is the wrong measurement gives you a balanced wrong measurement. So: write down in one sentence what the label literally records, without using the word it is named after. If good_hire records "received an offer from a panel of four managers", say that. Half the time, writing the sentence changes the project.
Ask, in the same breath, for the missingness rate of every feature broken down by group. Differential missingness is a common hidden proxy: a field is blank precisely for the people the old process never processed, and a tree model happily learns that "blank" predicts the outcome. The bad answer is "we impute the median" without a per-group table — imputing over a group with 40% missingness quietly assigns that whole group the majority's typical value.
Development: commit to the metric before you see the numbers
Every fairness metric passes for some model. Here is one classifier on two groups — 4,000 people in Group A, 1,000 in Group B, base rates 40% and 25%.
| Group | TP | FN | FP | TN | n |
|---|---|---|---|---|---|
| A | 1,200 | 400 | 240 | 2,160 | 4,000 |
| B | 150 | 100 | 75 | 675 | 1,000 |
Every standard metric follows from those cells:
| Metric | Group A | Group B | Gap or ratio | Verdict at the usual threshold |
|---|---|---|---|---|
| Selection rate (1,440/4,000 vs 225/1,000) | 0.360 | 0.225 | ratio 0.63 | Fails the 0.80 four-fifths rule |
| True positive rate (1,200/1,600 vs 150/250) | 0.750 | 0.600 | 15.0 pts | Fails a 5-point bound |
| False positive rate (240/2,400 vs 75/750) | 0.100 | 0.100 | 0.0 pts | Passes, exactly |
| Precision (1,200/1,440 vs 150/225) | 0.833 | 0.667 | 16.7 pts | Fails predictive parity |
| Accuracy (3,360/4,000 vs 825/1,000) | 0.840 | 0.825 | 1.5 pts | Passes |
A team reporting "false positive rates are identical to three decimal places, and accuracy differs by 1.5 points" is telling the literal truth about a model that misses 40% of qualified Group B applicants against 25% of Group A's. No fabrication is required. The only defence is procedural: the metric is chosen and recorded before the numbers exist, with the reason, and the review reports all of them anyway.
Choose the metric before you see the numbers, and write down who chose it and why. Otherwise the metric you report is simply the one that passed.
Two more questions with teeth here. "Show me the explanation a rejected person receives" — not the tooling, the actual sentence. A force plot is an engineering artefact; "declined mainly because of a debt-to-income ratio of 0.61 and two missed payments in 24 months" is an explanation, and if the model cannot produce the second, the stakes may justify a simpler model. And "what does it do with a missing or out-of-range field?" Feed it a negative income and a 200-year employment history: systems that degrade gracefully refuse, systems that fail dangerously return a confident score built on an imputed value.
Testing: does the evidence come from the world or from the model?
One question dominates here and is rarely asked: was this test data influenced by the model you are testing? If your evaluation set is made of applicants the current system already screened, you are measuring how well the model reproduces itself. The fix is a small randomly selected control slice held outside the model's influence, and it is the only clean data you will ever have.
Alongside it: report groups by intersection, not one attribute at a time — a model can show a 2-point gap by sex and a 3-point gap by skin tone while the darker-skinned-women cell is off by 30 points. Sort cells by size ascending and read the small ones first. Check the metrics on realistic input — typos, truncated fields, wrong units — not the cleaned test set. And for models trained on personal records, check whether training data can be recovered; memorisation is a privacy failure that never appears in an accuracy metric.
Deployment readiness: accountability you can test
| Question | What it catches | What a bad answer sounds like |
|---|---|---|
| Name the individual accountable for this system's outcomes. | Diffusion — a committee that meets quarterly and owns nothing | "The AI governance board." |
| How is it turned off, who is authorised, how long does it take, when was that last exercised end to end? | Kill switches that exist as a config flag nobody has ever flipped | "We'd push a change to disable the endpoint." (Never tested. Needs a release. Four hours.) |
| During the pilot, how many people appealed, and how many appeals were upheld? | Appeal routes that technically exist and are structurally unusable | "There's a contact address on the rejection notice." |
Two more: if a person corrects a field about themselves, does the decision get re-run — correction rights that update a database but not the outcome are not rights. And what is the fallback while the system is off, since a system nothing else can replace cannot really be switched off. Note too that a zero appeal rate across 10,000 pilot decisions is not evidence everyone was satisfied. It is evidence nobody could find the route.
After launch: the questions nobody schedules
A pre-launch review catches only what was visible before launch. Much real harm appears afterwards, through drift — the deployed population diverging from the training population — and through feedback, where the model's decisions shape the data it is next trained on.
The feedback effect is worth the arithmetic. Of every 5,000 applicants above, 4,000 are Group A and 1,000 Group B, so Group B is 20% of the pool. Repayment outcomes exist only for approved applicants, so the next training set gets 1,440 labelled rows from Group A and 225 from Group B — Group B's share of the labelled data being
One retraining cycle has cut Group B's representation from 20% to 13.5% — a relative drop of a third — with nobody changing a line of code. Fewer rows means a worse fit, which means lower scores, fewer approvals, and fewer rows again. The system is manufacturing the evidence for its own conclusions.
Anything you measure only before launch, you have measured on the one distribution the system was never going to face.
Signal Cadence Alarm when What it catches Selection rate per group Daily Any pairwise ratio drops below 0.80 Drift; pipeline changes TPR and FPR per group Weekly, as outcomes arrive Gap widens by over 5 points versus the certified baseline Degradation invisible in aggregate accuracy Group mix of the next training set Every retrain Diverges from the applicant pool by over 10% relative The feedback loop, before it compounds Human override rate per group Monthly One group overridden far more often Field evidence the model is wrong about a group Re-run the proxy probe quarterly too; new data sources bring new proxies. But the override row is the most under-used signal in production — a human overturning the model disproportionately for one group is direct field evidence of a disparity your test set did not contain. Two questions round this out: when was the incident plan last rehearsed (a plan never run is a document, not a capability), and what changed in the applicable regulation, since a system certified against last year's rules is not certified.
Impact assessment: turning "harm" into rows
Impact assessment answers something narrower than "is this ethical?": if this system behaves as designed, and if it fails, who is harmed, how badly, and what evidence would settle it? Severity must be anchored to consequence — Critical for irreversible harm to health, liberty or livelihood; High for a material opportunity lost and hard to recover, such as credit or employment; Medium for harm an appeal can undo; Low for friction with no lasting effect. Without those anchors, severity becomes a measure of how embarrassing something would be for the organisation. Each row must then be falsifiable: "mitigation" names a mechanism, "evidence" names something you can go and look at.
Harm Severity Evidence that would settle it Mitigation Qualified applicants declined at different rates by group High Per-group TPR with confidence intervals, on the control slice Constrained retraining with a 5-point bound; monthly monitoring Decision cannot be explained to the person it affects High Twenty sampled rejection notices, read by someone outside the team Reason codes generated at scoring time, stored with the decision Sensitive attributes recoverable from retained features Medium Proxy probe AUC on the live feature set Drop or orthogonalise the top leaking features; document the trade Store the whole assessment as a record whose fields are pointers to evidence, and gate deployment on the presence of that evidence rather than on somebody's verdict:
Python1REQUIRED_EVIDENCE = {2 "problem": ["decision_sentence", "non_model_baseline_score"],3 "data": ["subgroup_counts_path", "proxy_probe_auc",4 "missingness_by_group_path"],5 "fairness": ["metric_committed_on", "metric_chosen_by",6 "metric_frame_path", "smallest_reported_cell_n"],7 "explainability": ["sample_rejection_notice_path"],8 "privacy": ["dpia_path", "lawful_basis", "retention_days"],9 "accountability": ["accountable_person", "kill_switch_last_tested_on",10 "fallback_process_doc"],11 "monitoring": ["alert_config_path", "control_slice_pct", "review_due_on"],12}1314class BlockedDeployment(Exception):15 pass1617def gate(record: dict) -> str:18 """Fail on ABSENT EVIDENCE, never on a self-reported verdict."""19 missing = {area: [k for k in keys if not record.get(area, {}).get(k)]20 for area, keys in REQUIRED_EVIDENCE.items()}21 missing = {area: keys for area, keys in missing.items() if keys}22 if missing:23 raise BlockedDeployment(missing)24 return "clear to deploy"What matters is what the gate checks. A review object whose methods return
{"status": "passed"}is the signature-line problem rendered in code: it records that someone asserted a verdict. This version cannot be satisfied by assertion, only by producing a file, a number, or a name. It also composes with the paperwork you owe anyway — a regulatory risk assessment and a data protection impact assessment want most of these same fields, and should be one document.When the checklist has no right answer
Some conflicts are not oversights. Two stakeholders want genuinely incompatible things, and no checklist item resolves that — only an explicit design conversation does. Value-sensitive design, developed by Batya Friedman and colleagues, is that conversation held deliberately.
It begins by separating direct stakeholders, who use the system, from indirect stakeholders, affected without using it. In lending the loan officer is direct and the applicant indirect — exactly backwards from where the harm falls, which is why indirect stakeholders get designed for last, if at all. Then you map the tensions, writing down the mechanism rather than the sentiment:
Tension The mechanism underneath it Resolved badly A defensible resolution Accuracy vs. demographic parity The loss-minimising classifier is not the parity-satisfying one when base rates differ Demand exact parity, then abandon it when accuracy drops A slack bound justified in advance, with per-group outcomes reported beside the ratio Privacy vs. measurable fairness You cannot measure a gap on an attribute you refuse to collect Collect nothing, then claim fairness cannot be assessed Collect it under a separate lawful basis, hold it segregated, use it only for aggregate audit — never as a feature Automation speed vs. contestability An automated decision has no natural point for a person to intervene Add a review queue nobody has time to work Automate reversible decisions; route irreversible ones to a human with a time budget The third step is iteration with real stakeholders, and here is the failure mode that gives the field its bad name. Ethics-washing is rarely fabricated results. It is a genuine consultation, honestly conducted, held after the design is locked: everyone is sincere, nothing they say can change anything, and the exercise produces legitimacy instead of information.
If no possible outcome of the consultation would have changed the design, it was not a review. Ask what would have to be said in that room to delay the launch — if there is no answer, you already know.
Running this so that it changes decisions
When a question does fail, do not settle it in a corridor. State the fact and the conflict separately — "the TPR gap by age is 15 points" is the fact, "equalising it costs 2 points of accuracy" is the conflict, and most arguments happen because people dispute different halves. Write down at least three options, one of which is not shipping; if "do not ship" is never on the list, this is not a decision procedure. Check what the law allows: group-specific thresholds are unlawful in several jurisdictions for employment and credit however well they perform. Then decide, record the reasoning, and set a review date — when the question returns in eighteen months with different people in the room, that record is the difference between a decision and a rumour.
Checklists themselves fail in recognisable ways. Two are obvious once named: answers arriving as paragraphs of prose mean the review records opinions, not artefacts; and a review scheduled after the launch date has already been decided. The rest:
Symptom What it means What to do A 100% pass rate across two years of reviews The questions have no failing answers For each item, write the bad answer; delete any item where you cannot The reviewer built the thing Nobody is positioned to say no Rotate reviewers across teams; the reviewer's job is to ask for the artefact It has never stopped anything It is a signature-collection process The single best indicator; treat it as urgent Put the questions where the work happens, not in a policy document. Data questions belong in the dataset card that must exist before a training run; development questions in the pull request template; deployment questions in the release gate, as code that refuses to proceed when evidence is missing. A question someone has to remember to ask gets asked when there is time, which is never. Cap each gate at five to nine items and retire anything an automated test now covers — a list grown to 120 items, one per past incident, gets completed in ten minutes with a ruler.
Then accept the cost. The first time one of these questions delays a launch by three weeks, someone senior will ask whether the process has become too heavy. The answer is that a control which never fires is not a control, and three weeks is the whole price of finding out before a regulator does.