Course Content
Introduction to AI
4 sections · 10 lessons
Ethics, Bias, and Responsible AI
Most writing about AI ethics is either abstract enough to be useless or alarmed enough to be unhelpful. This lesson takes a different approach: concrete failures that actually happened, why they happened mechanically, and what you can do about them as someone who builds or buys these systems.
The starting point matters. Bias in AI is not usually the result of anyone behaving badly. It is generally the result of ordinary people building a reasonable-seeming system on data that recorded an unfair world. That makes it a design problem, which is good news — design problems can be addressed.
Where bias actually enters
There are four distinct entry points, and confusing them leads to fixes that do not work.
1. In the data, from history
The most common source by far.
A company trains a hiring model on ten years of its own decisions. Those decisions reflected whoever was hired and promoted historically. If the industry skewed heavily one way, the data records that pattern — and the model learns it as the definition of a good candidate.
The model is working perfectly. It found the pattern in the data. The pattern was unfair, and the model has now made it fast, cheap, consistent, and much harder to argue with.
This is the crucial thing to understand: the model does not create the bias. It concentrates it, scales it, and gives it the appearance of objectivity. A biased human decision can be challenged. "The algorithm scored you 34" feels like a fact.
2. In the data, from who is represented
Different mechanism, same effect. Some groups are underrepresented in your data, so the model has seen fewer examples and performs worse for them.
The 2018 Gender Shades study tested commercial facial analysis systems that guessed a person's gender from a photo. Error rates were under 1% for lighter-skinned men and above 30% for darker-skinned women — the same systems, on the same task. The system was not designed to work badly for some people. It simply had far less to learn from.
This is why headline accuracy is dangerous. A system that is "97% accurate" overall can be near-useless for a substantial minority of the people it will be applied to, and the average hides it completely.
3. In what you chose to measure
The subtlest source, and worth real attention.
A widely used US healthcare algorithm, studied by Obermeyer and colleagues in Science in 2019, was designed to identify patients who would benefit from extra care. It could not measure "medical need" directly, so it used healthcare spending as a stand-in. Reasonable on the face of it — sicker patients cost more.
But spending also reflects access. Patients who face barriers to care spend less while being equally or more ill. The model learned to predict spending accurately, and in doing so systematically under-identified the patients with the greatest need.
Nobody was careless. The proxy seemed sensible. The gap between "what we can measure" and "what we actually care about" is where a great deal of harm lives.
4. In how the output is used
A model can be reasonable and the system around it still cause harm.
Consider a risk score of 0.7. One product uses it to route a case to a human for careful review. Another uses it to deny an application automatically. Same model, same number, entirely different consequence for a real person.
Design decisions — thresholds, whether a human reviews, whether there is an appeal, whether the person is even told a model was involved — often matter more than model accuracy.
Why "just remove the sensitive attribute" does not work
This is the first thing everyone suggests and it is worth understanding precisely why it fails.
Remove race from the inputs. The model still receives postcode, education, employment history, shopping patterns, and dozens of other variables. In a society with a history of segregation, those variables correlate with race. The model reconstructs a usable approximation without ever being told.
These are proxy variables, and they are unavoidable in rich datasets.
Removing a protected attribute does not remove its influence. It removes your ability to measure that influence. You have made the system less fair to audit while leaving it just as biased.
The better approach is the opposite of the instinct: keep the sensitive attribute available for measurement, even where you exclude it from the model's inputs, so you can check whether outcomes differ across groups. You cannot fix what you have made yourself unable to see.
Fairness has no single definition
Here is a genuinely uncomfortable fact that is often left out of introductory material.
There are several reasonable mathematical definitions of fairness, and it has been proven that you generally cannot satisfy them all at once. This is not a gap in current research. It is a mathematical impossibility except in special cases.
| Definition | What it requires |
|---|---|
| Demographic parity | Each group receives positive outcomes at the same rate |
| Equal opportunity | Among those who genuinely qualify, each group is approved at the same rate |
| Predictive parity | A given score means the same thing regardless of group |
Each sounds obviously correct in isolation. When the underlying base rates differ between groups — which they often do, frequently because of past injustice — satisfying one forces you to violate another.
This matters practically because it means fairness is not a box an engineer can tick. Somebody has to decide which definition applies to this situation, and that is a value judgement requiring the people affected, domain experts, and often lawyers. An engineer choosing quietly and alone is making a policy decision without the authority to make it.
The explainability problem
Modern models are frequently unable to tell you why they produced a particular output. This collides with law and with basic decency.
Regulations in several jurisdictions give people a right to an explanation for significant automated decisions. "The network produced 0.34" is not an explanation.
Partial approaches exist. Some models are inherently readable — you can look at a decision tree and follow the path. Techniques exist to estimate which inputs pushed a particular decision, which is genuinely useful though approximate. And sometimes the honest answer is to accept a slightly less accurate but explainable model because the situation demands explanation.
The trade is real and should be made deliberately: a 92% accurate model you can explain may be worth more than a 94% model you cannot — in lending, hiring, healthcare, or criminal justice. Accuracy is not the only thing with value.
Privacy
Three concerns worth knowing by name.
Training data exposure. Large models can memorise fragments of their training data and reproduce them. If personal information was in the training set, it can sometimes be extracted.
Inference of what you did not disclose. Models can deduce sensitive attributes you never provided — health conditions, pregnancy, sexuality — from ordinary behavioural data. Consenting to share your shopping history is not consenting to have your medical status inferred, but the second follows from the first.
Re-identification. "Anonymised" datasets are often not. A handful of ordinary data points — postcode, date of birth, sex — is frequently enough to identify a specific person. Removing names is not anonymisation.
Accountability: who answers for it
When an automated decision harms someone, responsibility has a habit of evaporating. The data team points at the model team, who point at the product team, who point at the vendor.
The principle that resolves this: responsibility does not transfer to software. A named person or team must own each deployed system — its behaviour, its failures, and the decision to keep running it.
Practically, that means being able to answer:
- Who decided this system should be deployed, and on what evidence?
- Who monitors whether it is still performing acceptably?
- How does an affected person contest a decision, and who reviews the appeal?
- What triggers turning it off, and who has the authority to do so?
Teams that cannot answer these do not have a governance gap on paper — they have a system that will eventually harm someone with nobody positioned to notice or stop it.
Practical questions before deploying
Not a compliance checklist. These are the questions that catch real problems.
- Who is affected by this decision, and were any of them consulted? The people who bear the consequences usually see failure modes the builders do not.
- How does performance differ across groups? Break the numbers down. Never accept the average.
- What are we measuring, and is it what we actually care about? Interrogate every proxy. This is where the healthcare failure came from.
- What is the worst realistic outcome of a wrong answer? If it is irreversible, a human must be in the loop.
- Can we explain a decision to the person affected? If not, is that acceptable here — legally and ethically?
- How does someone appeal? Every consequential automated decision needs a route to a human.
- What would tell us this has started going wrong? Decide the signal and the threshold before launch.
- Should we build this at all? Sometimes the answer is no, and that has to remain a permitted conclusion.
The environmental cost
Briefly, because it is real and frequently omitted. Training large models consumes substantial energy and water for cooling. The cost of using a model at scale often exceeds the cost of training it, since inference runs continuously for years.
The practical point for a builder is not guilt but proportion: use the smallest model that does the job. Reaching for the largest available system to solve something a small purpose-trained model handles is wasteful in money and energy simultaneously — the incentives happen to align here.
What responsible practice actually looks like
Not a department. A set of habits distributed through how a team works.
- Document what you built. What data, what it does and does not cover, what performance looks like broken down by group, known limitations. Write it while you build, not afterwards.
- Test on subgroups by default. Make disaggregated evaluation the standard report, not a special investigation someone has to request.
- Involve affected people early. Not as a review at the end, when changing anything is expensive.
- Keep a human in the loop for consequential decisions. Automation should assist judgement in high-stakes settings, not replace it.
- Monitor for fairness alongside accuracy. Drift can affect groups unevenly; overall accuracy can hold steady while one group's outcomes deteriorate.
- Make disagreement safe. Most of these failures were foreseeable, and in several cases somebody did foresee them. The problem was not knowledge. It was that raising the concern was harder than staying quiet.
That last point is the one that generalises beyond AI, and it is probably the most important thing in this lesson. Every failure described here was preventable by someone asking an uncomfortable question early enough for the answer to matter.