AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

What ethical considerations should guide deployment in high-stakes domains?


What you need to know

The core questions

  • Fitness for the actual population: worst-slice performance, not average.
  • Distribution of costs and benefits: efficiency for the institution, risk for the applicant or patient, is an ethical problem that no accuracy metric shows.
  • Meaningful consent: consent that cannot be refused without losing the service is not meaningful.
  • Real oversight: reviewers with time, training, information and authority.
  • Contestability: people with the least time and power must still be able to challenge a decision.
  • Dignity and autonomy: people should know when AI is involved and not be manipulated.
  • Proportionality: is AI the right tool here at all, or is a simpler rule clearer and fairer?

Turning questions into practice

  1. Harm model — name who could be hurt, how, and how badly.
  2. Worst-slice evaluation — with thresholds that block launch.
  3. Conservative defaults — abstain or escalate when uncertain.
  4. Independent review — domain experts, ethics or risk review, and where possible people from affected communities.
  5. Staged rollout — pilot, then expand, with stop criteria agreed before launch.
  6. Outcome monitoring — real-world results, not only model metrics.
  7. Decommissioning plan — what would make you switch it off, and how.

Saying no

A senior answer includes the option of not deploying, or deploying only as an assistant where a human decides. State the trade-off you accepted, how you would know you were wrong, and what result would make you turn it off.

A real-life example

A state health department wants an AI tool to prioritise which rural patients get home visits from a limited team of community health workers. The model, trained on hospital records, ranks patients by predicted risk.

The ethics review asks who is missing from hospital records: people who never reached a hospital — often the poorest and most remote. The model would rank them low because it has little data on them, sending visits toward people already in the system. The team changes the plan: the tool suggests priorities, but health workers add patients from their own knowledge; 20% of visits are reserved for households with no hospital records; results are reviewed monthly by district, with a stop rule if visits to the most remote blocks fall. The tool launches as assistive only, with automated prioritisation deferred until a year of outcome data exists.

Follow-up questions to expect

  • "How do you weigh efficiency against fairness?" — Make the trade-off explicit with numbers for each group, decide it with the accountable owner and affected stakeholders, and write it down.
  • "Who should be on an ethics review?" — Domain experts, legal, risk, engineers who understand the system, and people who represent those affected.
  • "What is a stop criterion?" — A pre-agreed measurable condition, like "if the worst group's error rate exceeds 15% in any month, pause automated use".