AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do you run a post-mortem after an AI failure?


What you need to know

The standard sections

  • Timeline: when it started, was detected, was mitigated, was resolved. Detection lag is usually the most important number.
  • Impact: requests, users, money, and the kind of harm.
  • What happened: technical sequence.
  • Contributing factors: always more than one.
  • What went well: so you do not remove a control that helped.

AI-specific questions

QuestionWhy it matters
Which layer failed — model, prompt, retrieval, data, guardrail?Points the fix at the right place
Our change, or drift (vendor update, input shift, stale index)?Drift needs monitoring, not code review
Did the eval set contain anything like this case?Finds gaps in how eval sets are built
Which guardrail should have caught it, and why didn't it?Tests the defence-in-depth story
Was a human in the path, and would they have caught it?Tests oversight design

Outputs

  • The failing case, and variations of it, added to the regression suite.
  • Action items with an owner and a date each.
  • At least one detection improvement — an alert or check that would have caught it sooner.
  • A short summary shared widely so other teams learn.

A real-life example

An email assistant's "smart reply" feature sent a customer a reply that included a discount code meant only for internal staff. Timeline: the code was added to a shared knowledge base on Monday; first leak Tuesday 10:00; noticed Thursday when finance saw 600 redemptions; feature disabled Thursday 16:00. Detection lag: 54 hours.

Findings: the model did exactly what it was asked — use the knowledge base. The failure was that the knowledge base mixed internal and customer-facing documents with no label, and no output check looked for internal-only content. The eval set had 400 reply cases but none involving internal documents. Actions: split the knowledge base by audience (owner: platform lead, 2 weeks); an output check for strings tagged internal (owner: safety engineer, 1 week); an alert on unusual discount redemptions (owner: finance analytics, 1 week); 30 new eval cases with planted internal documents.

Follow-up questions to expect

  • "How is an AI post-mortem different from a normal one?" — The same process, plus questions about data, evals and drift, because AI failures often have no code change behind them.
  • "How do you avoid blame when a person made the change?" — Ask what allowed a single change to cause harm without being caught — missing review, missing test, missing alert.
  • "Who should attend?" — The engineers involved, the system owner, whoever handled users, and someone from a neighbouring team for fresh eyes.