Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do you run a post-mortem after an AI failure?
What you need to know
The standard sections
- Timeline: when it started, was detected, was mitigated, was resolved. Detection lag is usually the most important number.
- Impact: requests, users, money, and the kind of harm.
- What happened: technical sequence.
- Contributing factors: always more than one.
- What went well: so you do not remove a control that helped.
AI-specific questions
| Question | Why it matters |
|---|---|
| Which layer failed — model, prompt, retrieval, data, guardrail? | Points the fix at the right place |
| Our change, or drift (vendor update, input shift, stale index)? | Drift needs monitoring, not code review |
| Did the eval set contain anything like this case? | Finds gaps in how eval sets are built |
| Which guardrail should have caught it, and why didn't it? | Tests the defence-in-depth story |
| Was a human in the path, and would they have caught it? | Tests oversight design |
Outputs
- The failing case, and variations of it, added to the regression suite.
- Action items with an owner and a date each.
- At least one detection improvement — an alert or check that would have caught it sooner.
- A short summary shared widely so other teams learn.
A real-life example
An email assistant's "smart reply" feature sent a customer a reply that included a discount code meant only for internal staff. Timeline: the code was added to a shared knowledge base on Monday; first leak Tuesday 10:00; noticed Thursday when finance saw 600 redemptions; feature disabled Thursday 16:00. Detection lag: 54 hours.
Findings: the model did exactly what it was asked — use the knowledge base. The failure was that the knowledge base mixed internal and customer-facing documents with no label, and no output check looked for internal-only content. The eval set had 400 reply cases but none involving internal documents. Actions: split the knowledge base by audience (owner: platform lead, 2 weeks); an output check for strings tagged internal (owner: safety engineer, 1 week); an alert on unusual discount redemptions (owner: finance analytics, 1 week); 30 new eval cases with planted internal documents.
Follow-up questions to expect
- "How is an AI post-mortem different from a normal one?" — The same process, plus questions about data, evals and drift, because AI failures often have no code change behind them.
- "How do you avoid blame when a person made the change?" — Ask what allowed a single change to cause harm without being caught — missing review, missing test, missing alert.
- "Who should attend?" — The engineers involved, the system owner, whoever handled users, and someone from a neighbouring team for fresh eyes.