Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do you design an incident response plan for AI systems?
What you need to know
Severity by harm, not only availability
| Severity | Example |
|---|---|
| Sev-1 | Dangerous medical advice; exposure of personal data; wrong financial actions at scale |
| Sev-2 | Biased decisions on one group; widespread wrong answers on a key topic |
| Sev-3 | Tone problems, low-impact errors, isolated reports |
Detection
Alert on signals that move before users complain:
- Guardrail block rate and refusal rate (sudden rise or fall).
- Groundedness score and citation-support rate.
- Output length or toxicity distribution shifts.
- Tool error rate and unusual tool-call patterns.
- Cost and latency anomalies.
- Volume of user reports, with an external channel for security researchers.
Roles
On-call rotation, an incident commander, and pre-agreed contacts in legal, privacy (the DPO where you have one), communications and the system's accountable owner.
Containment tools (built in advance)
- A kill switch per capability ("turn off auto-send", "turn off refunds").
- Rollback to a pinned model/prompt/index/guardrail bundle.
- A degraded mode: answer from FAQ only, or route to humans.
All must be usable by on-call without a code deploy.
Notification templates and legal clocks
Prepare templates for internal updates, affected users and regulators. Write down the deadlines that may apply, and have legal confirm them:
- GDPR: personal-data breaches to the supervisory authority within 72 hours where required.
- India: DPDP breach reporting to the Data Protection Board and affected people; CERT-In's directions require certain cyber incidents to be reported within 6 hours.
- EU AI Act: providers of high-risk systems must report serious incidents to market surveillance authorities.
Evidence, recovery and review
Retention of traces and audit logs long enough to investigate, with a legal-hold procedure; a queue to re-process affected outputs; user remediation; a blameless post-mortem.
A real-life example
A healthcare symptom-checker's team runs a game day. At 11:00 the facilitator silently swaps in a prompt that under-triages chest pain. Targets: detect within 30 minutes, contain within 15 more.
Results: the "emergency escalation rate" alert fires at 11:22 because escalations drop from 4.1% to 1.3%. The on-call engineer finds the kill switch for "AI triage" — but it needs a deploy, taking 25 minutes. Legal's phone number in the runbook is outdated. Actions: the kill switch becomes a runtime flag; contacts are checked monthly by an automated reminder; a second game day a month later contains the same scenario in 6 minutes.
Follow-up questions to expect
- "What is a good first alert for an LLM system?" — Sudden changes in refusal or guardrail-block rate. They move in both directions when a prompt, model or input distribution changes.
- "How often should you rehearse?" — At least quarterly for high-risk systems, and after major architecture changes.
- "Who can press the kill switch?" — Any on-call engineer, without approval. Waiting for permission is how small incidents become big ones.