AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do you design an incident response plan for AI systems?


What you need to know

Severity by harm, not only availability

SeverityExample
Sev-1Dangerous medical advice; exposure of personal data; wrong financial actions at scale
Sev-2Biased decisions on one group; widespread wrong answers on a key topic
Sev-3Tone problems, low-impact errors, isolated reports

Detection

Alert on signals that move before users complain:

  • Guardrail block rate and refusal rate (sudden rise or fall).
  • Groundedness score and citation-support rate.
  • Output length or toxicity distribution shifts.
  • Tool error rate and unusual tool-call patterns.
  • Cost and latency anomalies.
  • Volume of user reports, with an external channel for security researchers.

Roles

On-call rotation, an incident commander, and pre-agreed contacts in legal, privacy (the DPO where you have one), communications and the system's accountable owner.

Containment tools (built in advance)

  • A kill switch per capability ("turn off auto-send", "turn off refunds").
  • Rollback to a pinned model/prompt/index/guardrail bundle.
  • A degraded mode: answer from FAQ only, or route to humans.

All must be usable by on-call without a code deploy.

Notification templates and legal clocks

Prepare templates for internal updates, affected users and regulators. Write down the deadlines that may apply, and have legal confirm them:

  • GDPR: personal-data breaches to the supervisory authority within 72 hours where required.
  • India: DPDP breach reporting to the Data Protection Board and affected people; CERT-In's directions require certain cyber incidents to be reported within 6 hours.
  • EU AI Act: providers of high-risk systems must report serious incidents to market surveillance authorities.

Evidence, recovery and review

Retention of traces and audit logs long enough to investigate, with a legal-hold procedure; a queue to re-process affected outputs; user remediation; a blameless post-mortem.

A real-life example

A healthcare symptom-checker's team runs a game day. At 11:00 the facilitator silently swaps in a prompt that under-triages chest pain. Targets: detect within 30 minutes, contain within 15 more.

Results: the "emergency escalation rate" alert fires at 11:22 because escalations drop from 4.1% to 1.3%. The on-call engineer finds the kill switch for "AI triage" — but it needs a deploy, taking 25 minutes. Legal's phone number in the runbook is outdated. Actions: the kill switch becomes a runtime flag; contacts are checked monthly by an automated reminder; a second game day a month later contains the same scenario in 6 minutes.

Follow-up questions to expect

  • "What is a good first alert for an LLM system?" — Sudden changes in refusal or guardrail-block rate. They move in both directions when a prompt, model or input distribution changes.
  • "How often should you rehearse?" — At least quarterly for high-risk systems, and after major architecture changes.
  • "Who can press the kill switch?" — Any on-call engineer, without approval. Waiting for permission is how small incidents become big ones.