AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

What steps would you take after an AI system makes a harmful decision?


What you need to know

  1. Contain — flip the feature flag, roll back to the previous model/prompt/index bundle, or route to a human path. Every minute of diagnosis before containment is more harm.
  2. Preserve evidence — freeze traces, inputs, retrieved documents, guardrail decisions, model and prompt versions. Suspend retention jobs that would delete them.
  3. Scope — replay logs to find every similar case. The reported case is rarely the only one.
  4. Remediate for people — correct each affected decision with a human reviewer, contact the people affected, reverse the consequence, open an appeal path.
  5. Escalate — legal, privacy, communications, and the accountable owner. Personal-data breaches have reporting duties (GDPR's 72 hours; DPDP reporting to the Data Protection Board and affected people; CERT-In's short deadline for certain cyber incidents in India).
  6. Fix and verify — root cause, then a permanent regression test, then a canary release.
  7. Learn — blameless post-mortem with owners, dates and at least one detection improvement.

What an interviewer listens for

  • Contain first. Candidates who start with "I'd debug the prompt" lose points.
  • The people affected. Fixing the system does not fix the decisions already made.
  • A test, not a promise. "We told the model not to do that again" is not a fix. The failing case becomes a test that blocks any release where it fails again.

A real-life example

An HR screening assistant at a large retailer rejects applicants automatically when its score is below 40. A candidate complains that she was rejected within a minute for a store-manager role she had done for six years. The HR tech lead:

  • Turns off automatic rejection within the hour; all candidates below 40 now go to a recruiter queue.
  • Freezes logs and finds the cause the next day: CVs uploaded as scanned PDFs lost their text, so the model scored an almost empty document.
  • Replays three weeks of logs: 212 applicants with scanned CVs were auto-rejected.
  • Has recruiters review all 212 by hand; 37 are invited to interview, with an apology.
  • Informs legal, because automated rejection can bring anti-discrimination and data-protection duties.
  • Adds a check that blocks scoring when extracted text is under 200 words, and adds 20 scanned CVs to the eval set.

Follow-up questions to expect

  • "How do you decide whether to shut the whole system down?" — By harm and scope: if you cannot bound which decisions are affected, turn off the automated path and fall back to humans until you can.
  • "Who do you tell first?" — The incident commander and the accountable owner, then legal and privacy in parallel; they decide on external notification.
  • "How do you know you found every affected case?" — Define the failure as a query over logs (for example, "extracted text under 200 words and auto-rejected"), and have a second person check the query.