Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What steps would you take after an AI system makes a harmful decision?
What you need to know
- Contain — flip the feature flag, roll back to the previous model/prompt/index bundle, or route to a human path. Every minute of diagnosis before containment is more harm.
- Preserve evidence — freeze traces, inputs, retrieved documents, guardrail decisions, model and prompt versions. Suspend retention jobs that would delete them.
- Scope — replay logs to find every similar case. The reported case is rarely the only one.
- Remediate for people — correct each affected decision with a human reviewer, contact the people affected, reverse the consequence, open an appeal path.
- Escalate — legal, privacy, communications, and the accountable owner. Personal-data breaches have reporting duties (GDPR's 72 hours; DPDP reporting to the Data Protection Board and affected people; CERT-In's short deadline for certain cyber incidents in India).
- Fix and verify — root cause, then a permanent regression test, then a canary release.
- Learn — blameless post-mortem with owners, dates and at least one detection improvement.
What an interviewer listens for
- Contain first. Candidates who start with "I'd debug the prompt" lose points.
- The people affected. Fixing the system does not fix the decisions already made.
- A test, not a promise. "We told the model not to do that again" is not a fix. The failing case becomes a test that blocks any release where it fails again.
A real-life example
An HR screening assistant at a large retailer rejects applicants automatically when its score is below 40. A candidate complains that she was rejected within a minute for a store-manager role she had done for six years. The HR tech lead:
- Turns off automatic rejection within the hour; all candidates below 40 now go to a recruiter queue.
- Freezes logs and finds the cause the next day: CVs uploaded as scanned PDFs lost their text, so the model scored an almost empty document.
- Replays three weeks of logs: 212 applicants with scanned CVs were auto-rejected.
- Has recruiters review all 212 by hand; 37 are invited to interview, with an apology.
- Informs legal, because automated rejection can bring anti-discrimination and data-protection duties.
- Adds a check that blocks scoring when extracted text is under 200 words, and adds 20 scanned CVs to the eval set.
Follow-up questions to expect
- "How do you decide whether to shut the whole system down?" — By harm and scope: if you cannot bound which decisions are affected, turn off the automated path and fall back to humans until you can.
- "Who do you tell first?" — The incident commander and the accountable owner, then legal and privacy in parallel; they decide on external notification.
- "How do you know you found every affected case?" — Define the failure as a query over logs (for example, "extracted text under 200 words and auto-rejected"), and have a second person check the query.