Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What is the “default to safe” principle in guardrail design?
What you need to know
Where "safe by default" shows up
- Allowlists, not denylists: list what is permitted; everything else is denied. Denylists are always incomplete.
- Unknown means no: an unrecognised tool, argument or intent is rejected, not guessed.
- Fail closed on high-risk paths: payments, external email, account changes.
- Weak evidence means abstain: low retrieval confidence gives "I don't know".
- Ambiguity escalates: "delete my old records" asks which records.
- Safe configuration: a missing environment variable turns guardrails on, not off.
Fail mode per path
1FAIL_MODE = {"payments": "closed", "external_email": "closed", "faq": "open"}23def moderate(route, text, classifier):4 try:5 return classifier(text) # True means allowed6 except TimeoutError:7 mode = FAIL_MODE.get(route, "closed") # unlisted route -> closed8 return mode == "open"910def slow(_):11 raise TimeoutError1213print(moderate("faq", "opening hours?", slow)) # True14print(moderate("payments", "pay Rs 900", slow)) # False15print(moderate("new_route", "anything", slow)) # FalseThe fail mode is an explicit decision per route, written down and reviewed. A route nobody listed gets the safe default. When the FAQ route fails open, you still log it so you know how often it happened.
The cost of over-refusal
A system that refuses too much fails users too. People stop using it, copy data into unapproved tools, or phrase requests to get around it. So track:
- Block rate on real traffic.
- False-block rate on a labelled benign set.
- User signals: rephrasing after a refusal, abandonment, complaints.
Tune thresholds per risk tier. "Fail closed everywhere" is not a safety win if it makes the product unusable.
A real-life example
During a big sale, an e-commerce company's moderation provider has a 20-minute outage. The shopping assistant's product-question route is set to fail open, so customers keep getting answers (logged for later review). The "apply store credit" route fails closed and shows "This action is temporarily unavailable — try again shortly".
A post-incident review finds one gap: a new "gift message" feature launched the week before was never added to the fail-mode table and had been coded to fail open. For 20 minutes, unmoderated gift messages were printed on packing slips. Two were abusive. The fix makes the table the only source of fail modes, with a CI test that fails if any route is missing from it.
Follow-up questions to expect
- "When is failing open correct?" — When blocking causes more harm than the risk, for example an emergency-information page, or low-risk reads where an outage would block thousands of legitimate users.
- "How do you avoid over-refusal?" — Measure false blocks on real benign examples, tune thresholds per category, and offer a safe alternative ("I can't do X, but I can do Y").
- "How does this apply to agents?" — If the agent's plan is unclear or a tool returns an unexpected result, stop and ask, rather than trying something else.