AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

What is the “default to safe” principle in guardrail design?


The moderation service times out mid-saleFail open: product questions• Answers keep flowing• Every pass is logged for review• Chosen because blocking harms more• Written down in the fail-mode tableFail closed: store credit• Action shows 'try again shortly'• No unchecked money movement• Any unlisted route defaults here• A CI test catches missing routes
The gift-message route was never listed and failed open, which is why unknown must mean closed.

What you need to know

Where "safe by default" shows up

  • Allowlists, not denylists: list what is permitted; everything else is denied. Denylists are always incomplete.
  • Unknown means no: an unrecognised tool, argument or intent is rejected, not guessed.
  • Fail closed on high-risk paths: payments, external email, account changes.
  • Weak evidence means abstain: low retrieval confidence gives "I don't know".
  • Ambiguity escalates: "delete my old records" asks which records.
  • Safe configuration: a missing environment variable turns guardrails on, not off.

Fail mode per path

Python
FAIL_MODE = {"payments": "closed", "external_email": "closed", "faq": "open"}def moderate(route, text, classifier):    try:        return classifier(text)                    # True means allowed    except TimeoutError:        mode = FAIL_MODE.get(route, "closed")      # unlisted route -> closed        return mode == "open"def slow(_):    raise TimeoutErrorprint(moderate("faq", "opening hours?", slow))     # Trueprint(moderate("payments", "pay Rs 900", slow))    # Falseprint(moderate("new_route", "anything", slow))     # False

The fail mode is an explicit decision per route, written down and reviewed. A route nobody listed gets the safe default. When the FAQ route fails open, you still log it so you know how often it happened.

The cost of over-refusal

A system that refuses too much fails users too. People stop using it, copy data into unapproved tools, or phrase requests to get around it. So track:

  • Block rate on real traffic.
  • False-block rate on a labelled benign set.
  • User signals: rephrasing after a refusal, abandonment, complaints.

Tune thresholds per risk tier. "Fail closed everywhere" is not a safety win if it makes the product unusable.

A real-life example

During a big sale, an e-commerce company's moderation provider has a 20-minute outage. The shopping assistant's product-question route is set to fail open, so customers keep getting answers (logged for later review). The "apply store credit" route fails closed and shows "This action is temporarily unavailable — try again shortly".

A post-incident review finds one gap: a new "gift message" feature launched the week before was never added to the fail-mode table and had been coded to fail open. For 20 minutes, unmoderated gift messages were printed on packing slips. Two were abusive. The fix makes the table the only source of fail modes, with a CI test that fails if any route is missing from it.

Follow-up questions to expect

  • "When is failing open correct?" — When blocking causes more harm than the risk, for example an emergency-information page, or low-risk reads where an outage would block thousands of legitimate users.
  • "How do you avoid over-refusal?" — Measure false blocks on real benign examples, tune thresholds per category, and offer a safe alternative ("I can't do X, but I can do Y").
  • "How does this apply to agents?" — If the agent's plan is unclear or a tool returns an unexpected result, stop and ask, rather than trying something else.