AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

What is the “defense-in-depth” approach in AI safety?


800 red-team prompts against a symptom-checker61%14 more22%not tested12%97 more9 of 9all 9Stopped here firstHarm if removedInput classifierPrompt refusalsOutput safety checkData-access checkAblation: switch one layer off in staging and count what now reaches users.
The layer that catches the most first is not always the one you can least afford to lose.

What you need to know

The layers in an LLM application

  1. Limit attempts — identity, rate limits, quotas.
  2. Check inputs — validation, moderation, injection screening.
  3. Separate trust — untrusted documents and tool results are delimited and never treated as instructions.
  4. Choose a safe model — safety-tuned, pinned version.
  5. Scope data — retrieval filtered to what this user may see.
  6. Scope actions — tool allowlists, argument checks, scoped credentials, human approval.
  7. Check outputs — groundedness, PII and secret scanning, safety classification.
  8. Watch — monitoring, audit logs, anomaly alerts, kill switch.
  9. Organise — red-teaming, incident response, rollback.

Independence is the whole point

Layers multiply only when failures are uncorrelated:

  • Correlated: an injection classifier and an LLM judge built on the same model family — the same clever text fools both.
  • Uncorrelated: a classifier (statistical) plus a permission check (deterministic) plus a human approval (a person). Each fails for a different reason.

How to verify it

  • Per-layer catch rate: run a red-team corpus and record which layer stopped each attack.
  • Ablation: switch off one layer in staging and re-run. If nothing changes, the layer may not be earning its latency. If everything gets through, you have found a single point of failure.
  • Correlation check: if two layers always catch the same cases, count them as one.

A real-life example

A healthcare symptom-checker is tested with 800 red-team prompts: requests for prescription-only drug doses, self-harm content, injected instructions in uploaded lab reports, and attempts to get other patients' data.

The per-layer report shows the input classifier stops 61%, the system prompt's refusals stop another 22%, the output medical-safety check stops 12%, and the data-access check stops all 9 cross-patient attempts on its own. Ablation shows that without the output check, 97 harmful responses reach users; without the input classifier, only 14 more get through, because later layers catch most of them. The team keeps both but moves the output check to the top of the latency budget, because it is the layer that actually protects patients.

Follow-up questions to expect

  • "Doesn't every layer add latency?" — Yes. Run independent checks in parallel with generation where possible, put cheap deterministic ones first, and remove layers that ablation shows add nothing.
  • "What is the most important layer?" — For agents, the tool and permission layer, because it bounds the worst case whatever the model does.
  • "How is this different from layered guardrails?" — Same idea; defence in depth adds the requirement that the layers are independent and that you prove each one works.