AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do input and output guardrails protect AI applications?


Two checkpoints the model cannot talk its way pastUser messageInput rails: limits,moderation, PII redactionModel callOutput rails:safety, PII,groundedness, schemaUser, or avalidatedtool callIn week one the output PII scan fired 14 times, all from one retrieval bug.
The system prompt asks; the rails on either side of the model enforce, and only enforcement survives a determined user.

What you need to know

Input guardrails (before the model)

  • Identity, rate limits and budgets: who is asking, and how much can they ask.
  • Validation: length caps, allowed file types, JSON schema for structured inputs.
  • Moderation: classify the request into policy categories (self-harm, hate, violence, illegal activity).
  • Injection and jailbreak screening: on the user's text and on retrieved documents or tool results.
  • PII redaction: remove card numbers, Aadhaar numbers and similar before the text leaves your boundary.
  • Topic scoping: decline off-domain requests cheaply ("I can only help with your account").

Output guardrails (after the model)

  • Safety classification of the reply.
  • PII and secret scanning: the model can echo data from context that this user should not see.
  • Groundedness: is each claim supported by the retrieved context?
  • Format: schema validation for structured output.
  • Tool-call validation: is this tool allowed, are the arguments in range, does it need human approval?

Common tools, described accurately

ToolWhat it is
Provider moderation endpointsOpenAI's moderation endpoint (omni-moderation-latest) returns a flag and per-category scores for text and images. Azure AI Content Safety returns severity levels for hate, sexual, violence and self-harm, and its Prompt Shields detect injection in user prompts and documents. AWS Bedrock Guardrails and Google's Gemini safety settings offer configurable content filters.
Llama Guard (Meta)An open-weight LLM fine-tuned as a classifier. It reads a user prompt or model response and outputs safe or unsafe plus the violated hazard codes (for example S1 for violent crimes). Llama Guard 3 and 4 follow the MLCommons hazard taxonomy; Meta's separate Prompt Guard models detect injection and jailbreak attempts.
NeMo Guardrails (NVIDIA)An open-source toolkit where you configure "rails" in YAML and the Colang language: input, output, dialog, retrieval and execution rails. It can call other checkers, such as Llama Guard or a self-check prompt.
Guardrails AIAn open-source Python library. A Guard wraps an LLM call and runs validators from its Hub (PII, toxicity, regex, schema), with an on_fail policy such as raise an exception, fix, re-ask or filter.

Design rules

  • Cheap first: regex and schema checks cost microseconds; a classifier costs tens to hundreds of milliseconds. Run the cheap ones first.
  • Fail closed on high-risk paths: if the moderation call times out on the payments route, block. Fail open only where blocking harms more, and write that decision down.
  • Streaming: either buffer to a sentence or paragraph boundary before checking, or check the stream and be ready to retract.

A real-life example

A bank's customer chatbot handles 40,000 chats a day. Input: a regex plus Luhn check removes card numbers before the text reaches the vendor model; Llama Guard screens for abuse; a small classifier rejects questions unrelated to banking. Output: a PII scanner blocks replies that contain another customer's account number, and a schema check ensures the "transfer" intent returns only a draft for the user to confirm.

In the first week, output PII scanning fires 14 times — all from one retrieval bug that returned a shared FAQ document containing a real customer's statement. Without the output guardrail, 14 customers would have seen someone else's data.

Follow-up questions to expect

  • "Why not just put the rules in the system prompt?" — Do both, but the prompt is persuasion: it lowers the rate of bad outputs. A guardrail in code is enforcement. Only enforcement survives a determined attacker.
  • "What does a guardrail add to latency?" — Deterministic checks add almost nothing; a classifier adds tens to a few hundred milliseconds. Run independent checks in parallel with the main call where possible.
  • "How do you measure a guardrail?" — Block rate, false-positive rate on a benign set, and recall on a labelled attack set, tracked per release.