Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do input and output guardrails protect AI applications?
What you need to know
Input guardrails (before the model)
- Identity, rate limits and budgets: who is asking, and how much can they ask.
- Validation: length caps, allowed file types, JSON schema for structured inputs.
- Moderation: classify the request into policy categories (self-harm, hate, violence, illegal activity).
- Injection and jailbreak screening: on the user's text and on retrieved documents or tool results.
- PII redaction: remove card numbers, Aadhaar numbers and similar before the text leaves your boundary.
- Topic scoping: decline off-domain requests cheaply ("I can only help with your account").
Output guardrails (after the model)
- Safety classification of the reply.
- PII and secret scanning: the model can echo data from context that this user should not see.
- Groundedness: is each claim supported by the retrieved context?
- Format: schema validation for structured output.
- Tool-call validation: is this tool allowed, are the arguments in range, does it need human approval?
Common tools, described accurately
| Tool | What it is |
|---|---|
| Provider moderation endpoints | OpenAI's moderation endpoint (omni-moderation-latest) returns a flag and per-category scores for text and images. Azure AI Content Safety returns severity levels for hate, sexual, violence and self-harm, and its Prompt Shields detect injection in user prompts and documents. AWS Bedrock Guardrails and Google's Gemini safety settings offer configurable content filters. |
| Llama Guard (Meta) | An open-weight LLM fine-tuned as a classifier. It reads a user prompt or model response and outputs safe or unsafe plus the violated hazard codes (for example S1 for violent crimes). Llama Guard 3 and 4 follow the MLCommons hazard taxonomy; Meta's separate Prompt Guard models detect injection and jailbreak attempts. |
| NeMo Guardrails (NVIDIA) | An open-source toolkit where you configure "rails" in YAML and the Colang language: input, output, dialog, retrieval and execution rails. It can call other checkers, such as Llama Guard or a self-check prompt. |
| Guardrails AI | An open-source Python library. A Guard wraps an LLM call and runs validators from its Hub (PII, toxicity, regex, schema), with an on_fail policy such as raise an exception, fix, re-ask or filter. |
Design rules
- Cheap first: regex and schema checks cost microseconds; a classifier costs tens to hundreds of milliseconds. Run the cheap ones first.
- Fail closed on high-risk paths: if the moderation call times out on the payments route, block. Fail open only where blocking harms more, and write that decision down.
- Streaming: either buffer to a sentence or paragraph boundary before checking, or check the stream and be ready to retract.
A real-life example
A bank's customer chatbot handles 40,000 chats a day. Input: a regex plus Luhn check removes card numbers before the text reaches the vendor model; Llama Guard screens for abuse; a small classifier rejects questions unrelated to banking. Output: a PII scanner blocks replies that contain another customer's account number, and a schema check ensures the "transfer" intent returns only a draft for the user to confirm.
In the first week, output PII scanning fires 14 times — all from one retrieval bug that returned a shared FAQ document containing a real customer's statement. Without the output guardrail, 14 customers would have seen someone else's data.
Follow-up questions to expect
- "Why not just put the rules in the system prompt?" — Do both, but the prompt is persuasion: it lowers the rate of bad outputs. A guardrail in code is enforcement. Only enforcement survives a determined attacker.
- "What does a guardrail add to latency?" — Deterministic checks add almost nothing; a classifier adds tens to a few hundred milliseconds. Run independent checks in parallel with the main call where possible.
- "How do you measure a guardrail?" — Block rate, false-positive rate on a benign set, and recall on a labelled attack set, tracked per release.