Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do tool-level guardrails differ from content-level guardrails?
What you need to know
Content-level guardrails
- Inspect what the model reads and writes
- Moderation, PII scan, groundedness, schema
- Probabilistic: have false positives and misses
- Can be fooled by rephrasing
Tool-level guardrails
- Inspect what the model tries to do
- Allowlists, argument limits, scopes, approvals
- Deterministic: true or false
- Cannot be argued with
A tool policy in code
1POLICY = {2 "search_docs": {"mode": "auto"},3 "send_email": {"mode": "approve", "allowed_domains": {"bank.example"}},4 "issue_refund": {"mode": "approve", "max_amount": 5000},5 "delete_record": {"mode": "deny"},6}78def check_tool_call(name, args):9 rule = POLICY.get(name)10 if rule is None or rule["mode"] == "deny":11 return "deny" # unknown tool -> no12 if "max_amount" in rule and args.get("amount", 0) > rule["max_amount"]:13 return "deny"14 if "allowed_domains" in rule:15 if args.get("to", "").rsplit("@", 1)[-1] not in rule["allowed_domains"]:16 return "deny"17 return rule["mode"] # "auto" or "approve"1819print(check_tool_call("issue_refund", {"amount": 1200})) # approve20print(check_tool_call("issue_refund", {"amount": 50000})) # deny21print(check_tool_call("send_email", {"to": "x@attacker.example"})) # deny22print(check_tool_call("transfer_funds", {})) # denyThis function runs in the tool-execution layer, after the model proposes a call and before anything happens. Unknown tools are denied by default. "approve" means the call is paused and shown to a person. Nothing the model writes can change this table.
What each layer is good for
- Tool level: money, deletion, external communication, changes to production, access to other tenants' data.
- Content level: toxicity, self-harm, tone, leaked PII in text, unsupported claims, format — things a permission table cannot express.
You need both. A tool guardrail cannot stop the bot from insulting a customer; a content guardrail cannot reliably stop a cleverly worded injection from triggering a refund.
Improper output handling
OWASP's LLM05 is the tool-level risk on the other side: passing model output straight into a shell, SQL query, HTML page or API. Treat model output as untrusted user input — parameterise queries, escape HTML, validate against a schema.
A real-life example
A bank's customer chatbot can look up transactions and raise disputes. The first design relied on a content filter to block "suspicious" dispute requests. A tester wrote a polite request to dispute "all transactions this year" and it went straight through: 1,400 disputes filed on one account.
The fix moved the rule into the tool layer: raise_dispute accepts one transaction ID that belongs to the logged-in customer, at most 5 disputes per day, and anything above Rs 50,000 goes to an agent for approval. The content filter stayed, but only for abusive language and PII in replies.
Follow-up questions to expect
- "Where does the approval UI come from?" — The tool layer returns "pending approval" to the agent and creates a card for the user or an operator showing the exact action and arguments. The agent resumes only after approval.
- "How do you choose what needs approval?" — By reversibility and impact: anything that moves money, deletes data, or leaves the organisation.
- "What about read-only tools?" — They still need scoping. A read tool with access to all customers is a data-leak tool if the output is shown to one customer.