AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do tool-level guardrails differ from content-level guardrails?


Guarding what it says versus what it doesContent-level• Checks text in and out• Moderation, PII, groundedness• Probabilistic, can be rephrased around• Handles tone, toxicity, leaked dataTool-level• Checks the action before it runs• Allowlist, limits, scopes, approval• Deterministic, true or false• Handles money, deletion, sending
A polite request filed 1,400 disputes past the content filter; a one-transaction limit in the tool would have stopped it at one.

What you need to know

Content-level guardrails

  • Inspect what the model reads and writes
  • Moderation, PII scan, groundedness, schema
  • Probabilistic: have false positives and misses
  • Can be fooled by rephrasing

Tool-level guardrails

  • Inspect what the model tries to do
  • Allowlists, argument limits, scopes, approvals
  • Deterministic: true or false
  • Cannot be argued with

A tool policy in code

Python
POLICY = {    "search_docs":   {"mode": "auto"},    "send_email":    {"mode": "approve", "allowed_domains": {"bank.example"}},    "issue_refund":  {"mode": "approve", "max_amount": 5000},    "delete_record": {"mode": "deny"},}def check_tool_call(name, args):    rule = POLICY.get(name)    if rule is None or rule["mode"] == "deny":        return "deny"                              # unknown tool -> no    if "max_amount" in rule and args.get("amount", 0) > rule["max_amount"]:        return "deny"    if "allowed_domains" in rule:        if args.get("to", "").rsplit("@", 1)[-1] not in rule["allowed_domains"]:            return "deny"    return rule["mode"]                            # "auto" or "approve"print(check_tool_call("issue_refund", {"amount": 1200}))    # approveprint(check_tool_call("issue_refund", {"amount": 50000}))   # denyprint(check_tool_call("send_email", {"to": "x@attacker.example"}))  # denyprint(check_tool_call("transfer_funds", {}))                # deny

This function runs in the tool-execution layer, after the model proposes a call and before anything happens. Unknown tools are denied by default. "approve" means the call is paused and shown to a person. Nothing the model writes can change this table.

What each layer is good for

  • Tool level: money, deletion, external communication, changes to production, access to other tenants' data.
  • Content level: toxicity, self-harm, tone, leaked PII in text, unsupported claims, format — things a permission table cannot express.

You need both. A tool guardrail cannot stop the bot from insulting a customer; a content guardrail cannot reliably stop a cleverly worded injection from triggering a refund.

Improper output handling

OWASP's LLM05 is the tool-level risk on the other side: passing model output straight into a shell, SQL query, HTML page or API. Treat model output as untrusted user input — parameterise queries, escape HTML, validate against a schema.

A real-life example

A bank's customer chatbot can look up transactions and raise disputes. The first design relied on a content filter to block "suspicious" dispute requests. A tester wrote a polite request to dispute "all transactions this year" and it went straight through: 1,400 disputes filed on one account.

The fix moved the rule into the tool layer: raise_dispute accepts one transaction ID that belongs to the logged-in customer, at most 5 disputes per day, and anything above Rs 50,000 goes to an agent for approval. The content filter stayed, but only for abusive language and PII in replies.

Follow-up questions to expect

  • "Where does the approval UI come from?" — The tool layer returns "pending approval" to the agent and creates a card for the user or an operator showing the exact action and arguments. The agent resumes only after approval.
  • "How do you choose what needs approval?" — By reversibility and impact: anything that moves money, deletes data, or leaves the organisation.
  • "What about read-only tools?" — They still need scoping. A read tool with access to all customers is a data-leak tool if the output is shown to one customer.