Course Content
AI Security: Preventing Prompt Injection
3 sections · 7 lessons
Policy-Based Guardrails
A B2B analytics company added an assistant to its product. Security was covered, the team said, because the assistant reused the web application's existing role system. A user with the support_agent role got the support_agent permissions. Nothing new was granted.
Six weeks later a customer's shared inbox received a spreadsheet containing 41,000 rows of another tenant's transaction history. Traced back: a support agent had pasted a customer's email thread into the assistant to ask for a summary. The thread contained instructions. The assistant called export_transactions with a tenant filter it had been told to use, then send_attachment, and both calls were authorised. The support_agent role really did include both permissions.
It had included them for three years without incident, because in the web application those permissions sat behind screens that only ever offered the current tenant, a maximum of 500 rows, and a recipient field pre-filled with the logged-in user's own address. The role was never the control. The user interface was the control. Hand the same role to something that calls the API directly, with arguments chosen by text a stranger wrote, and three years of apparent safety evaporates in a single request.
Guardrails are the machinery that replaces the missing user interface: the thing that decides, per call, with full knowledge of who is asking and why, whether this specific action is permitted. This lesson builds that machinery.
The architecture: four places a decision can be made
A guardrail system is a set of decision points arranged around the model. Each one sees different information and can enforce different things, and the reason to be explicit about the arrangement is that people habitually put all their logic at the wrong point.
| Point | Sees | Can enforce | Cannot enforce |
|---|---|---|---|
| Pre-request | Identity, entitlements, quota, the raw request | Who may use the feature, rate and cost caps, which tools are even offered | Anything about what the model will decide |
| Tool-call authorisation | Tool name, arguments, session context, history so far | The real decisions — ownership, limits, allow-lists, approval routing | Content policy on prose |
| Decode-time | Partial token stream | Output shape: valid JSON, schema conformance, stop conditions | Meaning, intent, or truth |
| Post-response | The finished text and the full call trace | Redaction, egress rules, content policy, audit | Anything already executed by a tool |
The row that carries the load is tool-call authorisation, and it is the one most often reduced to a single boolean check at startup: "does this user's role permit this tool?" That question is asked once, at the wrong time, with none of the information that matters — the arguments, the target record, the amount, the recipient.
Authorise the call, not the tool. "May this user use
export_transactions?" is a question with no useful answer. "May this session export tenant 8812's transactions, 41,000 rows, right now?" has an obvious one.
Layer 1: capabilities instead of roles
The idea in plain language
A role is a label on the person: "you are a support agent, and support agents may do these fourteen things." A capability is a token attached to a specific permission: "the bearer of this may read orders belonging to customer 4471, until 14:32, up to 20 times." The difference is that a role is a static bundle granted broadly in advance, while a capability is minted narrowly for one purpose and expires.
Think of a hotel. The role model is a master key stamped "housekeeping" that opens every room on three floors, all week. The capability model is a key card cut for room 412, valid until Thursday, that stops working the moment the guest checks out. If a housekeeper's key is stolen, the difference is whether the thief gets one room or ninety.
Why the distinction matters more for an LLM than for a web app
Count the surface. Suppose eight roles and forty tools. The role matrix has 8 × 40 = 320 cells, each a yes-or-no decision someone made at some point, most of them years ago for a reason nobody remembers. Reviewing that at two minutes a cell is nearly eleven hours of work, so it does not get reviewed, so every cell decays towards "yes" as one-off requests accumulate.
Now look at what a single session needs. The support agent summarising an email thread needed exactly three things: read this ticket, read the ticket's own customer record, post a reply to this ticket. The support_agent role offered fourteen tools. The over-grant factor is 14 / 3 ≈ 4.7× — and the eleven unnecessary tools included the two that caused the incident.
In the web application that over-grant was invisible, because the screens never offered those eleven. An LLM has no screens. Every tool in the set is one persuasive sentence away from being called.
1import time, secrets2from dataclasses import dataclass, field34@dataclass(frozen=True)5class Capability:6 action: str # "orders.read"7 resource: str # "customer:4471" or "ticket:99120"8 constraints: dict = field(default_factory=dict) # max_rows, max_amount...9 expires_at: float = 0.010 token: str = field(default_factory=lambda: secrets.token_urlsafe(16))1112 def permits(self, action: str, resource: str, args: dict) -> tuple[bool, str]:13 if time.time() > self.expires_at:14 return False, "capability expired"15 if action != self.action:16 return False, f"action {action} not granted"17 if resource != self.resource and not resource.startswith(self.resource + "/"):18 return False, f"resource {resource} outside grant {self.resource}"19 for key, limit in self.constraints.items():20 if key in args and args[key] > limit:21 return False, f"{key}={args[key]} exceeds limit {limit}"22 return True, "ok"232425def mint_for_session(session) -> list[Capability]:26 """Called once per request, server-side, from verified session state.27 Nothing the model or the user can write reaches this function."""28 now = time.time()29 caps = [30 Capability("tickets.read", f"ticket:{session.ticket_id}",31 expires_at=now + 300),32 Capability("orders.read", f"customer:{session.customer_id}",33 constraints={"max_rows": 50}, expires_at=now + 300),34 Capability("tickets.reply", f"ticket:{session.ticket_id}",35 constraints={"max_chars": 4000}, expires_at=now + 300),36 ]37 if session.user.has_clearance("finance") and session.approved_by_human:38 caps.append(Capability("refunds.issue", f"order:{session.order_id}",39 constraints={"max_amount": 500},40 expires_at=now + 120))41 return capsFour properties do the work, and each closes a specific failure from the opening story. Capabilities are minted per request from server-side session state, so the customer id in the grant cannot be influenced by anything in the email thread. They name a specific resource, so export_transactions for tenant 8812 has no matching grant regardless of how convincingly it is requested. They carry numeric constraints, so 41,000 rows fails against max_rows: 50. And they expire, so a payload that lies dormant in memory finds nothing to use when it wakes.
The tool set follows the capabilities
One more move makes the whole thing coherent: build the tool schema you send to the model from the minted capabilities. If the session has no refunds.issue capability, the model is never told that a refund tool exists. That is not security by obscurity — the authorisation check still runs — but it removes the temptation entirely, and it shrinks the prompt.
Layer 2: request-level guardrails
These run before the model and enforce properties of the request as a whole. They are cheap, deterministic, and they catch a class of abuse that per-call checks cannot see: the abuse that consists of doing permitted things too many times.
| Guardrail | Typical limit | Attack it stops |
|---|---|---|
| Requests per user per minute | 20 | Brute-forcing a probabilistic filter with variants |
| Tool calls per request | 10 | An agent looping through a whole table one record at a time |
| Total tokens per request | 50,000 | Denial of wallet through induced generation |
| Cost per user per day | A fixed ceiling | Sustained resource abuse |
| Wall-clock per request | 60 s | Runaway agent loops |
| Distinct resources touched | 5 | Bulk exfiltration — the export in the opening story |
| Input size | 32 KB | Payload burial in a huge document |
The "distinct resources touched" row is the underrated one. Legitimate sessions are narrow: a support agent works one ticket and one customer at a time. A session that reads its ninth distinct customer record is doing something no workflow requires, and that single counter catches slow, patient exfiltration that stays under every per-call limit.
Layer 3: decode-time constraints, and their honest scope
You can constrain generation itself — biasing or forbidding tokens, forcing output to match a grammar or JSON schema, setting stop sequences. This is genuinely powerful for one purpose and close to useless for another, and conflating the two is a common mistake.
| Use | Mechanism | Verdict |
|---|---|---|
| Force valid JSON matching a schema | Grammar-constrained decoding: only tokens keeping the output parseable are permitted | Strong and deterministic. The output cannot be malformed, so downstream parsing cannot be attacked |
| Restrict a field to an enum | Same, with the enum in the grammar | Strong. "category" can only be one of your five values, whatever the model was told |
| Stop generation at a marker | Stop sequences | Useful for bounding length and preventing role-play continuation |
| Ban specific words | Logit bias set to negative infinity on those token ids | Weak. Blocks the spelling, not the concept |
| Prevent harmful topics | Long ban lists | Do not. Degrades fluency and fails on synonyms |
The reason word banning fails is mechanical. A ban list operates on token ids, and any word has many token spellings — different casing, a leading space, split across subwords, a Unicode lookalike, the same word in another language, or a synonym you did not think of. Banning one form pushes the model to the next. Meanwhile every banned token slightly distorts generation everywhere else, so you pay a fluency cost on 100% of traffic for a control that fails on the traffic you care about.
Constraining structure is different in kind. If a quarantined model's only permitted output is an object matching {"category": one of five strings, "confidence": number}, then no instruction inside the document it read can make it emit anything else. The instruction may change which of the five it picks — that is a real risk and you handle it elsewhere — but it cannot make the model emit a URL, a command, or a paragraph of prose. That is a hard boundary produced by ordinary software, which is exactly the kind you want.
Constrain the shape of output with a grammar; never try to constrain its meaning with a token ban list. The first is enforcement, the second is a fluency tax with a bypass.
Layer 4: the policy engine
Scattering these checks through your tool implementations guarantees inconsistency — one tool checks ownership, another forgot, a third checks it differently. Centralise the decision and give it a single shape.
1from enum import Enum23class Effect(Enum):4 ALLOW = "allow"5 DENY = "deny"6 APPROVE = "require_human_approval"7 TRANSFORM = "allow_with_modification"89class Decision:10 def __init__(self, effect, reason, rule_id, policy_version,11 obligations=None, modified_args=None):12 self.effect = effect13 self.reason = reason # shown to the user AND logged14 self.rule_id = rule_id # which rule decided this15 self.policy_version = policy_version16 self.obligations = obligations or [] # e.g. ["audit", "notify_dpo"]17 self.modified_args = modified_args1819class PolicyEngine:20 def __init__(self, policy):21 self.policy = policy2223 def authorise(self, session, tool, args) -> Decision:24 v = self.policy["version"]2526 # 1. Capability check - deterministic, cannot be argued with.27 action, resource = TOOL_MAP[tool](args)28 why = "no capabilities minted for this session"29 for cap in session.capabilities:30 ok, why = cap.permits(action, resource, args)31 if ok:32 break33 else:34 return Decision(Effect.DENY, why, "cap.no_grant", v)3536 # 2. Rate and budget checks.37 if session.tool_calls >= self.policy["limits"]["calls_per_request"]:38 return Decision(Effect.DENY, "call budget exhausted",39 "limit.calls", v)40 if session.distinct_resources() >= self.policy["limits"]["resources"]:41 return Decision(Effect.DENY, "too many distinct records in one session",42 "limit.resources", v, obligations=["alert_security"])4344 # 3. Argument-level rules from the policy document.45 for rule in self.policy["rules"].get(tool, []):46 if matches(rule["when"], args, session):47 return Decision(Effect[rule["effect"]], rule["reason"],48 rule["id"], v,49 obligations=rule.get("obligations"),50 modified_args=apply_transform(rule, args))5152 # 4. Default. Never make this ALLOW.53 return Decision(Effect.ALLOW if tool in self.policy["default_allow"]54 else Effect.DENY,55 "default", "default", v)Three details are load-bearing. The default at the bottom is deny unless explicitly listed — a fail-open default means that every tool added in future is permitted until someone remembers to write a rule, which is the same decay that filled the role matrix. Every decision carries a rule id and a policy version, so an audit six months later can name the exact line that allowed something. And obligations lets a rule allow an action and require side effects — write an audit record, notify the data protection officer, alert security — which is how compliance requirements attach to permissions without being duplicated in every tool.
Approval that a compromised model cannot forge
When the effect is APPROVE, the human must see something the model did not write. Render the approval card from the resolved arguments: the recipient your code looked up, the amount as a number, the record id, the tenant name from your database. If you show the model's own summary of what it is about to do, an injection simply writes a reassuring summary, and the approval step becomes a rubber stamp with an audit trail — worse than no gate, because it manufactures evidence of human oversight that did not happen.
Versioning and safe rollout
A policy change is a production change with the same blast radius as a code deploy, and a bad one fails in a distinctive way: it does not error, it silently denies work that should have been allowed, and nobody notices until the complaints arrive.
Shadow mode, with the arithmetic that justifies it
Run the candidate policy alongside the live one, log what it would have decided, change nothing.
A team ran policy v5 in shadow for seven days across 240,000 requests. It would have denied 1,830 calls that v4 allowed. They sampled 100 of those and reviewed them by hand: 6 were genuinely abusive. Extrapolating, roughly 0.06 × 1,830 ≈ 110 real catches and about 1,720 new false denials — 94% of the new denials would have been legitimate work blocked, at a rate of about 246 a day.
Shipping that policy would have looked, from the dashboard, like a security improvement: denials up sharply, incidents flat. Shadow mode plus a hundred hand-reviewed samples cost an afternoon and prevented a fortnight of user-hostile behaviour.
The rollout sequence
| Stage | What runs | Promote when |
|---|---|---|
| Shadow | Candidate evaluated, decisions logged only | Divergences sampled and understood; false-denial rate acceptable |
| Canary | Live for 1% of sessions | No rise in support contacts or task-abandonment for that cohort |
| Staged | 10% → 50% → 100%, with a hold at each step | Metrics stable across a full weekday cycle |
| Full | Everyone; previous version retained | Rollback path verified and one-command |
Two rules make rollback actually work. Policies are immutable, versioned artefacts — you publish v5.1, you never edit v5 in place, so the version recorded in an old audit entry still describes real content. And the engine can hold multiple versions simultaneously, keyed by session, so rolling back is a routing change rather than a redeploy.
Escalation and compliance
Denials and approvals are not the end of the process; they are the start of a workflow that has to be designed or it will be improvised badly under pressure.
| Trigger | Route | Target | If nobody responds |
|---|---|---|---|
| Approval required, routine (refund under 500) | Team queue | 15 minutes | Expire and tell the user, do not auto-approve |
| Approval required, high value | Named approver plus a second | 1 hour | Expire; escalate to the manager |
| Repeated denials in one session | Security on-call | Immediate | Auto-suspend the session |
| Canary token in output | Page the security team | Immediate | Auto-disable the affected tool |
| Personal data leaving the region | Data protection officer | Same day | Block by default, never allow by default |
The right-hand column is the one people forget to specify, and it is where systems fail quietly. An approval queue with no expiry policy eventually gets an "auto-approve after 24 hours" rule added by someone trying to unblock a backlog, and at that point the gate is gone.
What auditors and regulators actually ask for
Across data protection regimes and the newer AI-specific rules, the recurring demands are consistent enough to design for directly:
- Show the decision. For a given past action, name the rule and the policy version that permitted it, and produce the inputs it was evaluated against.
- Show the human. Where oversight is claimed, evidence that a person saw the resolved facts and acted — with a timestamp and an identity, not a summary the system generated.
- Show the boundary. Which categories of personal data the system can reach, and what prevents it reaching more.
- Show the exceptions. Every override, who made it, on what grounds, and whether it was time-bounded.
- Show retention. How long decision logs live and how personal data inside them is minimised — audit logs that store raw prompts often become the largest uncontrolled copy of customer data in the company.
All five fall out of the design above almost for free, provided rule_id, policy_version, resolved arguments and approver identity are written on every decision from the start. Retrofitting them afterwards is the expensive path, because the decisions you most need to explain have already happened.
Misconceptions worth naming
| Belief | Correction |
|---|---|
| "We reuse the app's roles, so we inherit its security" | The UI was doing the constraining. An LLM calls the API with arguments an attacker can influence. |
| "The system prompt lists the rules" | Prose in the context is a preference. Guardrails run in code the model cannot reach. |
| "Checking the tool is permitted is enough" | The arguments are the attack. Authorise the call with its arguments, every time. |
| "Ban the dangerous words at decode time" | Bans block spellings, not concepts, and cost fluency on all traffic. Constrain structure instead. |
| "A human approves, so we're safe" | Only if the human sees resolved facts from your code rather than the model's description. |
| "Default-allow with a deny-list is more usable" | Every new tool is permitted until someone writes a rule. Default deny, then grant deliberately. |
| "Policy changes are low risk, they only tighten things" | Tightening blocks legitimate work silently. Shadow-run every change and sample the divergences. |
What this means when you build something
The most valuable half-hour you can spend on an LLM feature is drawing the capability table for a single real session. Take one concrete task — summarise this ticket and reply to it — and write down the minimum set of actions, each with its specific resource and its numeric limits. Then compare that list with what the role your feature currently uses actually permits. The ratio between them is your over-grant factor, and in most systems it is between 4× and 20×.
Everything else follows from closing that gap. Mint capabilities per request from server-side session state, so nothing in the model's context can widen them. Build the tool schema from the capabilities, so unavailable actions are not even described. Route every call through one policy engine that defaults to deny, records a rule id and a policy version, and can return "require approval" as easily as "allow". Add the session-level counters — calls, tokens, cost, distinct resources — because they catch the patient attacks that pass every individual check.
Then treat the policy itself as software. Version it, shadow-run every change against a week of real traffic, hand-review a sample of the divergences before promoting, and keep the previous version loadable so rollback is a routing decision. A team that does this ends up in a position where an injection that fully persuades the model still cannot export another tenant's data, because there is no capability naming that tenant, no grant permitting 41,000 rows, and no tool that will send an attachment to an address the model chose. The model was fooled; nothing happened. That is the entire objective.