AI Security: Preventing Prompt Injection

Policy-Based Guardrails


A B2B analytics company added an assistant to its product. Security was covered, the team said, because the assistant reused the web application's existing role system. A user with the support_agent role got the support_agent permissions. Nothing new was granted.

Six weeks later a customer's shared inbox received a spreadsheet containing 41,000 rows of another tenant's transaction history. Traced back: a support agent had pasted a customer's email thread into the assistant to ask for a summary. The thread contained instructions. The assistant called export_transactions with a tenant filter it had been told to use, then send_attachment, and both calls were authorised. The support_agent role really did include both permissions.

It had included them for three years without incident, because in the web application those permissions sat behind screens that only ever offered the current tenant, a maximum of 500 rows, and a recipient field pre-filled with the logged-in user's own address. The role was never the control. The user interface was the control. Hand the same role to something that calls the API directly, with arguments chosen by text a stranger wrote, and three years of apparent safety evaporates in a single request.

Guardrails are the machinery that replaces the missing user interface: the thing that decides, per call, with full knowledge of who is asking and why, whether this specific action is permitted. This lesson builds that machinery.

Inheriting a role versus granting a capabilityReused web-app role• support_agent can read any customer• The model inherits the whole role• An injection gets the entire set• Approval is text the model producedPer-request capability• This ticket, this customer, this action• Scoped to the subject of the request• An injection reaches only that scope• Approval signed outside the model
A role is a standing grant and an LLM request is an untrusted one, so the permission has to be minted per request rather than inherited from the caller.

The architecture: four places a decision can be made

A guardrail system is a set of decision points arranged around the model. Each one sees different information and can enforce different things, and the reason to be explicit about the arrangement is that people habitually put all their logic at the wrong point.

PointSeesCan enforceCannot enforce
Pre-requestIdentity, entitlements, quota, the raw requestWho may use the feature, rate and cost caps, which tools are even offeredAnything about what the model will decide
Tool-call authorisationTool name, arguments, session context, history so farThe real decisions — ownership, limits, allow-lists, approval routingContent policy on prose
Decode-timePartial token streamOutput shape: valid JSON, schema conformance, stop conditionsMeaning, intent, or truth
Post-responseThe finished text and the full call traceRedaction, egress rules, content policy, auditAnything already executed by a tool

The row that carries the load is tool-call authorisation, and it is the one most often reduced to a single boolean check at startup: "does this user's role permit this tool?" That question is asked once, at the wrong time, with none of the information that matters — the arguments, the target record, the amount, the recipient.

Authorise the call, not the tool. "May this user use export_transactions?" is a question with no useful answer. "May this session export tenant 8812's transactions, 41,000 rows, right now?" has an obvious one.

Layer 1: capabilities instead of roles

The idea in plain language

A role is a label on the person: "you are a support agent, and support agents may do these fourteen things." A capability is a token attached to a specific permission: "the bearer of this may read orders belonging to customer 4471, until 14:32, up to 20 times." The difference is that a role is a static bundle granted broadly in advance, while a capability is minted narrowly for one purpose and expires.

Think of a hotel. The role model is a master key stamped "housekeeping" that opens every room on three floors, all week. The capability model is a key card cut for room 412, valid until Thursday, that stops working the moment the guest checks out. If a housekeeper's key is stolen, the difference is whether the thief gets one room or ninety.

Why the distinction matters more for an LLM than for a web app

Count the surface. Suppose eight roles and forty tools. The role matrix has 8 × 40 = 320 cells, each a yes-or-no decision someone made at some point, most of them years ago for a reason nobody remembers. Reviewing that at two minutes a cell is nearly eleven hours of work, so it does not get reviewed, so every cell decays towards "yes" as one-off requests accumulate.

Now look at what a single session needs. The support agent summarising an email thread needed exactly three things: read this ticket, read the ticket's own customer record, post a reply to this ticket. The support_agent role offered fourteen tools. The over-grant factor is 14 / 3 ≈ 4.7× — and the eleven unnecessary tools included the two that caused the incident.

In the web application that over-grant was invisible, because the screens never offered those eleven. An LLM has no screens. Every tool in the set is one persuasive sentence away from being called.

Python
import time, secretsfrom dataclasses import dataclass, field@dataclass(frozen=True)class Capability:    action: str                     # "orders.read"    resource: str                   # "customer:4471" or "ticket:99120"    constraints: dict = field(default_factory=dict)   # max_rows, max_amount...    expires_at: float = 0.0    token: str = field(default_factory=lambda: secrets.token_urlsafe(16))    def permits(self, action: str, resource: str, args: dict) -> tuple[bool, str]:        if time.time() > self.expires_at:            return False, "capability expired"        if action != self.action:            return False, f"action {action} not granted"        if resource != self.resource and not resource.startswith(self.resource + "/"):            return False, f"resource {resource} outside grant {self.resource}"        for key, limit in self.constraints.items():            if key in args and args[key] > limit:                return False, f"{key}={args[key]} exceeds limit {limit}"        return True, "ok"def mint_for_session(session) -> list[Capability]:    """Called once per request, server-side, from verified session state.    Nothing the model or the user can write reaches this function."""    now = time.time()    caps = [        Capability("tickets.read",  f"ticket:{session.ticket_id}",                   expires_at=now + 300),        Capability("orders.read",   f"customer:{session.customer_id}",                   constraints={"max_rows": 50}, expires_at=now + 300),        Capability("tickets.reply", f"ticket:{session.ticket_id}",                   constraints={"max_chars": 4000}, expires_at=now + 300),    ]    if session.user.has_clearance("finance") and session.approved_by_human:        caps.append(Capability("refunds.issue", f"order:{session.order_id}",                               constraints={"max_amount": 500},                               expires_at=now + 120))    return caps

Four properties do the work, and each closes a specific failure from the opening story. Capabilities are minted per request from server-side session state, so the customer id in the grant cannot be influenced by anything in the email thread. They name a specific resource, so export_transactions for tenant 8812 has no matching grant regardless of how convincingly it is requested. They carry numeric constraints, so 41,000 rows fails against max_rows: 50. And they expire, so a payload that lies dormant in memory finds nothing to use when it wakes.

The tool set follows the capabilities

One more move makes the whole thing coherent: build the tool schema you send to the model from the minted capabilities. If the session has no refunds.issue capability, the model is never told that a refund tool exists. That is not security by obscurity — the authorisation check still runs — but it removes the temptation entirely, and it shrinks the prompt.

Layer 2: request-level guardrails

These run before the model and enforce properties of the request as a whole. They are cheap, deterministic, and they catch a class of abuse that per-call checks cannot see: the abuse that consists of doing permitted things too many times.

GuardrailTypical limitAttack it stops
Requests per user per minute20Brute-forcing a probabilistic filter with variants
Tool calls per request10An agent looping through a whole table one record at a time
Total tokens per request50,000Denial of wallet through induced generation
Cost per user per dayA fixed ceilingSustained resource abuse
Wall-clock per request60 sRunaway agent loops
Distinct resources touched5Bulk exfiltration — the export in the opening story
Input size32 KBPayload burial in a huge document

The "distinct resources touched" row is the underrated one. Legitimate sessions are narrow: a support agent works one ticket and one customer at a time. A session that reads its ninth distinct customer record is doing something no workflow requires, and that single counter catches slow, patient exfiltration that stays under every per-call limit.

Layer 3: decode-time constraints, and their honest scope

You can constrain generation itself — biasing or forbidding tokens, forcing output to match a grammar or JSON schema, setting stop sequences. This is genuinely powerful for one purpose and close to useless for another, and conflating the two is a common mistake.

UseMechanismVerdict
Force valid JSON matching a schemaGrammar-constrained decoding: only tokens keeping the output parseable are permittedStrong and deterministic. The output cannot be malformed, so downstream parsing cannot be attacked
Restrict a field to an enumSame, with the enum in the grammarStrong. "category" can only be one of your five values, whatever the model was told
Stop generation at a markerStop sequencesUseful for bounding length and preventing role-play continuation
Ban specific wordsLogit bias set to negative infinity on those token idsWeak. Blocks the spelling, not the concept
Prevent harmful topicsLong ban listsDo not. Degrades fluency and fails on synonyms

The reason word banning fails is mechanical. A ban list operates on token ids, and any word has many token spellings — different casing, a leading space, split across subwords, a Unicode lookalike, the same word in another language, or a synonym you did not think of. Banning one form pushes the model to the next. Meanwhile every banned token slightly distorts generation everywhere else, so you pay a fluency cost on 100% of traffic for a control that fails on the traffic you care about.

Constraining structure is different in kind. If a quarantined model's only permitted output is an object matching {"category": one of five strings, "confidence": number}, then no instruction inside the document it read can make it emit anything else. The instruction may change which of the five it picks — that is a real risk and you handle it elsewhere — but it cannot make the model emit a URL, a command, or a paragraph of prose. That is a hard boundary produced by ordinary software, which is exactly the kind you want.

Constrain the shape of output with a grammar; never try to constrain its meaning with a token ban list. The first is enforcement, the second is a fluency tax with a bypass.

Layer 4: the policy engine

Scattering these checks through your tool implementations guarantees inconsistency — one tool checks ownership, another forgot, a third checks it differently. Centralise the decision and give it a single shape.

Python
from enum import Enumclass Effect(Enum):    ALLOW    = "allow"    DENY     = "deny"    APPROVE  = "require_human_approval"    TRANSFORM = "allow_with_modification"class Decision:    def __init__(self, effect, reason, rule_id, policy_version,                 obligations=None, modified_args=None):        self.effect = effect        self.reason = reason            # shown to the user AND logged        self.rule_id = rule_id          # which rule decided this        self.policy_version = policy_version        self.obligations = obligations or []   # e.g. ["audit", "notify_dpo"]        self.modified_args = modified_argsclass PolicyEngine:    def __init__(self, policy):        self.policy = policy    def authorise(self, session, tool, args) -> Decision:        v = self.policy["version"]        # 1. Capability check - deterministic, cannot be argued with.        action, resource = TOOL_MAP[tool](args)        why = "no capabilities minted for this session"        for cap in session.capabilities:            ok, why = cap.permits(action, resource, args)            if ok:                break        else:            return Decision(Effect.DENY, why, "cap.no_grant", v)        # 2. Rate and budget checks.        if session.tool_calls >= self.policy["limits"]["calls_per_request"]:            return Decision(Effect.DENY, "call budget exhausted",                            "limit.calls", v)        if session.distinct_resources() >= self.policy["limits"]["resources"]:            return Decision(Effect.DENY, "too many distinct records in one session",                            "limit.resources", v, obligations=["alert_security"])        # 3. Argument-level rules from the policy document.        for rule in self.policy["rules"].get(tool, []):            if matches(rule["when"], args, session):                return Decision(Effect[rule["effect"]], rule["reason"],                                rule["id"], v,                                obligations=rule.get("obligations"),                                modified_args=apply_transform(rule, args))        # 4. Default. Never make this ALLOW.        return Decision(Effect.ALLOW if tool in self.policy["default_allow"]                        else Effect.DENY,                        "default", "default", v)

Three details are load-bearing. The default at the bottom is deny unless explicitly listed — a fail-open default means that every tool added in future is permitted until someone remembers to write a rule, which is the same decay that filled the role matrix. Every decision carries a rule id and a policy version, so an audit six months later can name the exact line that allowed something. And obligations lets a rule allow an action and require side effects — write an audit record, notify the data protection officer, alert security — which is how compliance requirements attach to permissions without being duplicated in every tool.

Approval that a compromised model cannot forge

When the effect is APPROVE, the human must see something the model did not write. Render the approval card from the resolved arguments: the recipient your code looked up, the amount as a number, the record id, the tenant name from your database. If you show the model's own summary of what it is about to do, an injection simply writes a reassuring summary, and the approval step becomes a rubber stamp with an audit trail — worse than no gate, because it manufactures evidence of human oversight that did not happen.

Versioning and safe rollout

A policy change is a production change with the same blast radius as a code deploy, and a bad one fails in a distinctive way: it does not error, it silently denies work that should have been allowed, and nobody notices until the complaints arrive.

Shadow mode, with the arithmetic that justifies it

Run the candidate policy alongside the live one, log what it would have decided, change nothing.

A team ran policy v5 in shadow for seven days across 240,000 requests. It would have denied 1,830 calls that v4 allowed. They sampled 100 of those and reviewed them by hand: 6 were genuinely abusive. Extrapolating, roughly 0.06 × 1,830 ≈ 110 real catches and about 1,720 new false denials — 94% of the new denials would have been legitimate work blocked, at a rate of about 246 a day.

Shipping that policy would have looked, from the dashboard, like a security improvement: denials up sharply, incidents flat. Shadow mode plus a hundred hand-reviewed samples cost an afternoon and prevented a fortnight of user-hostile behaviour.

The rollout sequence

StageWhat runsPromote when
ShadowCandidate evaluated, decisions logged onlyDivergences sampled and understood; false-denial rate acceptable
CanaryLive for 1% of sessionsNo rise in support contacts or task-abandonment for that cohort
Staged10% → 50% → 100%, with a hold at each stepMetrics stable across a full weekday cycle
FullEveryone; previous version retainedRollback path verified and one-command

Two rules make rollback actually work. Policies are immutable, versioned artefacts — you publish v5.1, you never edit v5 in place, so the version recorded in an old audit entry still describes real content. And the engine can hold multiple versions simultaneously, keyed by session, so rolling back is a routing change rather than a redeploy.

Escalation and compliance

Denials and approvals are not the end of the process; they are the start of a workflow that has to be designed or it will be improvised badly under pressure.

TriggerRouteTargetIf nobody responds
Approval required, routine (refund under 500)Team queue15 minutesExpire and tell the user, do not auto-approve
Approval required, high valueNamed approver plus a second1 hourExpire; escalate to the manager
Repeated denials in one sessionSecurity on-callImmediateAuto-suspend the session
Canary token in outputPage the security teamImmediateAuto-disable the affected tool
Personal data leaving the regionData protection officerSame dayBlock by default, never allow by default

The right-hand column is the one people forget to specify, and it is where systems fail quietly. An approval queue with no expiry policy eventually gets an "auto-approve after 24 hours" rule added by someone trying to unblock a backlog, and at that point the gate is gone.

What auditors and regulators actually ask for

Across data protection regimes and the newer AI-specific rules, the recurring demands are consistent enough to design for directly:

  • Show the decision. For a given past action, name the rule and the policy version that permitted it, and produce the inputs it was evaluated against.
  • Show the human. Where oversight is claimed, evidence that a person saw the resolved facts and acted — with a timestamp and an identity, not a summary the system generated.
  • Show the boundary. Which categories of personal data the system can reach, and what prevents it reaching more.
  • Show the exceptions. Every override, who made it, on what grounds, and whether it was time-bounded.
  • Show retention. How long decision logs live and how personal data inside them is minimised — audit logs that store raw prompts often become the largest uncontrolled copy of customer data in the company.

All five fall out of the design above almost for free, provided rule_id, policy_version, resolved arguments and approver identity are written on every decision from the start. Retrofitting them afterwards is the expensive path, because the decisions you most need to explain have already happened.

Misconceptions worth naming

BeliefCorrection
"We reuse the app's roles, so we inherit its security"The UI was doing the constraining. An LLM calls the API with arguments an attacker can influence.
"The system prompt lists the rules"Prose in the context is a preference. Guardrails run in code the model cannot reach.
"Checking the tool is permitted is enough"The arguments are the attack. Authorise the call with its arguments, every time.
"Ban the dangerous words at decode time"Bans block spellings, not concepts, and cost fluency on all traffic. Constrain structure instead.
"A human approves, so we're safe"Only if the human sees resolved facts from your code rather than the model's description.
"Default-allow with a deny-list is more usable"Every new tool is permitted until someone writes a rule. Default deny, then grant deliberately.
"Policy changes are low risk, they only tighten things"Tightening blocks legitimate work silently. Shadow-run every change and sample the divergences.

What this means when you build something

The most valuable half-hour you can spend on an LLM feature is drawing the capability table for a single real session. Take one concrete task — summarise this ticket and reply to it — and write down the minimum set of actions, each with its specific resource and its numeric limits. Then compare that list with what the role your feature currently uses actually permits. The ratio between them is your over-grant factor, and in most systems it is between 4× and 20×.

Everything else follows from closing that gap. Mint capabilities per request from server-side session state, so nothing in the model's context can widen them. Build the tool schema from the capabilities, so unavailable actions are not even described. Route every call through one policy engine that defaults to deny, records a rule id and a policy version, and can return "require approval" as easily as "allow". Add the session-level counters — calls, tokens, cost, distinct resources — because they catch the patient attacks that pass every individual check.

Then treat the policy itself as software. Version it, shadow-run every change against a week of real traffic, hand-review a sample of the divergences before promoting, and keep the previous version loadable so rollback is a routing decision. A team that does this ends up in a position where an injection that fully persuades the model still cannot export another tenant's data, because there is no capability naming that tenant, no grant permitting 41,000 rows, and no tool that will send an attachment to an address the model chose. The model was fooled; nothing happened. That is the entire objective.