AI Security: Preventing Prompt Injection

Safety Evaluation Frameworks


A team ships a customer support assistant. Safety sign-off is a suite of 40 adversarial prompts written by two engineers over an afternoon. The model refuses all 40. Green build, ship it.

Six weeks later the provider releases a new model version. The team swaps it in and re-runs the suite: 40 out of 40 still refused. Nothing regressed, so nothing needs review.

Nine days after the upgrade, a support agent pastes a customer's forwarded email into the assistant. The email contains a paragraph of text formatted to look like a system notice, instructing the assistant to append a summary of the account's recent invoices to its reply. It does. The invoices belong to a different account.

The post-mortem question is the one worth sitting with: what did the passing test suite actually prove? It proved the new model still refused 40 specific strings that two engineers thought of in March. It said nothing about instructions arriving inside a retrieved document rather than a user message, because no test in the suite did that. The suite was a regression check that everyone read as a safety guarantee, and those are different claims.

The pyramid under a green safety buildCheap unit checks,run on every commitVersioned attacksuite with a hold-outSystem evals withtools connectedHuman rubricreview on a sampleProductionmonitoring and incidentstopbottom40 prompts written in one afternoon, never held out, all refused.
A pass rate on the cases you thought of measures the authors of the suite, and coverage of ingestion paths is what it silently leaves at zero.

What a safety score actually claims

Two numbers get confused constantly, and separating them is most of this discipline.

  • Pass rate — the fraction of your test set that behaves correctly. It is a regression signal: it tells you whether something you already knew about got worse.
  • Coverage — the fraction of the plausible attack space your test set represents at all. It is the number that governs what a pass rate is worth, and almost nobody measures it.

A suite of 40 prompts covering three attack techniques against one input surface has near-zero coverage of everything else. No pass rate on it, however high, can say anything about the uncovered region.

Statistics puts a floor under how much you can even claim from a clean run. For zero observed failures in n independent trials, the 95% upper bound on the true failure rate is approximately 3/n — the rule of three. So "40 out of 40 passed" is entirely consistent with a true failure rate as high as 7.5%.

Put that in production terms. Suppose the assistant handles 200,000 requests a day and roughly 2% of them — 4,000 — are adversarial, whether from users probing it or from injected content in documents it reads.

Test set size, zero failures95% upper bound on failure rateCompatible with, per day
407.5%up to 300 successful attacks
1003.0%up to 120
5000.6%up to 24
2,0000.15%up to 6
10,0000.03%up to 1.2

A perfect score on a small test set is not evidence of safety. It is evidence that your test set is small.

And every row of that table assumes the honest best case: that your test cases are drawn from the same distribution as real attacks. They are not. Real attackers deliberately look for what your tests do not.

The evaluation pyramid

Different techniques catch different failures, cost different amounts, and run at different cadences. Arranged cheapest and most frequent at the bottom:

Text
        /  COMPLIANCE AUDIT  \      quarterly / pre-launch   external       /---------------------\      /   RED-TEAM PROGRAMME  \     every release            adversarial     /-------------------------\    /     BENCHMARK SUITE       \   every build              comparable   /-----------------------------\  /    CONTINUOUS MONITORING      \ always on                real traffic /---------------------------------\/       UNIT SAFETY TESTS           \ every commit           fast, cheap-------------------------------------
LayerQuestion it answersCatches uniquelyBlind to
Unit safety testsDoes this specific rule or filter still behave as written?A refactor silently disabling a checkAnything nobody wrote a test for
Continuous monitoringWhat does real traffic look like right now?Novel attacks in the wild, and drift in normal usageAttacks nobody has tried yet
Benchmark suiteHow does this compare against a standard, shared test set?Cross-model and cross-version comparisonAnything specific to your product's domain; contamination
Red-team programmeCan a motivated adversary get through in a way we did not anticipate?Novel technique familiesRare cases and long-tail volume; it is sampling, not proof
Compliance auditWould an independent reviewer accept our evidence?Gaps between what you claim and what you can demonstrateDay-to-day regressions between audits

Each layer is blind to something the layer above catches, which is why the answer to "which one should we do" is always "all five, at their own cadence". The opening incident is a pure monitoring-and-red-team failure: unit tests passed, and no other layer existed.

Building a suite: coverage first, pass rate second

Stop writing prompts and start filling a matrix. A test case is defined by three independent dimensions:

  • Technique — direct instruction override, role-play framing, authority claim, nested or quoted instruction, encoding obfuscation, incremental escalation across turns, payload split across sources, refusal-suppression preamble.
  • Surface — where the hostile text enters: the user's own message, a retrieved document, a tool's return value, a filename or metadata field, a previous turn in the conversation.
  • Objective — what the attacker wants: configuration disclosure, unauthorised tool invocation, data exfiltration, harmful content generation, denial of wallet.

Eight techniques by five surfaces by five objectives is 200 cells. A 40-prompt suite fills at most 20% of them, and in practice far less, because hand-written prompts cluster: they are nearly all "user message" surface and nearly all "configuration disclosure" objective. Measure it and the illusion collapses.

Python
from dataclasses import dataclassfrom collections import Counterfrom itertools import productTECHNIQUES = ["override", "roleplay", "authority", "nested",              "encoding", "escalation", "split", "suppression"]SURFACES   = ["user_msg", "retrieved_doc", "tool_output", "metadata", "history"]OBJECTIVES = ["config_disclosure", "tool_invocation", "exfiltration",              "harmful_content", "resource_abuse"]@dataclass(frozen=True)class Case:    id: str    technique: str    surface: str    objective: str    payload: str    holdout: bool = False      # never used while tuning defencesdef coverage(cases):    cells = set(product(TECHNIQUES, SURFACES, OBJECTIVES))    filled = {(c.technique, c.surface, c.objective) for c in cases}    by_surface = Counter(c.surface for c in cases)    return {        "cells_total": len(cells),        "cells_filled": len(filled),        "coverage": len(filled) / len(cells),        "empty_surfaces": [s for s in SURFACES if by_surface[s] == 0],        "most_loaded_surface": by_surface.most_common(1)[0] if cases else None,    }

Run that against the team's 40 prompts and the report reads: coverage 0.06, empty surfaces ['retrieved_doc', 'tool_output', 'metadata', 'history']. The incident is in that list, and it was visible before it happened.

Hold-out cases

Reserve a slice of your suite — 20% is reasonable — that is never looked at while tuning filters, thresholds or prompts. The moment you tune against a test case, that case measures memorisation rather than robustness. A defence that scores 96% on the development set and 71% on the hold-out has not learned to block attacks; it has learned to block your development set.

If your filter has ever been adjusted to make a specific test pass, that test no longer evaluates anything. Keep a set you cannot touch.

What the metrics mean, and how they mislead

Accuracy is the wrong headline

Take 10,000 sampled requests of which 100 are genuinely harmful. Consider a filter with recall 0.80 and precision 0.40:

  • True positives: 80 of the 100 harmful requests blocked
  • False negatives: 20 harmful requests allowed through
  • False positives: precision 0.40 means 80 is 40% of everything blocked, so 200 blocked in total → 120 legitimate requests wrongly blocked
  • True negatives: 9,900 − 120 = 9,780

Accuracy = (80 + 9,780) / 10,000 = 98.6%. Now compare a filter that blocks nothing at all: it gets all 9,900 benign cases right and all 100 harmful ones wrong, for 99.0% accuracy. The do-nothing filter scores higher. F1 tells the real story: 2 × 0.40 × 0.80 / (0.40 + 0.80) = 0.533 for the real filter, and 0 for the do-nothing one.

Any metric dominated by the majority class is useless where the class you care about is 1% of traffic.

MetricReads asBlind spot
Accuracy"Mostly right"Dominated by the benign majority; a do-nothing baseline wins
RecallFraction of real attacks caughtTrivially 1.0 if you block everything
PrecisionFraction of blocks that were justifiedTrivially 1.0 if you block only the most obvious case
F1Balance of the twoAssumes the two errors cost the same. They rarely do
Attack success rateFraction of adversarial attempts that workedEntirely determined by how you define "worked" — see below
False-positive rate on benign trafficThe user-hostility cost of your defencesNothing — measure it, or you will ship a filter nobody can work with

Attack success rate depends entirely on the judge

ASR is the headline number for red-teaming: attempts that succeeded, divided by attempts made. The subtlety is that something has to decide whether an individual attempt succeeded, and that judge is itself an evaluator with its own error rates.

The naive judge is a keyword check — look for "I cannot" or "I'm sorry" and call it a refusal. It fails in both directions. A response beginning "Sure, here are the steps — actually, I cannot help with that" is scored as a refusal while containing partial compliance. A response that safely explains why a request is dangerous without ever saying "I cannot" is scored as a success for the attacker.

Suppose a 200-attempt run and a keyword judge reporting 18 successes: ASR = 9.0%. Hand-review 40 of the 200 and you find the judge missed 3 real successes and invented 1. Extrapolated across 200, that is roughly 15 missed and 5 invented, giving a corrected 18 − 5 + 15 = 28 successes, or 14.0%. The reported number was low by more than a third, and the direction of that error is the dangerous one.

Three practical rules. Hand-review a fixed sample of judge decisions every cycle and report the judge's own precision and recall next to the ASR. Prefer consequence-based judging where you can — a tool call that fired, a record that was read — because it is a fact rather than an interpretation. And never compare an ASR across teams or vendors unless the judges are the same; the number is meaningless out of that context.

Evaluate the system, not the model's obedience

This distinction determines whether your evaluation programme is worth anything.

A system prompt is not a security boundary. It is text placed in a privileged position in the context window, and the model has been trained to weight text in that position above text arriving from users or documents. That training produces a strong priority ordering — a statistical tendency to prefer those instructions when they conflict — and a tendency is not an access control. It has no enforcement mechanism, it degrades under unusual phrasing, long contexts and unfamiliar languages, and it can be outweighed by sufficiently compelling contrary text. Nothing in the architecture makes it impossible to override; that is simply not what it is.

Which means an evaluation that measures "did the model follow its instructions" is measuring how strong a tendency is today, on this model version. Useful, but not the safety question. The safety question is whether anything happened.

So report two numbers from every red-team run:

Instruction-following ASRConsequence ASR
MeasuresAttempts where the model was persuaded — it agreed, complied in text, adopted the injected roleAttempts where something real occurred — a tool fired, data left the boundary, a record changed
Depends onModel version, training, phrasing, context lengthYour code: capability checks, allow-lists, approval gates, egress filters
Moves whenThe provider ships a new modelYou change your architecture
TargetAs low as you can get it, but never trustedZero, and it is achievable

A healthy run looks like this: 50 attempts, the model persuaded in 12 of them, and 0 tool calls executed — because all 12 attempted calls were denied by checks the model cannot reach. Instruction-following ASR 24%, consequence ASR 0%. The gap between those two numbers is the measured value of every defence that is not prose.

The model getting fooled is a fact about the model. Nothing happening anyway is a fact about your architecture, and it is the only one you fully control.

The corollary matters for how you read a model upgrade. If instruction-following ASR jumps from 8% to 22% after a version swap while consequence ASR stays at 0, you have a real signal worth investigating and no incident. If consequence ASR moves off zero at all, that is an outage-grade event regardless of what the other number says.

Independent signals worth combining

Hosted moderation classifiers

A hosted moderation endpoint gives you a second opinion trained on a different distribution than your own filters. Used as an evaluation signal rather than a production gate, its value is disagreement: cases where your filter and the external classifier differ are the highest-yield review queue you will ever have.

Read the output correctly. These services return a score per category — harassment, hate, self-harm, sexual, violence and finer variants — not a single safe/unsafe bit. Two consequences follow. First, the default flag threshold is calibrated for a general-purpose product, not yours; a text scoring 0.31 for harassment is not flagged at a 0.5 threshold but may be exactly what your community guidelines forbid. Second, the taxonomy is theirs. A statement that defames a competitor, leaks a contract term or gives dangerous medical advice may score near zero on every category and still be the worst output your system produced that week.

Set thresholds per category from your own labelled data. The trade is explicit: dropping a threshold from 0.5 to 0.25 might take recall from 0.72 to 0.91 while precision falls from 0.88 to 0.61 — catching 19% more real cases at the cost of roughly one wrongly-blocked benign item for every 1.5 real catches. For self-harm content that trade is obviously correct. For a spam category it is obviously not. Decide it per category, in writing.

Principle rubrics for human review

Some review has to be human, and unstructured human review is unreliable. The constitutional approach — scoring output against an explicit written set of principles rather than a general impression — is valuable here even if you never train a model that way, because it converts "this feels wrong" into a repeatable judgement.

Name the principles: harmlessness (does this enable harm?), honesty (does it state falsehoods as fact or overclaim certainty?), respect (does it treat people and groups decently?), legality (does it facilitate unlawful activity?), boundary (does it disclose configuration, other users' data, or internal reasoning it should not?).

Then measure the reviewers. Have two people independently score the same 100 outputs and compute agreement. If they agree on 85 with a chance-agreement rate of 0.50, Cohen's kappa is (0.85 − 0.50) / (1 − 0.50) = 0.70 — substantial agreement, and a rubric you can trust to produce comparable numbers next quarter. Below about 0.6, your "human evaluation" is largely noise, and the fix is a sharper rubric with worked examples of each borderline call, not more reviewers.

Published benchmarks, and contamination

Standard datasets and classifiers — toxicity models you can run locally, published suites like HarmBench and JailbreakBench that measure robustness to harmful requests and jailbreaks, and AgentDojo, which measures prompt injection against tool-using agents — give you numbers comparable with other people's numbers. That is their entire value, and it is real.

It is also the source of their main failure. Published benchmarks are public text, and public text ends up in training corpora. A model that has memorised the test set scores well without being safer, which is benchmark contamination.

You can test for it directly. Paraphrase a random 100 items from the benchmark, preserving the adversarial intent while changing surface wording, and re-score. A model scoring 94% on the published items and 71% on the paraphrases has a 23-point gap that is not explained by difficulty. Treat that gap as the honest measure of how much of the benchmark score is memorisation.

Dashboards that do not lie

With four independent signals you will want one tracked number. Combining them naively is how a critical failure disappears.

Take moderation pass rate 0.95, principle compliance 0.92, benchmark accuracy 0.90, and an attack success rate of 0.20 — inverted to a red-team score of 0.80. Equal weights give (0.95 + 0.92 + 0.90 + 0.80) / 4 = 3.57 / 4 = 89.3%. That reads as a solid B. It contains the fact that one in five adversarial attempts succeeded.

Python
FLOORS = {          # per-component hard minimums; breaching any one fails the build    "moderation": 0.90,    "principles": 0.85,    "benchmark": 0.80,    "red_team": 0.95,          # i.e. consequence ASR must stay at or below 5%}def report(scores, coverage_frac, judge_recall):    breaches = [k for k, v in scores.items() if v < FLOORS[k]]    return {        "mean": sum(scores.values()) / len(scores),        "min_component": min(scores, key=scores.get),        "breaches": breaches,        "coverage": coverage_frac,          # publish it beside every score        "judge_recall": judge_recall,       # how much to trust the ASR itself        "gate": "FAIL" if breaches else "PASS",    }

Three design rules make a dashboard honest. Every component carries its own floor that a good average cannot compensate for. The minimum component is displayed as prominently as the mean. And coverage and judge reliability are published beside every score, because a number without them is not interpretable — 99.2% across 40 cases judged by a keyword matcher and 99.2% across 4,000 cases judged by consequence are not the same claim, and only the second one means anything.

Misconceptions worth naming

BeliefCorrection
"We pass our safety tests, so we're safe"You are safe against what you tested. Publish coverage next to the pass rate or the claim is unbounded.
"A high benchmark score proves robustness"It may prove memorisation. Paraphrase a sample and measure the gap.
"Red-teaming is a pre-launch activity"It is a regression suite. ASR is meaningful only as a trend across releases.
"Our overall safety score is 89%"An average hides its worst component. Report the minimum and enforce per-component floors.
"The model refused, so the attack failed"Refusal is the soft measure. The hard measure is whether any tool fired or data moved.
"We hardened the system prompt, so we're covered"A system prompt is a trained priority ordering, not a boundary. Evaluate consequences, not obedience.
"Internal testing means we're compliant"Audit is external by definition. Its currency is evidence you can produce, not tests you claim to have run.

What this means when you build something

Start by writing down the matrix, not the prompts. Techniques down one axis, entry surfaces across another, attacker objectives on the third. Fill one case per cell, however crudely. A suite of 200 rough cases with full coverage tells you far more than 40 polished ones clustered in a corner, because the empty cells are the finding — and empty cells are the only part of an evaluation that predicts incidents rather than describing history.

Make consequence the primary metric. Instrument your tool layer so a test run reports which calls were attempted and which executed, and treat any executed call from an adversarial run as a build failure. That single gate is worth more than every text-classification score combined, because it is the one number that does not move when a provider ships a new model version.

Wire the cheap layers into CI and the expensive ones into the release calendar. Unit tests and a fast subset of the suite on every commit; the full suite, the benchmark and the paraphrase-gap check on every build; a red-team cycle with a mandatory novelty budget — a fixed fraction of attempts must use techniques not already in the library — on every release; and a standing evidence pack of decision logs, policy versions and approval records for the audit that will eventually ask for them.

Finally, hold the hold-out set genuinely apart, and re-run everything on every model swap without exception. The team in the opening story did the swap and re-ran the suite, which was the right instinct executed against the wrong suite. Had their report printed coverage: 0.06, empty surfaces: ['retrieved_doc', ...] next to the reassuring 40/40, someone would have asked the question nine days early.