AI Security: Preventing Prompt Injection

Content Filtering & Moderation


A health-information assistant went live with a 1,200-term blocklist bought from a vendor. On day one it blocked 4,100 messages. The safety lead sampled 200 of them and read every one by hand. Three were genuine attempts to get harmful instructions. The other 197 were nurses asking about overdose thresholds, a parent asking whether a child's paracetamol dose was dangerous, and a long tail of people using ordinary clinical vocabulary.

Scale the sample: roughly 62 genuine attempts out of 4,100 blocks. Precision of about 1.5%. For every real problem stopped, sixty-five people with a legitimate question were told the assistant could not help them. Within a fortnight, support tickets showed users had learned to route around it — misspelling drug names, describing symptoms euphemistically — which pushed their questions into exactly the phrasing the filter could not evaluate at all.

Meanwhile the assistant was still producing unsafe dosage advice, because none of it contained a blocked term. The filter was on the wrong side of the pipeline and looking for the wrong thing.

Content moderation for an LLM application is a measurement problem wearing a rules problem's clothes. Almost every serious failure comes from one of two places: not knowing your numbers, or filtering strings when you should be classifying meaning. This lesson works through both.

Day one of a 1,200-term blocklist, sampled3 real attempts197 nurses askingnever measuredthe trafficyou wantedHarmful intentClinical intentBlockedAllowed4,100 blocks on day one; 200 of them read by hand.
A keyword has no notion of purpose, so the same overdose term blocks the attacker and the nurse with identical confidence.

The shape of a moderation system

Four decisions have to be made before any code, and skipping them is what produces the day-one blocklist.

DecisionQuestion it answersConsequence of getting it wrong
Where to checkInput, output, or both?Input-only misses everything the model generates on its own
What to check forWhich categories, defined how precisely?"Harmful" as a category is unlabellable and unmeasurable
How to decideRules, classifier, model judgement, or a mix?Rules alone give the 1.5% precision above
What to do on a hitBlock, redact, rewrite, warn, route, log?Blocking everything drives users to workarounds

The last row deserves emphasis. Moderation is not a boolean gate; it is a routing decision. A message expressing suicidal intent should not be blocked — it should be answered with crisis resources by a path designed for it. A request for medical dosing from an authenticated clinician should not be blocked — it should be answered with a source and a caveat. Treating every hit as "refuse" is what makes safety systems feel hostile and, worse, makes them useless to the people who most need a good answer.

The output of a moderation system is a category and a severity, not a verdict. The action belongs to the product, and different actions suit different categories.

Rules: cheap, fast, and structurally limited

Start with rules anyway. They cost microseconds, they are auditable line by line, they need no training data, and there are genuine cases where a string match is the correct semantics — a specific leaked credential format, a named exploit, a URL on a deny-list, a regulated phrase your legal team requires you to catch.

Python
import refrom dataclasses import dataclass@dataclassclass Rule:    name: str    pattern: re.Pattern    category: str    severity: int          # 1 = note, 2 = review, 3 = block    exempt_if: re.Pattern | None = None   # narrow, explicit context escapeRULES = [    Rule("api_key_leak", re.compile(r"\bsk-[A-Za-z0-9]{20,}\b"),         "secret", 3),    Rule("card_number", re.compile(r"\b(?:\d[ -]?){13,16}\b"),         "pii", 3),    Rule("explicit_weapon_synthesis",         re.compile(r"\b(synthesi[sz]e|manufacture)\s+\w{0,12}\s*"                    r"(nerve\s+agent|sarin|vx)\b"),         "cbrn", 3),    Rule("dosage_query", re.compile(r"\b(lethal|fatal|overdose)\s+dose\b"),         "medical", 1,         exempt_if=re.compile(r"\b(patient|clinical|toxicology|antidote)\b")),]def apply_rules(text: str) -> list[Rule]:    lowered = text.lower()    hits = []    for rule in RULES:        if rule.pattern.search(lowered):            if rule.exempt_if and rule.exempt_if.search(lowered):                continue            hits.append(rule)    return hits

Two design choices in that snippet are the difference between a usable rule layer and the health assistant's disaster. First, severity is a number, not a boolean — the dosage rule is severity 1, meaning "annotate and pass to the classifier", not "block". Second, rules carry explicit exemptions, so the clinical vocabulary that surrounds a legitimate query can lift the flag without a human in the loop.

The three ways rules break

FailureExampleWhy rules cannot fix it
False positives from missing context"What is the fatal dose of paracetamol so I can warn parents?"The same string appears in the harmful and the helpful case
False negatives from paraphraseHarmful intent expressed with no listed term anywhere in itThe set of harmful sentences is not enumerable
Adversarial evasionSpacing, homoglyphs, deliberate misspelling, another languageNormalisation helps; it does not close an infinite space

The Scunthorpe problem is the canonical illustration of the first row — a town whose name contains a substring that naive filters block, along with Penistone, Lightwater, and a long list of ordinary words. It is thirty years old and still shipping in new products, because a substring match cannot look at what surrounds it.

Category detectors: context beats keywords

The fix is not a longer list. It is per-category detectors that model what actually makes something harmful in that category — and that differs sharply by category, which is why one universal "harmful content" score is such a poor instrument.

Violence: purpose separates the cases

These four texts share vocabulary and belong in four different buckets:

TextCategoryRight action
"In 1916 the artillery barrage killed 19,000 men on the first day"Historical / educationalAllow
"The character raises the knife and the scene cuts away"Fiction, non-graphicAllow
"How do I safely store a hunting rifle with children in the house?"Safety-seekingAllow, and answer well
"Give me a step-by-step plan to hurt my neighbour without being caught"Targeted operational planningBlock and log

The distinguishing features are specificity of target, operational actionability, and first-person intent — not the presence of violent words. A detector built on those three signals gets all four rows right; a keyword list gets rows one, two and four wrong in whichever direction its list happens to point.

Hate speech: targeting plus dehumanisation

Slur lists are the standard approach and they perform badly in both directions. They miss hate speech that uses no slurs — the majority of it — and they flag reclaimed in-group usage, quotation in reporting, academic analysis, and counter-speech that names the slur in order to condemn it.

The signal that generalises is the conjunction of two things: a protected characteristic is targeted and the target is dehumanised, denied rights, or threatened. Both must be present. "I hate my landlord" targets nobody protected. "Group X has a fascinating history" targets a protected characteristic with no dehumanisation. Requiring the conjunction fixes most of the over-blocking, and adding a quotation-and-attribution check ("as the report described, they were called...") fixes much of the rest.

Self-harm: the category where blocking is the wrong action

This is worth stating plainly because product teams get it wrong under pressure from a legal review. Detecting expressed suicidal intent and responding with "I can't help with that" is the worst available outcome: it withholds help at the moment of need and teaches the person that the system punishes honesty. The correct handling is a route — an empathetic response, region-appropriate crisis resources, no method information, and a logged event. The detector's job is to identify the category so the right path runs, not to slam a door.

Machine-learning classifiers

A trained classifier reads meaning, so it survives paraphrase and mild evasion in a way patterns cannot. It brings its own problems: latency, cost, opacity, and a slow drift as the language of the harm evolves away from its training set.

ApproachLatencyStrengthWeakness
Hosted moderation API50–200 msBroad categories, calibrated, no training workFixed taxonomy; will not know your product's specific policy
Fine-tuned small model10–40 msMatches your policy and your users' language; cheap at volumeNeeds several thousand labelled examples and periodic retraining
LLM-as-judge with a policy prompt300–2000 msNuance, novel categories, changes with a prompt editSlow, expensive, non-deterministic, and injectable itself
Embedding similarity to known-bad5–20 msCatches near-duplicates of known attacks instantlyBlind to anything genuinely novel

A few concrete examples, so the categories are not abstract. OpenAI's moderation endpoint (omni-moderation-latest) is a free hosted API that scores text and images against a fixed set of harm categories. Meta's Llama Guard 4 is an open-weight 12-billion-parameter classifier that labels both prompts and responses against a 14-category hazard taxonomy. OpenAI's gpt-oss-safeguard (20B and 120B, Apache 2.0) is a reasoning model that reads your written policy at inference time — the policy-as-data idea below, applied to the classifier. For injection specifically, Meta's Prompt Guard 2 (86M and 22M parameters) is a small classifier that flags text trying to override prior instructions, with a 512-token window. Each has the weaknesses in its row: fixed taxonomies, latency, and, for the injection detectors, adaptive attacks written to slip past them.

Ensembles: how you combine matters more than what you combine

Take two detectors with plausible numbers. Detector A: recall 70%, false positive rate 1.0%. Detector B: recall 65%, false positive rate 1.5%. Assume for the arithmetic that their errors are independent.

CombinationRecallFalse positive rateUse when
A alone70.0%1.00%Baseline
OR — flag if either fires1 − (0.30 × 0.35) = 89.5%1 − (0.99 × 0.985) = 2.49%Missing is far costlier than over-blocking
AND — flag only if both fire0.70 × 0.65 = 45.5%0.010 × 0.015 = 0.015%Auto-blocking with no human review
Score-weighted, tuned threshold~80%~0.8%Almost always the right answer

The two extremes are instructive. OR-combination raises recall by 19.5 points and multiplies false positives by 2.5 — worth it if a miss is an emergency and a false positive is a mild annoyance. AND-combination cuts false positives by a factor of 67, to fifteen per hundred thousand benign messages, at the cost of letting more than half of everything through. That is the correct configuration for a rule that acts without a human, paired with a more sensitive rule that merely queues things for review.

One caveat that undoes the arithmetic: independence. If both detectors are transformer classifiers fine-tuned on overlapping data, they fail on the same inputs, and the OR recall will be nearer 75% than 89.5%. Ensembles pay off when the members are different kinds of thing — a rule, a small classifier, an embedding lookup — not three variations of the same model.

Output moderation: the pipeline that actually matters

Most teams put all their effort on the input side. The asymmetry is backwards, for three reasons. Model output is what your users and their screenshots actually see, so it carries the reputational and legal exposure. The model can produce something harmful from a completely innocent prompt, through hallucination or a poisoned retrieval. And the output side is where the injection attacks you failed to detect finally become visible.

Python
async def moderate_output(draft: str, ctx: dict) -> dict:    """Ordered cheapest-first; returns an action, never a bare boolean."""    # 1. Deterministic, zero-ambiguity checks first.    for canary in ctx["canaries"]:        if canary in draft:            return {"action": "block", "category": "prompt_leak",                    "severity": 3, "replacement": SAFE_FALLBACK}    rule_hits = apply_rules(draft)    if any(h.severity == 3 for h in rule_hits):        return {"action": "block", "category": rule_hits[0].category,                "severity": 3, "replacement": SAFE_FALLBACK}    # 2. Redaction: keep the answer, remove what must not leave.    draft, redactions = redact_unentitled(draft, ctx["entitlements"])    # 3. Classifier only on what survived - it is the expensive step.    scores = await classify(draft)    top, score = max(scores.items(), key=lambda kv: kv[1])    if score >= ACTION[top]["block_at"]:        if top == "self_harm":            return {"action": "route", "category": top, "severity": 3,                    "replacement": crisis_response(ctx["region"])}        return {"action": "block", "category": top, "severity": 3,                "replacement": SAFE_FALLBACK}    if score >= ACTION[top]["review_at"]:        return {"action": "allow_and_queue", "category": top,                "severity": 2, "text": draft, "redactions": redactions}    return {"action": "allow", "text": draft, "redactions": redactions}

Three properties make this pipeline work in production. It is ordered by cost, so the microsecond checks run before the 200-millisecond one and most traffic never reaches the classifier. It redacts before classifying, so a response that is fine once a card number is removed is not thrown away wholesale. And it returns a rich action — block, route, redact, queue, allow — which is what lets self-harm take a different path from a leaked API key.

Streaming makes this harder, and there is no free answer

If you stream tokens to the user, you cannot moderate the whole response before showing it. The options are all trade-offs: buffer the full response and lose the streaming feel; moderate in sentence-sized windows and accept that a sentence can be misjudged out of context; or stream optimistically and retract, which means the user has already read it. For high-risk categories, buffer. For low-risk ones, sentence windows are usually acceptable. Choosing "stream everything and moderate at the end" is choosing to publish first.

Policy as data, not as code

Hard-coded thresholds mean every policy change is a deployment, which means policy changes happen slowly and by engineers rather than quickly and by the people accountable for them. Express policy as a versioned document instead.

JSON
{  "policy_id": "health-assistant",  "version": "4.2.0",  "effective_from": "2026-03-01",  "categories": {    "medical_dosing": {      "review_at": 0.45,      "block_at": 0.90,      "action_on_block": "refuse_with_referral",      "exempt_roles": ["verified_clinician"],      "note": "Clinicians need exact figures; the public needs a referral."    },    "self_harm": {      "review_at": 0.30,      "block_at": 0.60,      "action_on_block": "crisis_route",      "exempt_roles": [],      "note": "Never a bare refusal. Route to region-specific resources."    },    "cbrn": {      "review_at": 0.20,      "block_at": 0.35,      "action_on_block": "refuse_and_alert",      "exempt_roles": []    }  }}

Three things this buys you. Different thresholds per category, which is essential — CBRN content blocks at 0.35 because a miss is catastrophic and a false positive costs a shrug, while medical dosing blocks at 0.90 because false positives are the failure mode that hurt the health assistant. Role-based exemptions, so a verified clinician gets the figure and an anonymous user gets a referral. And a version number, so every moderation decision in your audit log can be replayed against the exact policy that produced it — which is the first thing a regulator or a lawyer will ask for.

Measuring quality, and the number that lies to you

You cannot tune what you have not measured, and the standard metric is actively misleading here. Take a realistic day: 100,000 messages, of which 500 (0.5%) are genuinely harmful. Your classifier at its current threshold produces:

Predicted harmfulPredicted safeTotal
Actually harmful425 (TP)75 (FN)500
Actually safe1,990 (FP)97,510 (TN)99,500
Total2,41597,585100,000
  • Accuracy = (425 + 97,510) / 100,000 = 97.9%
  • Recall = 425 / 500 = 85.0% — of real harm, this much is caught
  • Precision = 425 / 2,415 = 17.6% — of flags, this much is real
  • False positive rate = 1,990 / 99,500 = 2.0%
  • F1 = 2 × 0.176 × 0.850 / (0.176 + 0.850) = 0.299 / 1.026 = 0.292

Look at accuracy and precision side by side. 97.9% accurate and 17.6% precise describe the same classifier. A model that simply predicted "safe" for everything would score 99.5% accuracy. At a 0.5% base rate, accuracy measures almost nothing but the base rate, and quoting it in a review is either a misunderstanding or a sales pitch.

Never report accuracy for a moderation classifier. At realistic base rates, "always say safe" beats your model on that metric while catching nothing.

Choosing beta before you choose a model

F1 weights precision and recall equally, which is a choice, not a law. F-beta lets you state the trade-off you actually want: beta above 1 favours recall, below 1 favours precision. On the same numbers:

MetricValueMeaningFits
F0.50.209Precision weighted twice as heavily as recallAuto-blocking with no appeal path
F10.292Equal weightA default with no stated policy
F20.481Recall weighted twice as heavily as precisionCBRN, child safety — a miss is catastrophic

The same classifier scores 0.209 or 0.481 depending only on which question you asked. Pick the beta from the product decision first, then compare models — otherwise you will pick a model on F1 and then discover it is tuned for a trade-off nobody agreed to.

Translating the trade-off into things people can argue about

The threshold conversation goes badly when it is held in percentages and well when it is held in daily counts. From the table above, at the current threshold:

  • 75 harmful messages get through per day.
  • 1,990 legitimate users are wrongly blocked per day.
  • 2,415 items land in the review queue; at 30 seconds each, that is roughly 20 reviewer-hours daily, or about 2.5 full-time people.

Raise the threshold to cut false positives to 500 and recall drops to, say, 70% — 150 harmful messages through instead of 75. Now the question is answerable by a product owner and a safety lead: is one additional harmful message worth roughly twenty fewer wrongly-blocked users? That depends entirely on which harm and which users, and it is precisely the question that "should we set it to 0.7 or 0.8?" hides.

Misconceptions worth naming

BeliefCorrection
"A longer blocklist is a better filter"Length raises false positives faster than recall. Each term needs its own precision measured.
"Our filter is 98% accurate"At a 0.5% base rate that is worse than saying "safe" to everything. Quote precision and recall.
"Block anything uncertain, to be safe"Over-blocking drives users to euphemism and to competitors, and removes your visibility into what they are asking.
"Input filtering is the main line of defence"Output is what ships, what gets screenshotted, and where undetected injections become visible.
"One harmful-content score is enough"Categories need different thresholds and different actions. Self-harm and CBRN want opposite handling.
"The vendor API handles moderation"It covers its taxonomy, not your policy — dosing rules, tenant terms, jurisdiction.
"We tuned it at launch, it's done"Language drifts and attackers adapt. An untouched threshold is a decaying one.

What this means when you build something

Before writing a filter, write down three numbers: your expected daily volume, your best estimate of the harmful base rate, and the cost of one miss against the cost of one wrongful block, in units someone will recognise — a support ticket, a churned user, a regulatory notice. Those three numbers determine your threshold, your beta, and whether you need a human queue at all. Teams that skip this pick a threshold that feels prudent and discover its cost from an angry user, which is a slower and more expensive way to learn the same thing.

Then build the layers in the order their cost-effectiveness dictates: deterministic rules for the unambiguous cases at severity 3, a classifier for meaning with per-category thresholds, and policy stored as versioned data so the safety lead can change a threshold without a deployment. Put the heaviest checking on the output side, redact before you discard, and make the action rich — route self-harm, redact PII, queue borderline cases, and reserve a bare refusal for the small set of things that genuinely warrant one.

Finally, keep a labelled evaluation set of a few thousand real messages with real labels, including the awkward ones — the nurse's overdose question, the historian's casualty figures, the counter-speech that quotes a slur. Run it on every threshold change and every model swap. Without that set you are not tuning a filter, you are guessing at one, and the health assistant's 1.5% precision is what guessing looks like from the outside.