Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Controls That Map to Threats
The CISO's team had seen many AI security designs. They had a standard first question, and they asked it about every threat on Meridian's list: "What stops this if the model does exactly what the attacker says?" Answers that began with "the system prompt tells the model to..." were marked as not a control.
That question is the right one, because there is currently no complete fix for prompt injection. Language models cannot reliably separate instructions from the data they are reading; to the model, it is all text. Guards and careful prompts reduce the rate of success, but a determined attacker, or plain bad luck, will sometimes get through. A design that is safe only if the model always resists is not safe.
This lesson chooses controls for Meridian's threats that hold even when the model is fully turned, ranks controls by strength, and states honestly what risk remains. It produces part B of MER-10.
Assume the model can be turned
Treat the model's output, for any request that included untrusted text, as if the attacker wrote it. Then ask what that output can reach.
At Meridian the answer is short, and the design made it short on purpose. Model output can become text shown to a staff member, or text inside a draft letter. It cannot choose a customer, because case binding is in code. It cannot send anything, because no send tool exists. It cannot change a figure, because figures come from the calculator and are checked. It cannot make the browser fetch anything, because of the output controls below. The worst outcome of a fully successful injection is misleading text in front of a trained person, with checks and review between it and any customer.
Controls have different strengths
Not all controls are equal. Safety engineering has long ranked controls, and the same ranking works for AI.
| Strength | Type | Meridian example |
|---|---|---|
| Strongest | Eliminate the capability | No send, waive or update tools exist |
| Constrain with code | Customer ID bound from the case; template allow-list; role filter before retrieval | |
| Detect with code | Figure checks, citation checks, link removal, record ownership check | |
| Detect with a model | Guard classifier flags instruction-like text in notes | |
| Human review | Letter review with highlights and drills | |
| Weakest | Instructions and policy | System prompt rules; staff training |
Every threat should have at least one control from the top three rows. The lower rows still matter, as extra layers and as sources of signal, but none of them should be the only thing between an attacker and harm.
Model-based guards deserve a fair description. Meridian's guard classifier catches 95% of the known attack set, flags 0.4% of normal traffic by mistake, and runs in under 150 milliseconds. That is useful: it stops unsophisticated attempts and, just as valuable, it counts them, so the security team can see attack trends. But attackers adapt to guards, and a guard trained on last year's attacks misses next year's. It is a layer, not a wall. When it blocks honest text, such as an angry customer email pasted by a staff member, the panel asks the staff member to summarise the email in their own words, which is a small cost.
Output handling in code
T-02, exfiltration through rendering, is closed mostly by code that treats model output as untrusted before the panel shows it.
1import re2from urllib.parse import urlparse34ALLOWED_HOSTS = {"policyhub.meridian.internal", "atlas.meridian.internal"}5IMAGE = re.compile(r"!\[[^\]]*\]\([^)]*\)")6LINK = re.compile(r"\[([^\]]+)\]\(([^)\s]+)\)")7BARE_URL = re.compile(r"https?://[^\s)]+")89def safe_output(text: str) -> tuple[str, list[str]]:10 """Strip images and links to hosts outside the allow-list. Report what was removed."""11 findings: list[str] = []12 if IMAGE.search(text):13 findings.append("image_removed")14 text = IMAGE.sub("", text)1516 def keep_link(match: re.Match) -> str:17 host = urlparse(match.group(2)).hostname or ""18 if host in ALLOWED_HOSTS:19 return match.group(0)20 findings.append(f"link_removed:{host}")21 return match.group(1) # keep the label, drop the address2223 text = LINK.sub(keep_link, text)24 for url in BARE_URL.findall(text):25 if (urlparse(url).hostname or "") not in ALLOWED_HOSTS:26 findings.append("bare_url_removed")27 text = text.replace(url, "[link removed]")28 return text, findingssafe_output removes every image, keeps links only to the bank's own policy and case systems, and replaces any other address with a marker. It returns what it removed as findings. Those findings are a detective control as well: a legitimate policy answer never contains an outside address, so any removal is logged as a likely injection attempt and counted on the security dashboard. Behind this code, the panel's content security policy also blocks the browser from fetching anything outside the bank's domains, so even a bug in this function would not reopen the path. Two independent controls, both code.
The controls for Meridian's top threats
T-01 Injection through notes
- Preventive: untrusted labelling; no effectful tools reachable; figures from calculator
- Detective: claims check links agreements to records; guard flags instruction-like text
- Corrective: review, kill switch, add case to adversarial set
T-04 Restricted policy disclosure
- Preventive: role filter before ranking; restricted policies in a separate index
- Detective: access tests per role on every release
- Corrective: remove from index, review audit log for who asked what
The claims check in T-01 needs a sentence of explanation. When a summary or letter states that the customer agreed something, such as a promise to pay, code looks for a matching structured record or a staff-authored note, not a customer-authored one. If none exists, the statement is marked "unverified: from customer correspondence" rather than stated as fact.
Residual risk, stated honestly
After controls, each threat is re-rated. Some drop a long way. T-01 falls from 9 to 3: injection attempts still happen, so likelihood stays at 3, but the impact drops to 1 because the model cannot act, figures are fixed and agreements are checked. What remains is real, though: injected text can still produce a misleading sentence in a summary shown to staff.
The architect's job is to state that plainly, with the controls that limit it, and hand it to the right person. Residual risk is accepted by the business owner, the Head of Servicing, as the decision-rights table in section 1 set out. It is not accepted by the architect and not hidden in an appendix.
| Threat | Preventive | Detective | Corrective | Test | Owner | Residual |
|---|---|---|---|---|---|---|
| T-01 Injection | Labels; no effectful tools; calculator figures | Claims check; guard | Kill switch; review | 150 adversarial cases per release | Service owner | 3, accepted |
| T-02 Exfiltration | Output filter; content security policy | Removal findings alert | Block and investigate | Exfiltration cases in CI | Security | 1 |
| T-03 Wrong customer | Case binding in code | Record ownership check | SEV1 runbook | Cross-case tests | Eng lead | 1 |
| T-04 Restricted policy | Role pre-filter; separate index | Role access tests | Remove, audit | Tests per role | Platform team | 1 |
| T-05 Insider lookup | Delegated token; queue rules | Volume anomaly alerts | Existing HR process | Access review quarterly | Ops risk | 2 |
| T-06 Poisoned policy | Four-eyes publishing | Scan for hidden or instruction-like text | Quarantine chunk | Seeded test document | Policy team | 1 |
| T-07 Consumption | 20 requests a minute and 400 a day per user; gateway budgets | Spend over 150% of forecast | Throttle | Load test | Platform team | 1 |
| T-10 Supply chain | Pinned versions; verified hashes; safe weight formats | Dependency scanning | Roll back bundle | Build pipeline checks | Platform team | 2 |
Check your understanding
0 of 3 answered
1.A proposed control for T-01 is "the system prompt tells the model to ignore instructions in case notes." How should the security review treat it?
2.Why is removing outside links in safe_output both a preventive and a detective control?
3.After controls, T-01's residual score is 3: injected text can still produce a misleading sentence shown to staff. Who accepts that risk?