Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
A layered guardrail architecture, and the road ahead
Every section of this course added a defence: a schema, a country filter, a grounding check, a confirmation card, a threshold on the sensitive route, a canary, a redactor. Each one, measured alone, lets something through. The NLI check misses some paraphrased errors. The redactor misses 9% of names. The injection filter at ingest misses instructions phrased as policy. If PolicyPal depended on any single one, it would fail regularly.
Safety engineers describe this with the Swiss cheese model: every layer has holes, but the holes are in different places, so an accident needs the holes in every layer to line up. A problem that gets past 90% of cases at each of three independent layers gets through only 0.1% of the time, if the layers really are independent. The design work is making sure they are.
This final lesson assembles PolicyPal's layers into one architecture, puts a cost and a failure behaviour on each, and then looks at what comes after launch.
The layers
| Layer | What it does | Stops | Added latency |
|---|---|---|---|
| 1. Identity and access | SSO; identity from the session; country and audience filters in retrieval | Other people's data, restricted documents | about 0 ms |
| 2. Input checks | Length limit, tag stripping, redaction, per-user rate limit, injection score | Oversized and abusive input, stored PII | about 20 ms |
| 3. Routing | Sensitive topics to people; out-of-scope to a fixed reply | Automated answers where harm is likely | 15 ms |
| 4. Constrained model | Versioned prompt, schema, allow-listed tools with enums, budgets | Wrong shapes, invented actions, runaway loops | none extra |
| 5. Actions | Confirmation card; at most 5 tickets per user per day | Unwanted or spammed side effects | a human click |
| 6. Output checks | Schema and citation validation, NLI grounding, PII scan, link allow-list, canary | Unsupported claims, leaks, exfiltration | about 120 ms |
| 7. People and monitoring | Review queue, red-team suite per release, drift and cost alerts | Whatever the other six miss, found later | none in the request |
The input and output layers are the ones built in this section, and they are short.
1# policypal/guards.py2from dataclasses import dataclass, field34from policypal.grounding import check5from policypal.pii import redact6from policypal.redteam.run import ALLOWED_LINK, CANARY, URL78@dataclass9class Guarded:10 ok: bool11 text: str12 reasons: list[str] = field(default_factory=list)13 tools_allowed: bool = True1415def guard_input(text: str, user, limiter, injection_score) -> Guarded:16 if len(text) > 4000:17 return Guarded(False, "", ["too long"])18 if not limiter.allow(user.employee_id):19 return Guarded(False, "", ["rate limited"])20 risky = injection_score(text) > 0.821 return Guarded(True, redact(text), ["injection suspected"] if risky else [],22 tools_allowed=not risky)2324def guard_output(answer: str, sources: list[dict]) -> Guarded:25 if CANARY in answer:26 return Guarded(False, "", ["canary leaked"])27 reasons = []28 for url in URL.findall(answer):29 if not ALLOWED_LINK.match(url):30 answer = answer.replace(url, "[link removed]")31 reasons.append("external link removed")32 unsupported = check(answer, sources) # the NLI check from Section 333 reasons += [f["problem"] for f in unsupported]34 return Guarded(not unsupported, redact(answer), reasons)Look at how guard_input uses the injection score. Injection classifiers have false positives: a policy question that quotes "ignore the previous version of this form" can look like an attack. Blocking those questions would annoy honest users every day. So a high score does not block. It reduces privilege: the turn runs with no tools, so a successful injection could at most produce a wrong answer, which the output checks and citations then constrain. This is a pattern worth copying: when you are unsure, shrink what the model can do rather than refusing the user.
In guard_output, a failed grounding check is not the end; it hands back to the drop, regenerate or hand-over policy from Section 3. The canary is the one hard block, because a leaked prompt has no partial fix.
Fail open or fail closed?
Each guard is itself software that can break. Decide in advance what happens when it does.
| If this guard is down | Behaviour | Why |
|---|---|---|
| Identity provider | Fail closed: no answers | Without identity, every other layer is meaningless |
| Redactor | Answer, but store nothing | Users still get help; nothing unredacted is kept |
| Injection scorer | Treat every turn as risky: no tools | Answers continue with less privilege |
| NLI grounding check | Answer with a "sources not verified" banner, and alert | Citations remain; the banner keeps trust honest |
| Router | Prompted fallback router (Section 7) | Slower, same behaviour |
Failing closed everywhere would make PolicyPal fragile; failing open everywhere would make it unsafe. The table is a set of explicit decisions, tested in the same way as the degraded modes in Section 7.
The road ahead
Launch is the start of the work, not the end. The habits that keep PolicyPal good are the ones built in this course, run on a schedule.
- Every change — the eval suite and flip list, the red-team suite, and a release manifest.
- Every week — the review queue, with every confirmed failure turned into an eval case.
- Every month — judge calibration on 30 human-labelled cases, cost and drift review, router health checks.
- Every quarter — refresh a tenth of the eval set from new traffic, re-run the model pilot against newer models, and review the privacy map.
The models will keep changing, and some of your choices should change with them. Longer context windows and cheaper tokens will make some chunking and budget decisions less tight, but they do not remove the reasons for retrieval: permissions, citations, fresh documents and cost at scale. Better tool use will make agents more reliable, which may move some workflows up the autonomy ladder, but only when your eval says so. What does not change is the method: measure on your own task, prefer construction to persuasion, and let failures you can see drive the next piece of work.
If you want to go deeper next, four directions follow naturally from this course. Tool protocols such as the Model Context Protocol, for connecting assistants to many internal systems in a standard way. Serving open models with engines such as vLLM, when a self-hosted component like the router grows. Advanced evaluation, including simulated users for multi-turn conversations. And AI governance, the risk registers and review processes that larger organisations need around systems like this one.
Check your understanding
0 of 3 answered
1.Three independent layers each stop 90% of a certain attack. Roughly how often does the attack get through all three, and what is the catch?
2.When the injection score is high, why does PolicyPal run the turn with tools disabled instead of refusing to answer?
3.The NLI grounding service is down. What does PolicyPal do, and why?