Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

A layered guardrail architecture, and the road ahead


Every section of this course added a defence: a schema, a country filter, a grounding check, a confirmation card, a threshold on the sensitive route, a canary, a redactor. Each one, measured alone, lets something through. The NLI check misses some paraphrased errors. The redactor misses 9% of names. The injection filter at ingest misses instructions phrased as policy. If PolicyPal depended on any single one, it would fail regularly.

Safety engineers describe this with the Swiss cheese model: every layer has holes, but the holes are in different places, so an accident needs the holes in every layer to line up. A problem that gets past 90% of cases at each of three independent layers gets through only 0.1% of the time, if the layers really are independent. The design work is making sure they are.

This final lesson assembles PolicyPal's layers into one architecture, puts a cost and a failure behaviour on each, and then looks at what comes after launch.

PolicyPal's defences, in orderIdentity, access filters, input checksRouting: sensitive topics to peopleConstrained model, confirmed actionsOutput: grounding, links, canaryPeople: review queue, red-team suite
Every layer has holes; different mechanisms keep the holes from lining up, and each layer's failure mode is decided in advance.

The layers

LayerWhat it doesStopsAdded latency
1. Identity and accessSSO; identity from the session; country and audience filters in retrievalOther people's data, restricted documentsabout 0 ms
2. Input checksLength limit, tag stripping, redaction, per-user rate limit, injection scoreOversized and abusive input, stored PIIabout 20 ms
3. RoutingSensitive topics to people; out-of-scope to a fixed replyAutomated answers where harm is likely15 ms
4. Constrained modelVersioned prompt, schema, allow-listed tools with enums, budgetsWrong shapes, invented actions, runaway loopsnone extra
5. ActionsConfirmation card; at most 5 tickets per user per dayUnwanted or spammed side effectsa human click
6. Output checksSchema and citation validation, NLI grounding, PII scan, link allow-list, canaryUnsupported claims, leaks, exfiltrationabout 120 ms
7. People and monitoringReview queue, red-team suite per release, drift and cost alertsWhatever the other six miss, found laternone in the request

The input and output layers are the ones built in this section, and they are short.

Python
# policypal/guards.pyfrom dataclasses import dataclass, fieldfrom policypal.grounding import checkfrom policypal.pii import redactfrom policypal.redteam.run import ALLOWED_LINK, CANARY, URL@dataclassclass Guarded:    ok: bool    text: str    reasons: list[str] = field(default_factory=list)    tools_allowed: bool = Truedef guard_input(text: str, user, limiter, injection_score) -> Guarded:    if len(text) > 4000:        return Guarded(False, "", ["too long"])    if not limiter.allow(user.employee_id):        return Guarded(False, "", ["rate limited"])    risky = injection_score(text) > 0.8    return Guarded(True, redact(text), ["injection suspected"] if risky else [],                   tools_allowed=not risky)def guard_output(answer: str, sources: list[dict]) -> Guarded:    if CANARY in answer:        return Guarded(False, "", ["canary leaked"])    reasons = []    for url in URL.findall(answer):        if not ALLOWED_LINK.match(url):            answer = answer.replace(url, "[link removed]")            reasons.append("external link removed")    unsupported = check(answer, sources)             # the NLI check from Section 3    reasons += [f["problem"] for f in unsupported]    return Guarded(not unsupported, redact(answer), reasons)

Look at how guard_input uses the injection score. Injection classifiers have false positives: a policy question that quotes "ignore the previous version of this form" can look like an attack. Blocking those questions would annoy honest users every day. So a high score does not block. It reduces privilege: the turn runs with no tools, so a successful injection could at most produce a wrong answer, which the output checks and citations then constrain. This is a pattern worth copying: when you are unsure, shrink what the model can do rather than refusing the user.

In guard_output, a failed grounding check is not the end; it hands back to the drop, regenerate or hand-over policy from Section 3. The canary is the one hard block, because a leaked prompt has no partial fix.

Fail open or fail closed?

Each guard is itself software that can break. Decide in advance what happens when it does.

If this guard is downBehaviourWhy
Identity providerFail closed: no answersWithout identity, every other layer is meaningless
RedactorAnswer, but store nothingUsers still get help; nothing unredacted is kept
Injection scorerTreat every turn as risky: no toolsAnswers continue with less privilege
NLI grounding checkAnswer with a "sources not verified" banner, and alertCitations remain; the banner keeps trust honest
RouterPrompted fallback router (Section 7)Slower, same behaviour

Failing closed everywhere would make PolicyPal fragile; failing open everywhere would make it unsafe. The table is a set of explicit decisions, tested in the same way as the degraded modes in Section 7.

The road ahead

Launch is the start of the work, not the end. The habits that keep PolicyPal good are the ones built in this course, run on a schedule.

  1. Every change — the eval suite and flip list, the red-team suite, and a release manifest.
  2. Every week — the review queue, with every confirmed failure turned into an eval case.
  3. Every month — judge calibration on 30 human-labelled cases, cost and drift review, router health checks.
  4. Every quarter — refresh a tenth of the eval set from new traffic, re-run the model pilot against newer models, and review the privacy map.

The models will keep changing, and some of your choices should change with them. Longer context windows and cheaper tokens will make some chunking and budget decisions less tight, but they do not remove the reasons for retrieval: permissions, citations, fresh documents and cost at scale. Better tool use will make agents more reliable, which may move some workflows up the autonomy ladder, but only when your eval says so. What does not change is the method: measure on your own task, prefer construction to persuasion, and let failures you can see drive the next piece of work.

If you want to go deeper next, four directions follow naturally from this course. Tool protocols such as the Model Context Protocol, for connecting assistants to many internal systems in a standard way. Serving open models with engines such as vLLM, when a self-hosted component like the router grows. Advanced evaluation, including simulated users for multi-turn conversations. And AI governance, the risk registers and review processes that larger organisations need around systems like this one.

Check your understanding

0 of 3 answered

1.Three independent layers each stop 90% of a certain attack. Roughly how often does the attack get through all three, and what is the catch?

2.When the injection score is high, why does PolicyPal run the turn with tools disabled instead of refusing to answer?

3.The NLI grounding service is down. What does PolicyPal do, and why?