Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Tracing, Drift Detection and Incident Response
At 10:05 on a Wednesday, a team lead messaged the service owner: two of her agents had been told by the assistant that payment holidays were not available on car finance loans. They are available, under conditions. The questions came immediately. How many staff got this answer? Since when? Did any customer hear it? Which policy version, which prompt, which model? Is it still happening right now?
On Meridian's pilot, with only request and response logs, answering those questions for a similar problem had taken two days of searching text files. On the production design, it took twenty minutes. The difference was not better people. It was observability designed for a system that fails fluently.
This lesson designs what Meridian records for every request, which signals warn of drift, and what happens when something goes wrong. It produces part A of MER-09.
What a trace records
A trace is the record of one request as it passes through every step, with each step as a timed span. For an AI request, the trace must answer: what did the system see, what did it decide, and why?
| Span | Key attributes |
|---|---|
| Request | Capability, user role, hashed case reference, bundle version |
| Guard and rewrite | Input label, rewritten query, guard model version |
| Retrieval | Index version, candidate count, top chunk IDs and scores, filters applied |
| Pack | Context manifest: pieces, versions, tokens, dropped items |
| Model call | Route, model version, prompt version, input and output tokens, latency, finish reason |
| Checks | Each deterministic check with pass or fail |
| Response | Degraded mode or not, what was shown, feedback when it arrives |
Use the OpenTelemetry generative AI semantic conventions where they fit, such as gen_ai.request.model and gen_ai.usage.input_tokens. They are still marked as in development, so expect some names to change, but using them means tracing tools can read your data without custom mapping.
Content needs a separate decision. Full prompt and response text contains customer data and, in notes, sometimes health information. Meridian keeps two stores. Traces hold metadata and the context manifest, not full text; they are kept for 30 days and are open to the engineering team. The audit store holds full inputs and outputs, is kept for the records retention period, is readable only by named roles, and logs every access. A trace links to its audit record by ID, so an engineer with the right role can move from one to the other.
Drift signals
The quality objectives in section 8 are lagging signals: they depend on grading, so they move over days. Drift detection adds leading signals: cheap numbers computed on every request that move first.
- Topic mix, from the small classifier that labels each question.
- Retrieval top score. A falling score means staff are asking about things the index does not cover well.
- Decline rate. How often the assistant says it cannot find a supporting policy.
- Answer length and citation count.
- Deterministic check failure rate.
- Fallback and degraded-mode share.
- Unhelpful ratings from staff.
To compare distributions such as topic mix, Meridian uses the population stability index (PSI), which bank model monitoring teams already know from credit scorecards. It compares this week's distribution with a baseline.
1from math import log23def psi(baseline: list[float], current: list[float], floor: float = 1e-4) -> float:4 """Population stability index between two binned distributions (each sums to 1)."""5 total = 0.06 for b, c in zip(baseline, current):7 b, c = max(b, floor), max(c, floor)8 total += (c - b) * log(c / b)9 return total1011topics = ["arrears", "holidays", "fees", "interest", "other"]12baseline = [0.30, 0.25, 0.20, 0.15, 0.10]13this_week = [0.22, 0.24, 0.18, 0.14, 0.22]14print(round(psi(baseline, this_week), 2)) # 0.12The function adds up, for each bin, the change in share multiplied by the log of the ratio. The floor stops an empty bin from dividing by zero. A common rule of thumb reads PSI below 0.1 as stable, 0.1 to 0.25 as a moderate shift worth investigating, and above 0.25 as a significant shift. Here the result is 0.12, and almost all of it comes from the "other" topic growing from 10% to 22%. That is the pattern from section 8's case study: a new product that the classifier had no label for. Drift detection would have raised it days before the correctness objective breached.
Severities and kill switches
A normal incident scale is based on outages. AI needs a scale based on harm, because the worst incidents happen while every system is up.
| Severity | Example | Response |
|---|---|---|
| SEV1 | Wrong figure in a sent letter; restricted data shown; data about the wrong customer | Kill switch now, page on-call, tell risk and the data protection officer within an hour |
| SEV2 | A wrong answer repeated on one topic, no known customer impact yet | Topic-level kill switch within 30 minutes, fix within one working day |
| SEV3 | Quality objective breached; provider degraded | Service owner queue, business hours |
| SEV4 | A single wrong answer reported | Add to golden set, review weekly |
For a SEV1 involving personal data, the data protection officer decides whether it is a notifiable breach. Under GDPR, notifiable breaches must generally be reported to the supervisory authority within 72 hours of becoming aware of them, so that clock starts early.
Kill switches are how containment happens in minutes. Meridian has them at three levels: per capability (policy answers to search mode, summaries to facts only, letters to template mode), per topic (force search mode for questions classified to one policy area), and per model route. Each is a configuration flag the service owner can flip in under five minutes, with no deployment. Each is tested monthly, because a kill switch nobody has used is a kill switch that might not work.
The runbook for "it said something wrong"
- Capture — get the trace IDs from the report; the Atlas panel shows one on every answer.
- Scope — query traces for the same topic, the same cited sections and the same bundle over the last 30 days: how many answers, to how many staff?
- Contain — flip the narrowest kill switch that stops the harm.
- Diagnose — retrieval, policy version, prompt, model or check? The trace shows which step went wrong.
- Correct — tell affected staff; if customers were affected, route them into the existing complaints and remediation process.
- Learn — add golden cases, fix the cause, and hold a blameless review within five working days.
Part A of MER-09 records the trace schema, the signals with their thresholds, the severity table, the kill switches with their owners and test dates, and the runbook above.
Check your understanding
0 of 3 answered
1.Why does Meridian keep full prompt and response text in a separate audit store rather than in traces?
2.Topic mix PSI rises to 0.12, driven by the "other" topic doubling. What is the best response?
3.A team lead reports a wrong answer about car finance payment holidays. What is the right first containment action once the scope is known?