Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Quality Service-Level Objectives
Meridian's operations director asked a simple question in the design review: "On a Tuesday afternoon six months from now, how will I know the assistant is still giving good answers?" The dashboard the team had planned showed availability at 99.8% and p95 latency at 2.1 seconds. Both were green. Neither answered her question. The assistant could be available, fast and wrong.
Site reliability engineering has a well-tested vocabulary for "good enough, measured continuously": service-level indicators, objectives and error budgets. It was built for availability and latency, but it works just as well for quality, with one extra complication: many quality measures come from samples, not from every request.
This lesson defines Meridian's quality objectives, deals honestly with sampling, separates objectives from invariants, and writes down what happens when things go wrong. It produces part B of MER-08.
Indicators, objectives and budgets
Quality indicators come in two kinds.
- Census indicators are computed on every event. "Letter drafts passing all deterministic checks on first generation" is known for every single draft, because layer 1 checks run in production.
- Sampled indicators need grading, so they are measured on a sample. "Policy answers that are correct and supported by their citations" needs a judge or an expert, so Meridian grades 60 production answers a day with the calibrated judge, and 50 a week by analysts.
Sampled indicators carry uncertainty, just like the pass rates in section 1. Over a 28-day window, Meridian grades about 1,400 answers, which gives a range of roughly plus or minus 1.5 points around the measured rate. That is small enough to be useful against a 10-point budget. With only 100 graded answers a month, the range would be about 6 points, and the objective would be mostly noise.
Deciding whether an objective is met
Because of that uncertainty, "met" and "missed" need care. Meridian reports three states: met, at risk and breached. The code below reuses the Wilson interval from section 1.
1from math import sqrt23def wilson(passed: int, total: int, z: float = 1.96) -> tuple[float, float]:4 p = passed / total5 denom = 1 + z * z / total6 centre = (p + z * z / (2 * total)) / denom7 half = z * sqrt(p * (1 - p) / total + z * z / (4 * total * total)) / denom8 return centre - half, centre + half910def slo_status(passed: int, total: int, target: float) -> str:11 low, high = wilson(passed, total)12 if high < target:13 return "breached" # even the optimistic estimate misses the target14 if low < target:15 return "at risk" # we cannot yet show the target is met16 return "met"1718for passed in (1302, 1260, 1190):19 print(passed, "of 1400:", slo_status(passed, 1400, 0.90))slo_status compares the range, not the single measured rate, with the target. Running it prints the three cases below.
| Graded correct, of 1,400 | Measured | Range | Status against 90% |
|---|---|---|---|
| 1,302 | 93.0% | 91.5% to 94.2% | Met |
| 1,260 | 90.0% | 88.3% to 91.5% | At risk |
| 1,190 | 85.0% | 83.0% to 86.8% | Breached |
"At risk" is the honest answer when the data cannot yet tell you. It triggers attention, not alarm.
Meridian's quality objectives
| Indicator | Kind | Window | Objective |
|---|---|---|---|
| Policy answers correct and supported | Sampled: 60 a day by judge, 50 a week by analysts | 28 days | 90% |
| Out-of-scope questions correctly declined | Sampled: 20 a week, seeded | 28 days | 90% |
| Citations resolve to a current section | Census | 7 days | 99.9% |
| Summaries complete on first generation | Census, must-include check | 7 days | 98% |
| Letter drafts pass all checks on first generation | Census | 7 days | 97% |
| Answers rated unhelpful by staff | Census, optional feedback | 28 days | 8% or less |
| Policy answer p95 time to first text | Census | 28 days | 2.5 s |
Notice the summary and letter objectives are about first generation. A summary that fails the must-include check is regenerated or completed from records, so staff never see an incomplete one. The objective measures how often that safety net is needed. If it is needed more and more, something is drifting, even though no user is harmed yet.
Invariants are not objectives
Some things must have no error budget at all. A letter shown to a reviewer with a figure that differs from the calculator. A restricted policy shown to someone not entitled to it. A summary about the wrong customer. These are invariants: conditions the design guarantees with code, where a single violation is an incident, not a percentage.
Never write an invariant as an objective. "99.9% of letters have correct figures" sounds strict, but it quietly accepts one wrong letter in every thousand, about 17 a year at Meridian's volume. Write "zero, enforced by code, any violation is a severity 1 incident" instead, and design so it is true.
When the budget runs out
An objective without a consequence is a decoration. Meridian agreed the following policy with the Head of Servicing and model risk before launch, so nobody has to negotiate it during a bad week.
- More than half the budget left — normal releases of prompts, models and index changes.
- Between a quarter and a half left — releases need a second reviewer, and the next sprint prioritises quality work.
- Budget exhausted — freeze all changes except fixes, and open a problem review within two working days.
- Seven-day correctness below 85% — switch policy answers to search mode until a fix passes the release gate.
Quality alerts go to the service owner's queue during business hours. Nobody is woken at night because the sampled correctness rate dipped. Invariant violations are different: they open an incident immediately, which section 9 covers.
The objectives, the three-state reporting rule and the budget policy together form part B of MER-08. Section 9 connects them to tracing, alerts and the incident runbook.
Check your understanding
0 of 3 answered
1.The assistant is 99.8% available with p95 latency of 2.1 seconds. Why is that not enough to answer the operations director?
2.Why should "letter figures match the calculator" not be written as a 99.9% objective?
3.1,260 of 1,400 graded answers were correct against a 90% objective. What status should the dashboard show?