Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Quality Service-Level Objectives


Meridian's operations director asked a simple question in the design review: "On a Tuesday afternoon six months from now, how will I know the assistant is still giving good answers?" The dashboard the team had planned showed availability at 99.8% and p95 latency at 2.1 seconds. Both were green. Neither answered her question. The assistant could be available, fast and wrong.

Site reliability engineering has a well-tested vocabulary for "good enough, measured continuously": service-level indicators, objectives and error budgets. It was built for availability and latency, but it works just as well for quality, with one extra complication: many quality measures come from samples, not from every request.

This lesson defines Meridian's quality objectives, deals honestly with sampling, separates objectives from invariants, and writes down what happens when things go wrong. It produces part B of MER-08.

Graded answers against a 90% objective93.0%91.5to 94.2met90.0%88.3to 91.5at risk85.0%83.0to 86.8breachedMeasuredRangeStatus1,302 of 1,4001,260 of 1,4001,190 of 1,400
Judging the range rather than the point turns a coin-flip reading into an honest "at risk" that asks for attention, not alarm.

Indicators, objectives and budgets

Quality indicators come in two kinds.

  • Census indicators are computed on every event. "Letter drafts passing all deterministic checks on first generation" is known for every single draft, because layer 1 checks run in production.
  • Sampled indicators need grading, so they are measured on a sample. "Policy answers that are correct and supported by their citations" needs a judge or an expert, so Meridian grades 60 production answers a day with the calibrated judge, and 50 a week by analysts.

Sampled indicators carry uncertainty, just like the pass rates in section 1. Over a 28-day window, Meridian grades about 1,400 answers, which gives a range of roughly plus or minus 1.5 points around the measured rate. That is small enough to be useful against a 10-point budget. With only 100 graded answers a month, the range would be about 6 points, and the objective would be mostly noise.

Deciding whether an objective is met

Because of that uncertainty, "met" and "missed" need care. Meridian reports three states: met, at risk and breached. The code below reuses the Wilson interval from section 1.

Python
from math import sqrtdef wilson(passed: int, total: int, z: float = 1.96) -> tuple[float, float]:    p = passed / total    denom = 1 + z * z / total    centre = (p + z * z / (2 * total)) / denom    half = z * sqrt(p * (1 - p) / total + z * z / (4 * total * total)) / denom    return centre - half, centre + halfdef slo_status(passed: int, total: int, target: float) -> str:    low, high = wilson(passed, total)    if high < target:        return "breached"   # even the optimistic estimate misses the target    if low < target:        return "at risk"    # we cannot yet show the target is met    return "met"for passed in (1302, 1260, 1190):    print(passed, "of 1400:", slo_status(passed, 1400, 0.90))

slo_status compares the range, not the single measured rate, with the target. Running it prints the three cases below.

Graded correct, of 1,400MeasuredRangeStatus against 90%
1,30293.0%91.5% to 94.2%Met
1,26090.0%88.3% to 91.5%At risk
1,19085.0%83.0% to 86.8%Breached

"At risk" is the honest answer when the data cannot yet tell you. It triggers attention, not alarm.

Meridian's quality objectives

IndicatorKindWindowObjective
Policy answers correct and supportedSampled: 60 a day by judge, 50 a week by analysts28 days90%
Out-of-scope questions correctly declinedSampled: 20 a week, seeded28 days90%
Citations resolve to a current sectionCensus7 days99.9%
Summaries complete on first generationCensus, must-include check7 days98%
Letter drafts pass all checks on first generationCensus7 days97%
Answers rated unhelpful by staffCensus, optional feedback28 days8% or less
Policy answer p95 time to first textCensus28 days2.5 s

Notice the summary and letter objectives are about first generation. A summary that fails the must-include check is regenerated or completed from records, so staff never see an incomplete one. The objective measures how often that safety net is needed. If it is needed more and more, something is drifting, even though no user is harmed yet.

Invariants are not objectives

Some things must have no error budget at all. A letter shown to a reviewer with a figure that differs from the calculator. A restricted policy shown to someone not entitled to it. A summary about the wrong customer. These are invariants: conditions the design guarantees with code, where a single violation is an incident, not a percentage.

Never write an invariant as an objective. "99.9% of letters have correct figures" sounds strict, but it quietly accepts one wrong letter in every thousand, about 17 a year at Meridian's volume. Write "zero, enforced by code, any violation is a severity 1 incident" instead, and design so it is true.

When the budget runs out

An objective without a consequence is a decoration. Meridian agreed the following policy with the Head of Servicing and model risk before launch, so nobody has to negotiate it during a bad week.

  1. More than half the budget left — normal releases of prompts, models and index changes.
  2. Between a quarter and a half left — releases need a second reviewer, and the next sprint prioritises quality work.
  3. Budget exhausted — freeze all changes except fixes, and open a problem review within two working days.
  4. Seven-day correctness below 85% — switch policy answers to search mode until a fix passes the release gate.

Quality alerts go to the service owner's queue during business hours. Nobody is woken at night because the sampled correctness rate dipped. Invariant violations are different: they open an incident immediately, which section 9 covers.

The objectives, the three-state reporting rule and the budget policy together form part B of MER-08. Section 9 connects them to tracing, alerts and the incident runbook.

Check your understanding

0 of 3 answered

1.The assistant is 99.8% available with p95 latency of 2.1 seconds. Why is that not enough to answer the operations director?

2.Why should "letter figures match the calculator" not be written as a 99.9% objective?

3.1,260 of 1,400 graded answers were correct against a 90% objective. What status should the dashboard show?