Course Content
AI Monitoring and Observability
3 sections · 7 lessons
Why Observability Matters
A support team ships a ticket-triage classifier. Offline accuracy on the held-out test set: 94.2%. It goes live in January. For five months every dashboard is green — uptime 99.97%, p95 latency 240 ms, HTTP error rate 0.02%, CPU comfortably under half. Nobody is paged. Nobody is worried.
In June, a quarterly quality audit samples 500 tickets by hand and finds that 31% of billing complaints have been routed to the wrong queue since March. The cause turns out to be small: in March the company renamed its paid tier from "Pro" to "Business", and customers started writing "my Business plan was charged twice" instead of "my Pro plan was charged twice". The model had never seen the word "Business" used that way. It did not crash. It did not time out. For every one of those tickets it returned a confident label in about 190 ms with a 200 OK.
Three months of misrouted billing complaints. Zero alerts. This is the failure mode that defines the whole discipline: AI systems fail while succeeding. The request completes, the status code is fine, the latency is fine, and the answer is wrong. Every monitoring habit you brought from ordinary web services is tuned to catch the opposite kind of failure — the loud kind — and it will sit quietly through this one for as long as you let it.
Why your existing monitoring cannot see this
Classical service monitoring rests on an assumption so deep that people rarely say it out loud: the system knows when it is broken. A database that loses its connection raises an exception. A disk that fills up returns ENOSPC. A service that is overloaded returns 503s and its latency histogram shifts right. The failure announces itself, and your job is only to notice the announcement quickly.
A model has no equivalent of an exception. Its output for a request it understands perfectly and its output for a request that is completely outside its training distribution have exactly the same shape: a label, a score, a block of text. The confidence number does not help as much as you would hope — models are routinely confident and wrong, and a shift in vocabulary of the kind above often raises average confidence rather than lowering it, because the new phrasing lands squarely in some other class's region of the feature space.
The defining property of an AI failure is that it is indistinguishable from success at the level of the request. You cannot detect it one request at a time; you can only detect it in the statistics of many requests.
That single sentence explains most of what follows. If failures were visible per request, you would just check each request. Because they are only visible in aggregate, you need to be capturing aggregate signals continuously, comparing them to a baseline, and reasoning about whether a difference is real or noise. That is a statistical problem, not a plumbing problem — and treating it as plumbing is the most common way teams end up in June with three months of bad routing behind them.
Monitoring and observability are not synonyms
The two words get used interchangeably, and the distinction is worth getting right because it changes what you build.
Monitoring is checking known signals against known thresholds. You decided in advance that p95 latency matters and that 500 ms is the line. You wrote the check. It answers the question you already thought to ask.
Observability is a property of a system: how much you can infer about its internal state purely from the outputs it emits, including for questions you did not anticipate. The term is borrowed from control theory, where a system is called observable if its internal state can be reconstructed from its external measurements. In practice, observability means that when someone asks "why did Spanish-language requests from mobile clients start getting truncated answers last Tuesday afternoon?", the data to answer it is already there — nobody has to ship a new build to find out.
| Monitoring | Observability | |
|---|---|---|
| Question shape | "Is X above the line?" | "Why is X behaving like that?" |
| Handles | Known unknowns — failures you predicted | Unknown unknowns — failures you did not |
| Data shape | Pre-aggregated counters and gauges | High-dimensional events you can slice after the fact |
| Cost profile | Cheap, bounded | More expensive, must be managed |
| Fails when | The failure is one you did not predict | You did not record the dimension you now need |
| Output | Alerts | Answers |
You need both, and they trade off. Monitoring gives you the page at 3 a.m.; observability gives you the ability to explain what the page meant by 3:20. A team with alerts but no observability gets woken up and then stares at a red graph with no idea what to do. A team with observability but no alerts has all the answers and finds out three months late that anyone was asking.
The three pillars, and what each one is actually for
Logs, metrics, and traces are usually presented as a list. They are better understood as three different answers to "what resolution do I need, and can I afford it?"
| Pillar | What it is | Answers | Cost driver | Blind spot |
|---|---|---|---|---|
| Metrics | Numbers aggregated over time windows — counters, gauges, histograms | "Is something wrong, and since when?" | Number of distinct label combinations (cardinality) | Cannot tell you which request; aggregation destroys the individual |
| Logs | Discrete timestamped events, ideally structured as JSON | "What exactly happened to this one request?" | Bytes ingested and retained | Expensive to aggregate; no causal structure between entries |
| Traces | A causally linked tree of timed operations across services | "Where did the 4 seconds go?" | Spans stored; usually requires sampling | Sampling means the request you care about may not be there |
A concrete way to feel the difference. Your p95 latency jumps from 800 ms to 4.2 s at 14:05.
- The metric tells you it happened and pins the start to 14:05. It cannot tell you why.
- The trace of one slow request shows: retrieval 180 ms, reranking 90 ms, LLM call 3,900 ms, formatting 30 ms. The LLM call is the problem.
- The log line for that LLM call carries
retry_count: 2,provider_region: "eu-west", andfinish_reason: "length". Now you know: retries against a degraded region, and the model is running to its token limit.
Each pillar alone leaves you stuck. The metric without the trace gives you a red graph. The trace without the log gives you a slow box with no reason. The log without the metric means you never looked, because nothing told you to.
Where AI systems break differently
The word "AI-specific" gets waved around loosely. Here is the concrete list, with what each failure looks like from the outside.
| Failure | What it is | Externally visible as | Signal that catches it |
|---|---|---|---|
| Data drift | Input distribution moves away from training data | Nothing — all 200s | Distribution distance on input features (e.g. PSI) against a fixed baseline |
| Concept drift | The correct answer for a given input changes | Nothing, until labels arrive | Accuracy on delayed ground truth; complaint and thumbs-down rates |
| Prediction drift | Output mix shifts (e.g. class 3 goes from 12% to 34% of predictions) | Nothing | Prediction histogram vs baseline; mean confidence |
| Training–serving skew | Feature computed one way offline, another way online | Live accuracy far below offline accuracy from day one | Comparing feature statistics computed in both paths |
| Silent upstream schema change | A field starts arriving null, or in a new unit | Nothing, if the model imputes | Null-rate and range checks per feature |
| Prompt/context bloat | Retrieved context grows, tokens per call climb | Slightly slower, much more expensive | Tokens-per-request metric, cost per request |
| Hallucination / groundedness loss | Generated text not supported by retrieved sources | Fluent, confident, wrong | Groundedness scoring on sampled outputs; citation coverage |
| Provider-side model change | The hosted model you call is updated underneath you | Output style and length shift overnight | Output length distribution, refusal rate, eval-set score run on a schedule |
| Feedback loop | Model outputs become tomorrow's training data | Metrics improve while reality degrades | Held-out human-labelled sample that never enters training |
Notice the third column. In seven of those nine rows the externally visible symptom is nothing. That is the whole argument for a separate discipline.
The one that catches everyone: feedback loops
A recommender surfaces items it scores highly. Users click what is surfaced, because they cannot click what they never saw. Those clicks become the next training set. Click-through-rate on the dashboard rises quarter after quarter while catalogue coverage collapses and the model gets steadily worse at anything outside a narrowing band. Every metric you are watching says you are winning.
The only defence is a measurement channel the model cannot influence: a small random-exposure holdout, or a human-labelled sample drawn independently of what the model recommended. If every number you track is downstream of the model's own decisions, you have built a mirror, not a monitor.
What it costs to find out in June
Two arithmetics, because "monitoring pays for itself" is usually asserted rather than shown.
Silent cost regression
An assistant serves 500,000 requests a day. Average 1,200 input tokens and 350 output tokens per request. At 3 dollars per million input tokens and 15 dollars per million output tokens:
input: 500,000 x 1,200 = 600,000,000 tok/day = 600 MTok -> 600 x 3 = 1,800/dayoutput: 500,000 x 350 = 175,000,000 tok/day = 175 MTok -> 175 x 15 = 2,625/daytotal daily = 4,425 monthly (30d) = 132,750Now someone widens the retrieval window from 3 chunks to 5. Each request carries roughly 400 more input tokens. Latency rises by perhaps 40 ms — invisible against 800 ms of noise. Quality is unchanged or slightly better, so nobody complains.
extra input: 500,000 x 400 = 200,000,000 tok/day = 200 MTokextra cost: 200 x 3 = 600/day -> 18,000/monthEighteen thousand dollars a month, invisible to latency monitoring, error monitoring, and quality review. The only signal that catches it is a boring per-request token-count metric with a baseline. That metric costs almost nothing to emit.
Silent quality regression
Back to the triage classifier. Billing complaints are 8% of 4,000 daily tickets = 320 per day. A 31% misroute rate means about 99 tickets per day land in the wrong queue. Each costs roughly 20 extra minutes of handling and re-routing. Over the 90 days before the audit:
99 tickets/day x 90 days = 8,910 misrouted tickets8,910 x 20 min = 178,200 min = 2,970 agent-hoursNearly three thousand hours of avoidable work, plus the churn from customers whose billing problem sat in the wrong queue. A daily prediction-distribution check would have flagged the shift within about a week of March.
Monitoring is not an operational nicety bolted on after launch. For a model, it is the only mechanism that connects the thing you deployed to the thing that is actually happening.
The four layers worth measuring
Teams tend to instrument layer one thoroughly and layers two through four not at all. Layers two through four are where the money and the damage live.
| Layer | Representative metrics | Typical cadence | Who cares |
|---|---|---|---|
| 1. System health | Request rate, error rate by class, p50/p95/p99 latency, queue depth, saturation (CPU, memory, GPU utilisation, connection pool) | Seconds | On-call |
| 2. Model performance | Accuracy / precision / recall / F1 on delayed labels, mean confidence, calibration error, prediction class mix, human thumbs-up rate | Hours to days (label lag) | ML team, product |
| 3. Data quality | Per-feature null rate, out-of-range rate, cardinality of categoricals, schema conformance, PSI vs training baseline, freshness of feature values | Minutes to hours | Data engineering, ML team |
| 4. Cost & efficiency | Tokens in / out per request, cost per request, cache hit rate, retry rate, tokens wasted on truncated responses, cost per resolved ticket | Minutes | Whoever owns the budget |
For layer one specifically, the industry shorthand worth knowing is the four golden signals: latency, traffic, errors, saturation. Two related mnemonics: RED (Rate, Errors, Duration) for request-driven services, and USE (Utilisation, Saturation, Errors) for resources like GPUs and connection pools. They are checklists, not theories — their value is that they stop you from instrumenting five things and forgetting saturation.
A threshold is a statistical claim, and most people get the arithmetic wrong
The moment you have metrics, someone will say "alert me when latency is more than three standard deviations above normal". It sounds rigorous. Work out what it actually produces.
Suppose your metric really is roughly normal and you check it once a minute. That is:
60 x 24 x 7 = 10,080 checks per weekFor a normal distribution the probability of landing above +3σ on the upper tail is 0.00135. So on a perfectly healthy system:
10,080 x 0.00135 = 13.6 false alarms per week (about 2 per day)Thirteen pages a week from one metric, on a system where nothing is wrong. Now multiply by the number of metrics you are watching. Forty metrics, each with its own 3σ rule:
40 x 13.6 = 544 false alarms per week (about 78 per day)This is not a hypothetical failure — it is the single most common way monitoring dies. Within two weeks the team mutes the channel, and the alert that finally matters arrives in a feed nobody reads. The right response is not to give up on statistics; it is to use them properly, by combining looser thresholds with persistence requirements (fire only if the breach holds for several consecutive checks) so that false alarms fall faster than detection power does.
Every threshold you set is a bet about a probability distribution. If you never compute the false-alarm rate that bet implies, you have not set a threshold — you have set a trap for your own team.
Instrumenting one endpoint properly
The smallest useful thing you can build. Note that it touches all four layers, and that none of it is exotic.
1import time, uuid, logging, json2from prometheus_client import Counter, Histogram34# Layer 1 + 4: bounded labels only. Never put user_id or prompt text in a label.5REQUESTS = Counter(6 "llm_requests_total", "LLM requests", ["model", "endpoint", "status"]7)8LATENCY = Histogram(9 "llm_latency_seconds", "End-to-end latency", ["model", "endpoint"],10 buckets=(0.1, 0.25, 0.5, 1, 2, 4, 8, 16),11)12TOKENS = Histogram(13 "llm_tokens_total", "Tokens per request", ["model", "direction"],14 buckets=(64, 128, 256, 512, 1024, 2048, 4096, 8192),15)16CONFIDENCE = Histogram(17 "model_confidence", "Top-class probability", ["model"],18 buckets=(0.1, 0.3, 0.5, 0.7, 0.8, 0.9, 0.95, 0.99),19)2021log = logging.getLogger("triage")2223def handle(request, model_name="triage-v3"):24 request_id = str(uuid.uuid4())25 started = time.perf_counter()26 status = "ok"27 try:28 result = model.predict(request.text) # your inference call29 return result30 except Exception as exc:31 status = "error"32 log.exception("inference failed", extra={"request_id": request_id})33 raise34 finally:35 elapsed = time.perf_counter() - started36 REQUESTS.labels(model_name, "/triage", status).inc()37 LATENCY.labels(model_name, "/triage").observe(elapsed)38 if status == "ok":39 TOKENS.labels(model_name, "in").observe(result.input_tokens)40 TOKENS.labels(model_name, "out").observe(result.output_tokens)41 CONFIDENCE.labels(model_name).observe(result.confidence)42 # Layer 2 + 3: one structured event per request, sliceable later.43 log.info(json.dumps({44 "request_id": request_id,45 "model": model_name,46 "model_version": model.version,47 "predicted_class": result.label,48 "confidence": round(result.confidence, 4),49 "input_chars": len(request.text),50 "input_tokens": result.input_tokens,51 "output_tokens": result.output_tokens,52 "latency_ms": round(elapsed * 1000, 1),53 "locale": request.locale,54 "client": request.client_type,55 }))Three things in that snippet are doing heavy lifting and are easy to skip.
The finally block. Metrics are emitted whether or not the call succeeded. Instrumentation that lives on the happy path goes blind precisely when the system is unhealthy — which is the moment you need it.
The label discipline. Labels are bounded: four models, a dozen endpoints, a few statuses. Add user_id and every combination becomes its own stored time series. With 50,000 users that is 240 series turning into 12,000,000, and at roughly 3.5 KB of memory per active series your metrics backend needs about 42 GB of RAM to hold them. It will not; it will fall over. High-cardinality dimensions belong in the structured log line, where storage is linear in events rather than multiplicative in label values.
The model_version field. Without it, a regression that appears the day after a deploy is indistinguishable from one caused by a change in traffic. With it, one group-by answers the question.
What this changes about how you build
The practical consequence is that instrumentation stops being a post-launch chore and becomes part of the definition of "the feature is built".
Before you deploy any model, you should be able to answer four questions, and if you cannot, the deployment is not finished:
- What is the baseline? Not "normal-ish" — a stored, versioned snapshot of input feature distributions, prediction mix, latency percentiles, and tokens per request, captured on known-good traffic. Drift is meaningless without a fixed thing to drift from, and a rolling baseline that updates daily will happily follow a slow degradation all the way down.
- How will I learn the answer was wrong? Trace the ground-truth path concretely: who labels it, how long the lag is, how the label gets back to a store you can join against predictions. If the honest answer is "we won't", you need a proxy — thumbs-down rate, escalation rate, retry rate, a weekly human-labelled sample — and you need it wired up before launch, not after the audit.
- What dimension will I want to slice by? The bug will be in Spanish, or on mobile, or for one customer tier, or for one model version. Every one of those has to be a field on the event at write time. You cannot add a dimension retroactively to data already collected.
- What is the false-alarm rate of every threshold I just set? Compute it. Multiply it by the number of rules. If the total is more than a handful per week, you have not built monitoring — you have built a reason for your team to stop reading alerts.
The team in the opening scenario had excellent monitoring for a web service and none at all for a model. Their uptime number was true the entire time. It just never had anything to do with whether the system was doing its job.