AI Monitoring and Observability

Why Observability Matters


A support team ships a ticket-triage classifier. Offline accuracy on the held-out test set: 94.2%. It goes live in January. For five months every dashboard is green — uptime 99.97%, p95 latency 240 ms, HTTP error rate 0.02%, CPU comfortably under half. Nobody is paged. Nobody is worried.

In June, a quarterly quality audit samples 500 tickets by hand and finds that 31% of billing complaints have been routed to the wrong queue since March. The cause turns out to be small: in March the company renamed its paid tier from "Pro" to "Business", and customers started writing "my Business plan was charged twice" instead of "my Pro plan was charged twice". The model had never seen the word "Business" used that way. It did not crash. It did not time out. For every one of those tickets it returned a confident label in about 190 ms with a 200 OK.

Three months of misrouted billing complaints. Zero alerts. This is the failure mode that defines the whole discipline: AI systems fail while succeeding. The request completes, the status code is fine, the latency is fine, and the answer is wrong. Every monitoring habit you brought from ordinary web services is tuned to catch the opposite kind of failure — the loud kind — and it will sit quietly through this one for as long as you let it.

Five months of green, and what green could not seeWhat every dashboard reported• Uptime 99.97 percent• p95 latency 240 ms, flat• HTTP error rate 0.02 percent• CPU comfortably under halfWhat nobody was measuring• 31 percent of billing tickets misrouted• "Pro" renamed "Business" in March• Confidence stayed high on wrong labels• No accuracy check on live traffic
Every wrong prediction was returned as a healthy 200 OK — infrastructure monitoring cannot see inside the response.

Why your existing monitoring cannot see this

Classical service monitoring rests on an assumption so deep that people rarely say it out loud: the system knows when it is broken. A database that loses its connection raises an exception. A disk that fills up returns ENOSPC. A service that is overloaded returns 503s and its latency histogram shifts right. The failure announces itself, and your job is only to notice the announcement quickly.

A model has no equivalent of an exception. Its output for a request it understands perfectly and its output for a request that is completely outside its training distribution have exactly the same shape: a label, a score, a block of text. The confidence number does not help as much as you would hope — models are routinely confident and wrong, and a shift in vocabulary of the kind above often raises average confidence rather than lowering it, because the new phrasing lands squarely in some other class's region of the feature space.

The defining property of an AI failure is that it is indistinguishable from success at the level of the request. You cannot detect it one request at a time; you can only detect it in the statistics of many requests.

That single sentence explains most of what follows. If failures were visible per request, you would just check each request. Because they are only visible in aggregate, you need to be capturing aggregate signals continuously, comparing them to a baseline, and reasoning about whether a difference is real or noise. That is a statistical problem, not a plumbing problem — and treating it as plumbing is the most common way teams end up in June with three months of bad routing behind them.

Monitoring and observability are not synonyms

The two words get used interchangeably, and the distinction is worth getting right because it changes what you build.

Monitoring is checking known signals against known thresholds. You decided in advance that p95 latency matters and that 500 ms is the line. You wrote the check. It answers the question you already thought to ask.

Observability is a property of a system: how much you can infer about its internal state purely from the outputs it emits, including for questions you did not anticipate. The term is borrowed from control theory, where a system is called observable if its internal state can be reconstructed from its external measurements. In practice, observability means that when someone asks "why did Spanish-language requests from mobile clients start getting truncated answers last Tuesday afternoon?", the data to answer it is already there — nobody has to ship a new build to find out.

MonitoringObservability
Question shape"Is X above the line?""Why is X behaving like that?"
HandlesKnown unknowns — failures you predictedUnknown unknowns — failures you did not
Data shapePre-aggregated counters and gaugesHigh-dimensional events you can slice after the fact
Cost profileCheap, boundedMore expensive, must be managed
Fails whenThe failure is one you did not predictYou did not record the dimension you now need
OutputAlertsAnswers

You need both, and they trade off. Monitoring gives you the page at 3 a.m.; observability gives you the ability to explain what the page meant by 3:20. A team with alerts but no observability gets woken up and then stares at a red graph with no idea what to do. A team with observability but no alerts has all the answers and finds out three months late that anyone was asking.

The three pillars, and what each one is actually for

Logs, metrics, and traces are usually presented as a list. They are better understood as three different answers to "what resolution do I need, and can I afford it?"

PillarWhat it isAnswersCost driverBlind spot
MetricsNumbers aggregated over time windows — counters, gauges, histograms"Is something wrong, and since when?"Number of distinct label combinations (cardinality)Cannot tell you which request; aggregation destroys the individual
LogsDiscrete timestamped events, ideally structured as JSON"What exactly happened to this one request?"Bytes ingested and retainedExpensive to aggregate; no causal structure between entries
TracesA causally linked tree of timed operations across services"Where did the 4 seconds go?"Spans stored; usually requires samplingSampling means the request you care about may not be there

A concrete way to feel the difference. Your p95 latency jumps from 800 ms to 4.2 s at 14:05.

  • The metric tells you it happened and pins the start to 14:05. It cannot tell you why.
  • The trace of one slow request shows: retrieval 180 ms, reranking 90 ms, LLM call 3,900 ms, formatting 30 ms. The LLM call is the problem.
  • The log line for that LLM call carries retry_count: 2, provider_region: "eu-west", and finish_reason: "length". Now you know: retries against a degraded region, and the model is running to its token limit.

Each pillar alone leaves you stuck. The metric without the trace gives you a red graph. The trace without the log gives you a slow box with no reason. The log without the metric means you never looked, because nothing told you to.

Where AI systems break differently

The word "AI-specific" gets waved around loosely. Here is the concrete list, with what each failure looks like from the outside.

FailureWhat it isExternally visible asSignal that catches it
Data driftInput distribution moves away from training dataNothing — all 200sDistribution distance on input features (e.g. PSI) against a fixed baseline
Concept driftThe correct answer for a given input changesNothing, until labels arriveAccuracy on delayed ground truth; complaint and thumbs-down rates
Prediction driftOutput mix shifts (e.g. class 3 goes from 12% to 34% of predictions)NothingPrediction histogram vs baseline; mean confidence
Training–serving skewFeature computed one way offline, another way onlineLive accuracy far below offline accuracy from day oneComparing feature statistics computed in both paths
Silent upstream schema changeA field starts arriving null, or in a new unitNothing, if the model imputesNull-rate and range checks per feature
Prompt/context bloatRetrieved context grows, tokens per call climbSlightly slower, much more expensiveTokens-per-request metric, cost per request
Hallucination / groundedness lossGenerated text not supported by retrieved sourcesFluent, confident, wrongGroundedness scoring on sampled outputs; citation coverage
Provider-side model changeThe hosted model you call is updated underneath youOutput style and length shift overnightOutput length distribution, refusal rate, eval-set score run on a schedule
Feedback loopModel outputs become tomorrow's training dataMetrics improve while reality degradesHeld-out human-labelled sample that never enters training

Notice the third column. In seven of those nine rows the externally visible symptom is nothing. That is the whole argument for a separate discipline.

The one that catches everyone: feedback loops

A recommender surfaces items it scores highly. Users click what is surfaced, because they cannot click what they never saw. Those clicks become the next training set. Click-through-rate on the dashboard rises quarter after quarter while catalogue coverage collapses and the model gets steadily worse at anything outside a narrowing band. Every metric you are watching says you are winning.

The only defence is a measurement channel the model cannot influence: a small random-exposure holdout, or a human-labelled sample drawn independently of what the model recommended. If every number you track is downstream of the model's own decisions, you have built a mirror, not a monitor.

What it costs to find out in June

Two arithmetics, because "monitoring pays for itself" is usually asserted rather than shown.

Silent cost regression

An assistant serves 500,000 requests a day. Average 1,200 input tokens and 350 output tokens per request. At 3 dollars per million input tokens and 15 dollars per million output tokens:

Text
input:   500,000 x 1,200 = 600,000,000 tok/day = 600 MTok  ->  600 x 3  = 1,800/dayoutput:  500,000 x   350 = 175,000,000 tok/day = 175 MTok  ->  175 x 15 = 2,625/daytotal daily = 4,425      monthly (30d) = 132,750

Now someone widens the retrieval window from 3 chunks to 5. Each request carries roughly 400 more input tokens. Latency rises by perhaps 40 ms — invisible against 800 ms of noise. Quality is unchanged or slightly better, so nobody complains.

Text
extra input: 500,000 x 400 = 200,000,000 tok/day = 200 MTokextra cost:  200 x 3 = 600/day  ->  18,000/month

Eighteen thousand dollars a month, invisible to latency monitoring, error monitoring, and quality review. The only signal that catches it is a boring per-request token-count metric with a baseline. That metric costs almost nothing to emit.

Silent quality regression

Back to the triage classifier. Billing complaints are 8% of 4,000 daily tickets = 320 per day. A 31% misroute rate means about 99 tickets per day land in the wrong queue. Each costs roughly 20 extra minutes of handling and re-routing. Over the 90 days before the audit:

Text
99 tickets/day x 90 days = 8,910 misrouted tickets8,910 x 20 min = 178,200 min = 2,970 agent-hours

Nearly three thousand hours of avoidable work, plus the churn from customers whose billing problem sat in the wrong queue. A daily prediction-distribution check would have flagged the shift within about a week of March.

Monitoring is not an operational nicety bolted on after launch. For a model, it is the only mechanism that connects the thing you deployed to the thing that is actually happening.

The four layers worth measuring

Teams tend to instrument layer one thoroughly and layers two through four not at all. Layers two through four are where the money and the damage live.

LayerRepresentative metricsTypical cadenceWho cares
1. System healthRequest rate, error rate by class, p50/p95/p99 latency, queue depth, saturation (CPU, memory, GPU utilisation, connection pool)SecondsOn-call
2. Model performanceAccuracy / precision / recall / F1 on delayed labels, mean confidence, calibration error, prediction class mix, human thumbs-up rateHours to days (label lag)ML team, product
3. Data qualityPer-feature null rate, out-of-range rate, cardinality of categoricals, schema conformance, PSI vs training baseline, freshness of feature valuesMinutes to hoursData engineering, ML team
4. Cost & efficiencyTokens in / out per request, cost per request, cache hit rate, retry rate, tokens wasted on truncated responses, cost per resolved ticketMinutesWhoever owns the budget

For layer one specifically, the industry shorthand worth knowing is the four golden signals: latency, traffic, errors, saturation. Two related mnemonics: RED (Rate, Errors, Duration) for request-driven services, and USE (Utilisation, Saturation, Errors) for resources like GPUs and connection pools. They are checklists, not theories — their value is that they stop you from instrumenting five things and forgetting saturation.

A threshold is a statistical claim, and most people get the arithmetic wrong

The moment you have metrics, someone will say "alert me when latency is more than three standard deviations above normal". It sounds rigorous. Work out what it actually produces.

Suppose your metric really is roughly normal and you check it once a minute. That is:

Text
60 x 24 x 7 = 10,080 checks per week

For a normal distribution the probability of landing above +3σ on the upper tail is 0.00135. So on a perfectly healthy system:

Text
10,080 x 0.00135 = 13.6 false alarms per week  (about 2 per day)

Thirteen pages a week from one metric, on a system where nothing is wrong. Now multiply by the number of metrics you are watching. Forty metrics, each with its own 3σ rule:

Text
40 x 13.6 = 544 false alarms per week  (about 78 per day)

This is not a hypothetical failure — it is the single most common way monitoring dies. Within two weeks the team mutes the channel, and the alert that finally matters arrives in a feed nobody reads. The right response is not to give up on statistics; it is to use them properly, by combining looser thresholds with persistence requirements (fire only if the breach holds for several consecutive checks) so that false alarms fall faster than detection power does.

Every threshold you set is a bet about a probability distribution. If you never compute the false-alarm rate that bet implies, you have not set a threshold — you have set a trap for your own team.

Instrumenting one endpoint properly

The smallest useful thing you can build. Note that it touches all four layers, and that none of it is exotic.

Python
import time, uuid, logging, jsonfrom prometheus_client import Counter, Histogram# Layer 1 + 4: bounded labels only. Never put user_id or prompt text in a label.REQUESTS = Counter(    "llm_requests_total", "LLM requests", ["model", "endpoint", "status"])LATENCY = Histogram(    "llm_latency_seconds", "End-to-end latency", ["model", "endpoint"],    buckets=(0.1, 0.25, 0.5, 1, 2, 4, 8, 16),)TOKENS = Histogram(    "llm_tokens_total", "Tokens per request", ["model", "direction"],    buckets=(64, 128, 256, 512, 1024, 2048, 4096, 8192),)CONFIDENCE = Histogram(    "model_confidence", "Top-class probability", ["model"],    buckets=(0.1, 0.3, 0.5, 0.7, 0.8, 0.9, 0.95, 0.99),)log = logging.getLogger("triage")def handle(request, model_name="triage-v3"):    request_id = str(uuid.uuid4())    started = time.perf_counter()    status = "ok"    try:        result = model.predict(request.text)     # your inference call        return result    except Exception as exc:        status = "error"        log.exception("inference failed", extra={"request_id": request_id})        raise    finally:        elapsed = time.perf_counter() - started        REQUESTS.labels(model_name, "/triage", status).inc()        LATENCY.labels(model_name, "/triage").observe(elapsed)        if status == "ok":            TOKENS.labels(model_name, "in").observe(result.input_tokens)            TOKENS.labels(model_name, "out").observe(result.output_tokens)            CONFIDENCE.labels(model_name).observe(result.confidence)            # Layer 2 + 3: one structured event per request, sliceable later.            log.info(json.dumps({                "request_id": request_id,                "model": model_name,                "model_version": model.version,                "predicted_class": result.label,                "confidence": round(result.confidence, 4),                "input_chars": len(request.text),                "input_tokens": result.input_tokens,                "output_tokens": result.output_tokens,                "latency_ms": round(elapsed * 1000, 1),                "locale": request.locale,                "client": request.client_type,            }))

Three things in that snippet are doing heavy lifting and are easy to skip.

The finally block. Metrics are emitted whether or not the call succeeded. Instrumentation that lives on the happy path goes blind precisely when the system is unhealthy — which is the moment you need it.

The label discipline. Labels are bounded: four models, a dozen endpoints, a few statuses. Add user_id and every combination becomes its own stored time series. With 50,000 users that is 240 series turning into 12,000,000, and at roughly 3.5 KB of memory per active series your metrics backend needs about 42 GB of RAM to hold them. It will not; it will fall over. High-cardinality dimensions belong in the structured log line, where storage is linear in events rather than multiplicative in label values.

The model_version field. Without it, a regression that appears the day after a deploy is indistinguishable from one caused by a change in traffic. With it, one group-by answers the question.

What this changes about how you build

The practical consequence is that instrumentation stops being a post-launch chore and becomes part of the definition of "the feature is built".

Before you deploy any model, you should be able to answer four questions, and if you cannot, the deployment is not finished:

  1. What is the baseline? Not "normal-ish" — a stored, versioned snapshot of input feature distributions, prediction mix, latency percentiles, and tokens per request, captured on known-good traffic. Drift is meaningless without a fixed thing to drift from, and a rolling baseline that updates daily will happily follow a slow degradation all the way down.
  2. How will I learn the answer was wrong? Trace the ground-truth path concretely: who labels it, how long the lag is, how the label gets back to a store you can join against predictions. If the honest answer is "we won't", you need a proxy — thumbs-down rate, escalation rate, retry rate, a weekly human-labelled sample — and you need it wired up before launch, not after the audit.
  3. What dimension will I want to slice by? The bug will be in Spanish, or on mobile, or for one customer tier, or for one model version. Every one of those has to be a field on the event at write time. You cannot add a dimension retroactively to data already collected.
  4. What is the false-alarm rate of every threshold I just set? Compute it. Multiply it by the number of rules. If the total is more than a handful per week, you have not built monitoring — you have built a reason for your team to stop reading alerts.

The team in the opening scenario had excellent monitoring for a web service and none at all for a model. Their uptime number was true the entire time. It just never had anything to do with whether the system was doing its job.