AI Monitoring and Observability

Drift Detection and Anomaly Monitoring


A delivery platform runs a model that predicts how long an order will take. One input is distance_km. In April a logistics partner rewrites their integration and — without telling anyone, because to them it was a formatting detail — starts sending distances in miles.

Nothing breaks. Miles are perfectly good floating-point numbers. The schema validator passes: the field is present, numeric, positive, and inside the allowed range of 0 to 60. The model happily consumes them. The mean value of the feature moves from 6.4 to 3.98, because 6.4 km is 3.98 miles, and every estimated delivery time comes out roughly 38% too short for the affected partner.

Customer complaints climb for six weeks before anyone connects them to a model. There was no error to catch, no exception to log, no null to flag. The only visible fingerprint of the entire incident was that one number's distribution moved — and nobody was watching distributions.

That is drift monitoring in a sentence: watching the shape of your data and your outputs, so that changes with no error signature still produce a signal. It is entirely a statistics problem, and it is a statistics problem where the obvious approaches fail in specific, learnable ways.

Kilometres reported as miles, seen through PSI18460.26334310.00328150.0811460.068620.044Reference pctApril pctPSI term0 to 2 km2 to 5 km5 to 10 km10 to 20 kmover 20 kmThe terms total 0.46, well past the 0.25 major-shift threshold, months before a single label arrived.
Dividing every distance by 1.61 crowds the whole population into the shortest bucket — one bin carries most of the signal.

Four different things people call "drift"

The word gets used for several distinct phenomena that have different causes, different detection methods, and different responses. Confusing them is the first mistake.

What movesFormallyTypical causeDetect withResponse
Data drift (covariate shift)The inputsP(X)P(X) changes, P(Y∣X)P(Y \mid X) stays the sameNew user segment, new device, unit change, upstream bug, seasonalityPer-feature distribution distance vs a baselineOften none — the model may still be right. Investigate, do not auto-retrain
Concept driftThe relationshipP(Y∣X)P(Y \mid X) changesFraudsters adapt, a policy changes, meaning of a word shiftsAccuracy on delayed labels; proxy quality signalsRetraining is genuinely required — old labels are now wrong
Prediction driftThe outputsP(Y^)P(\hat{Y}) changesDownstream of either of the above, or a model swapPrediction histogram, class mix, mean confidence, output lengthCheap early warning — fires before labels arrive
Upstream/schema driftThe data contractNot a distribution change at allRenamed field, changed unit, new enum value, null-rate jumpNull rate, range, cardinality, type and enum checksFix the pipeline. Retraining on corrupted data makes it permanent

The delivery example is the fourth row masquerading as the first. This distinction matters enormously in practice, because the automated response to "data drift detected" is often "retrain on recent data" — and retraining on miles-labelled-as-kilometres bakes the bug into the weights, where it will survive the pipeline fix.

Data drift means your inputs changed. Concept drift means the truth changed. Only the second one is a reason to retrain, and telling them apart requires labels — which is why label acquisition, not detection, is usually the bottleneck.

Measuring how far a distribution has moved: PSI

The Population Stability Index is the workhorse. Bin the feature, compare the proportion of mass in each bin now against a reference, and sum a symmetric divergence term:

PSI=∑i=1B(ai−ei)ln⁡ ⁣(aiei)\text{PSI} = \sum_{i=1}^{B} (a_i - e_i)\ln\!\left(\frac{a_i}{e_i}\right)

where eie_i is the reference (expected) proportion in bin ii and aia_i is the current (actual) proportion. Both are proportions, so they sum to 1.

Worked, on prompt token length binned into five quantile buckets of the reference:

Bin (tokens)Reference eie_iCurrent aia_iai−eia_i - e_iln⁡(ai/ei)\ln(a_i/e_i)Contribution
0–2000.100.05−0.05−0.69310.03466
201–5000.250.15−0.10−0.51080.05108
501–9000.300.28−0.02−0.06900.00138
901–15000.250.32+0.07+0.24690.01728
1501+0.100.20+0.10+0.69310.06931
PSI0.1737

The conventional bands, which come from credit-risk practice and are worth treating as starting points rather than laws:

PSIReadingAction
< 0.10No meaningful shiftNothing
0.10 – 0.25Moderate shiftInvestigate; check whether accuracy has moved
> 0.25Major shiftTreat as an incident; likely retrain or pipeline fix

So 0.174 says: something real happened to prompt lengths — the short-prompt bucket halved and the long-prompt bucket doubled — but it is not yet catastrophic. Because PSI is a sum over bins, you also get the diagnosis for free: the two largest contributions (0.0693 and 0.0511) point straight at the 1501+ and 201–500 buckets.

Two ways PSI is computed wrongly

Bins recomputed on the current data. If you take quantile bins of today's data instead of the reference's fixed edges, each bin will contain roughly 1/B of today's mass by construction, and PSI will be near zero no matter what happened. The reference bin edges must be computed once, stored, and reused. This is the single most common implementation bug, and it produces a detector that never fires.

Empty bins. If a bin has zero current mass, ln⁡(0)=−∞\ln(0) = -\infty. The usual patch is to substitute a small epsilon, but watch what that does. With a reference proportion of 0.10 and an epsilon of 0.0001:

Text
(0.0001 - 0.10) x ln(0.0001 / 0.10)  = (-0.0999) x ln(0.001)  = (-0.0999) x (-6.9078)  = 0.6901

One empty bin contributes 0.69 on its own — nearly three times the "major shift" threshold — and if you had picked an epsilon of 0.00001 instead it would have contributed 0.92. The answer is dominated by a number you chose arbitrarily. The fix is not a better epsilon; it is to use 10 quantile bins from the reference so every bin starts with about 10% of the mass, and to require a minimum sample size (a few hundred at least) before computing PSI at all.

Python
import numpy as npdef fit_bins(reference: np.ndarray, n_bins: int = 10):    """Compute edges ONCE on the reference. Store them alongside the model."""    qs = np.linspace(0, 100, n_bins + 1)    edges = np.unique(np.percentile(reference, qs))    edges[0], edges[-1] = -np.inf, np.inf    expected = np.histogram(reference, bins=edges)[0] / len(reference)    return edges, expecteddef psi(current: np.ndarray, edges, expected, floor: float = 0.005):    if len(current) < 200:        return None                       # too few samples to be meaningful    actual = np.histogram(current, bins=edges)[0] / len(current)    a = np.clip(actual, floor, None)      # floor, not epsilon: 0.5% not 0.01%    e = np.clip(expected, floor, None)    a, e = a / a.sum(), e / e.sum()    return float(np.sum((a - e) * np.log(a / e)))

The p-value trap, and why effect sizes win

The instinct of anyone with statistical training is to run a hypothesis test. For a continuous feature that means the two-sample Kolmogorov–Smirnov test, whose statistic is the largest vertical gap between the two empirical cumulative distributions:

D=sup⁡x∣Fref(x)−Fcur(x)∣D = \sup_x \left| F_{\text{ref}}(x) - F_{\text{cur}}(x) \right|

and whose critical value at the 5% level is c(α)(n+m)/(nm)c(\alpha)\sqrt{(n+m)/(nm)} with c(0.05)=1.358c(0.05) = 1.358. Work out what that means at two sample sizes:

Text
n = m = 1,000     sqrt(2,000 / 1,000,000)   = 0.04472   x 1.358 = 0.0607n = m = 50,000    sqrt(100,000 / 2.5e9)     = 0.006325  x 1.358 = 0.0086

At 50,000 samples per side, a D of 0.01 is "statistically significant". D = 0.01 means the two cumulative distributions never differ by more than one percentage point anywhere. Nobody in the world would call that a problem — and your alert just fired.

The same thing happens with categorical features and chi-squared, and there the algebra makes the reason explicit. Take a four-category feature over 2,000 current samples:

CategoryReference shareExpected countObserved count(O−E)2/E(O-E)^2/E
A0.4080070012.500
B0.306005602.667
C0.204004609.000
D0.1020028032.000
χ² (3 d.f.)56.17

The critical value for 3 degrees of freedom at the 5% level is 7.815. Our 56.17 blows past it; the p-value is roughly 4×10−124 \times 10^{-12}. Screaming significance.

Now compute PSI on exactly the same data. The observed shares are 0.35, 0.28, 0.23, 0.14:

Text
(0.35-0.40) x ln(0.35/0.40) = (-0.05) x (-0.13353) = 0.006677(0.28-0.30) x ln(0.28/0.30) = (-0.02) x (-0.06899) = 0.001380(0.23-0.20) x ln(0.23/0.20) = (+0.03) x (+0.13976) = 0.004193(0.14-0.10) x ln(0.14/0.10) = (+0.04) x (+0.33647) = 0.013459                                              PSI  = 0.0257

PSI = 0.026, comfortably inside "no meaningful shift". Chi-squared says the drift is one of the most certain facts in the universe; PSI says it barely registers. Both are correct, because they answer different questions.

The reason falls out of the algebra. For small relative differences, ln⁡(a/e)≈(a−e)/e\ln(a/e) \approx (a-e)/e, so PSI ≈∑(ai−ei)2/ei\approx \sum (a_i-e_i)^2/e_i — which is exactly χ2/n\chi^2/n. Check it: 56.17/2000=0.028156.17 / 2000 = 0.0281, against a PSI of 0.0257. Same quantity, up to the approximation.

Chi-squared is n times an effect size. Double your traffic and chi-squared doubles while nothing about your data changed. A drift monitor built on p-values does not measure drift — it measures how much data you have.

The practical rule: alert on effect sizes with magnitude thresholds, not on p-values. Use PSI, or the KS statistic D itself with a threshold like 0.1, or Jensen–Shannon distance with a threshold like 0.1. Keep the p-value if you like, as a tiebreaker for small samples — but never as the trigger.

MeasureData typeRangeSample-size sensitive?Sensible threshold
PSIBinned numeric or categorical0 to ∞No0.1 warn / 0.25 act
KS statistic DDContinuous, univariate0 to 1No (the p-value is)0.1 warn / 0.2 act
Jensen–Shannon distanceAny binned distribution0 to 1, symmetric, boundedNo0.1 warn / 0.2 act
Wasserstein distanceContinuous; respects ordering0 to ∞, in feature unitsNoSet relative to feature s.d.
Chi-squared / KS p-valueAny0 to 1Yes, severelyDo not alert on it

Choosing the reference window

Every drift number is a comparison, so it is only as good as the thing you compare against. Two choices, and the wrong one silently disables the whole system.

Fixed referenceRolling reference
What it isTraining data, or a frozen snapshot of known-good production trafficThe trailing 7 or 30 days
DetectsTotal divergence from the world the model was fitted toSudden changes
MissesNothing, but flags benign long-term change tooSlow degradation — it chases the drift
Use forDeciding whether the model is still validIncident detection: did something break today?

The rolling-reference failure is worth naming because it is seductive. If the feature moves 1% a day and your reference is the last 7 days, then each day's comparison sees a 1% difference and reports PSI near zero — every single day, forever, while the feature drifts 30% away from training over a month. The detector is working exactly as specified and is completely blind. This is the boiling-frog failure, and the fix is to run both: a rolling reference for "did something break in the last hour", and a fixed training-data reference for "is this model still operating in the world it was built for".

Multivariate drift: what per-feature checks cannot see

Check each feature independently and you will miss any change that preserves the marginals but breaks the joint distribution. A toy version: suppose age and product_tier each keep their exact historical distributions, but the correlation between them inverts — young users used to buy the cheap tier and now buy the premium one. Every univariate test passes. The model, which learned the old relationship, is now systematically wrong.

The most practical detector is the domain classifier (also called adversarial validation). Label reference rows 0 and current rows 1, mix them, and train a gradient-boosted tree to tell them apart, scoring by cross-validated AUC.

Python
import numpy as npfrom sklearn.ensemble import HistGradientBoostingClassifierfrom sklearn.model_selection import cross_val_scorefrom sklearn.inspection import permutation_importancedef drift_auc(reference: np.ndarray, current: np.ndarray, seed: int = 0):    n = min(len(reference), len(current))          # balance, or AUC is confounded    rng = np.random.default_rng(seed)    ref = reference[rng.choice(len(reference), n, replace=False)]    cur = current[rng.choice(len(current), n, replace=False)]    X = np.vstack([ref, cur])    y = np.r_[np.zeros(n), np.ones(n)]    clf = HistGradientBoostingClassifier(max_depth=4, max_iter=150)    auc = cross_val_score(clf, X, y, cv=5, scoring="roc_auc").mean()    clf.fit(X, y)    imp = permutation_importance(clf, X, y, n_repeats=5, random_state=seed)    return auc, imp.importances_mean

Reading the result:

  • AUC ≈ 0.50 — the classifier cannot distinguish the two samples. The joint distributions are, as far as a strong learner can tell, the same.
  • AUC 0.55 – 0.65 — mild separability. Common and often benign.
  • AUC > 0.70 — the two periods are materially different data. Take it seriously.
  • AUC > 0.90 — usually a pipeline bug or a leaked timestamp-like feature, not organic drift.

The permutation importances are the reason this technique earns its keep: they rank which features the classifier used to separate the periods, which is a direct list of what changed, including interactions no univariate test would surface.

Apply the effect-size discipline here too. Under the null hypothesis of no drift, the variance of AUC is approximately (n+m+1)/(12nm)(n+m+1)/(12nm). With 5,000 rows a side:

Text
Var = 10,001 / (12 x 5,000 x 5,000) = 10,001 / 300,000,000 = 3.334e-5sd  = 0.00577an AUC of 0.52 sits (0.52 - 0.50) / 0.00577 = 3.5 standard errors above chance

Statistically significant, operationally meaningless. Threshold on the AUC value, not on its significance.

Drift in text and embeddings

For LLM inputs there are no tabular features to bin, so the same ideas move into embedding space:

  • Cheap scalar proxies first. Prompt length in tokens, detected language mix, share of prompts containing code fences or URLs, share matching known templates, mean cosine similarity to the reference centroid. These are one-dimensional and go straight into PSI. They catch a surprising fraction of real incidents.
  • Cluster-share drift. Cluster the reference embeddings once into, say, 50 clusters. Assign each new prompt to its nearest centroid and compute PSI over the cluster-share vector. This gives you an interpretable answer: "cluster 17 — password-reset questions — went from 3% of traffic to 19%".
  • Maximum Mean Discrepancy for a proper multivariate test on raw embeddings, with a permutation test for the threshold. More rigorous, and much harder to explain to anyone when it fires.

Anomaly detection: spikes, not shifts

Drift is a slow change in shape. An anomaly is a single point that does not belong. The default tool is the z-score, (x−μ)/σ(x - \mu)/\sigma, and its default implementation has a fatal flaw: the mean and standard deviation are themselves destroyed by the anomalies you are trying to find.

Concretely. You monitor mean output length per hour. Take a 24-hour baseline where 23 hours scatter around 340 tokens with a standard deviation of 25, and one hour — an incident last Tuesday — hit 1,200.

Text
mean = (23 x 340 + 1200) / 24 = 9,020 / 24 = 375.83sum of squared deviations  from the 23 normal hours : 22 x 25^2                    =  13,750  their offset from the new mean : 23 x (340 - 375.83)^2  =  29,533  the outlier : (1200 - 375.83)^2                         = 679,251                                                    total = 722,534sample variance = 722,534 / 23 = 31,415      sd = 177.2

Now a genuinely anomalous hour arrives at 448 tokens:

Text
classical z = (448 - 375.83) / 177.2 = 0.41      ->  silent

A z of 0.41 is utterly ordinary. One prior incident sitting in the baseline inflated the standard deviation by a factor of seven and blinded the detector.

The robust version replaces the mean with the median and the standard deviation with the median absolute deviation, MAD =median(∣xi−median(x)∣)= \text{median}(|x_i - \text{median}(x)|), rescaled by 1.4826 so that it estimates σ\sigma for normally distributed data. The median has a breakdown point of 50%: one outlier in 24 moves it essentially not at all. Here the median stays 340 and the MAD is about 17.

Text
sigma_hat = 1.4826 x 17 = 25.20robust z  = (448 - 340) / 25.20 = 4.29           ->  fires clearly

Same data, same anomaly: 0.41 versus 4.29. Use the median and MAD by default for any detector whose baseline is computed from production data, because production data contains incidents.

Python
import numpy as npdef robust_z(x: float, baseline: np.ndarray) -> float:    med = np.median(baseline)    mad = np.median(np.abs(baseline - med))    if mad == 0:                       # degenerate: near-constant series        mad = np.mean(np.abs(baseline - med)) or 1e-9    return float((x - med) / (1.4826 * mad))

Seasonality will fake every anomaly you have

Traffic at 03:00 on a Sunday is not comparable to traffic at 14:00 on a Tuesday. A detector that pools all hours will flag every weekday morning ramp as an anomaly and miss a genuine Sunday outage entirely, because the Sunday baseline was diluted by weekday volume. Two fixes, in increasing order of effort:

  1. Same-slot comparison. Compare Tuesday 14:00 against the last eight Tuesdays at 14:00. Crude, cheap, and removes both daily and weekly seasonality at once.
  2. Decompose, then detect on the residual. Fit trend plus daily plus weekly components (STL, or a seasonal-naive baseline), and run the robust z-score on what is left over. Better resolution, but needs several weeks of history before it works at all.

Performance drift, and the label-lag problem

Everything above is a proxy. The metric you actually care about is whether the model is still right, and that requires ground truth — which arrives late, if at all.

SystemLabel sourceTypical lagCoverage
Ad click predictionThe click, or its absence after a timeoutMinutesEffectively 100%
Ticket triageAgent re-assigns the ticketHoursOnly wrong ones are visible — biased
Credit defaultMissed payment60–90 daysOnly for approved applicants — biased
LLM answer qualityHuman review or an LLM judgeDaysWhatever you sample

Two consequences follow. First, with a 90-day lag your accuracy dashboard describes a model that was serving traffic three months ago; it is an audit record, not a monitor. Second, several of these label channels are systematically biased — the credit model only ever gets labels for applicants it approved, so its measured performance is conditioned on its own decisions.

So you build a ladder of signals, ordered by how fast they arrive and how much you can trust them:

  1. Immediate, weak: input drift (PSI), prediction-mix drift, mean confidence, output length, refusal rate, retry rate. Minutes. Correlated with quality, not equal to it.
  2. Fast, indirect: user behaviour the model did not choose — thumbs-down rate, edit rate on generated text, escalation rate, session abandonment, "talk to a human" clicks. Hours.
  3. Slow, trustworthy: a stratified human-labelled sample. A few hundred examples a week, sampled independently of the model's own outputs, is enough to detect a several-point accuracy move and is the only number in the stack that cannot be gamed by a feedback loop.
  4. Slowest, definitive: real outcome labels, when they exist.

Every fast quality signal is a proxy, and every proxy eventually decouples from the thing it proxies. Keep one slow, expensive, human-labelled channel alive permanently — it is the only thing that tells you when your fast signals started lying.

What this means when you build one

A drift monitor is a scheduled job, a stored artefact, and a small number of thresholds. The parts that decide whether it works are unglamorous.

Version and store the reference alongside the model. Bin edges, expected proportions, per-feature medians and MADs, prediction-mix baseline, embedding centroids — all of it, in the model registry, tagged with the model version. If your reference lives in a notebook, your drift numbers are not reproducible and nobody will trust them during an incident.

Count your checks before you set thresholds. Suppose you run a KS test on 30 features every hour at the 5% level. On perfectly stable data:

Text
30 features x 24 hours = 720 tests/day720 x 0.05 = 36 false positives per day = 252 per week

Two hundred and fifty-two alerts a week from a system where nothing is wrong. A Bonferroni correction to α/30=0.00167\alpha/30 = 0.00167 brings it to 1.2 a day, which is still 8 a week and comes at a real cost in detection power. The better answer is to abandon per-test significance entirely: use magnitude thresholds on effect sizes, require the breach to persist across two or more consecutive check intervals, and monitor a curated set of features that actually matter rather than all 300 columns in the table.

Route drift signals to a human, not to a retraining job. The wiring "PSI > 0.25 triggers retrain" is the most dangerous pattern in this whole area, because the most common cause of a large PSI is an upstream data bug — and retraining on corrupted data converts a reversible pipeline problem into a model that has learned the corruption. Drift opens a ticket with the top contributing features attached. A human decides whether the correct response is a retrain, a pipeline fix, or nothing at all.

Had the delivery platform stored a fixed reference for distance_km and computed PSI daily against it, the miles-for-kilometres switch would have produced a PSI well above 0.25 on day one, with the diagnosis attached: the two lowest bins gained enormous mass and the top bins emptied. Six weeks of short delivery estimates would have been six hours.