Course Content
AI Monitoring and Observability
3 sections · 7 lessons
Drift Detection and Anomaly Monitoring
A delivery platform runs a model that predicts how long an order will take. One input is distance_km. In April a logistics partner rewrites their integration and — without telling anyone, because to them it was a formatting detail — starts sending distances in miles.
Nothing breaks. Miles are perfectly good floating-point numbers. The schema validator passes: the field is present, numeric, positive, and inside the allowed range of 0 to 60. The model happily consumes them. The mean value of the feature moves from 6.4 to 3.98, because 6.4 km is 3.98 miles, and every estimated delivery time comes out roughly 38% too short for the affected partner.
Customer complaints climb for six weeks before anyone connects them to a model. There was no error to catch, no exception to log, no null to flag. The only visible fingerprint of the entire incident was that one number's distribution moved — and nobody was watching distributions.
That is drift monitoring in a sentence: watching the shape of your data and your outputs, so that changes with no error signature still produce a signal. It is entirely a statistics problem, and it is a statistics problem where the obvious approaches fail in specific, learnable ways.
Four different things people call "drift"
The word gets used for several distinct phenomena that have different causes, different detection methods, and different responses. Confusing them is the first mistake.
| What moves | Formally | Typical cause | Detect with | Response | |
|---|---|---|---|---|---|
| Data drift (covariate shift) | The inputs | P(X) changes, P(Y∣X) stays the same | New user segment, new device, unit change, upstream bug, seasonality | Per-feature distribution distance vs a baseline | Often none — the model may still be right. Investigate, do not auto-retrain |
| Concept drift | The relationship | P(Y∣X) changes | Fraudsters adapt, a policy changes, meaning of a word shifts | Accuracy on delayed labels; proxy quality signals | Retraining is genuinely required — old labels are now wrong |
| Prediction drift | The outputs | P(Y^) changes | Downstream of either of the above, or a model swap | Prediction histogram, class mix, mean confidence, output length | Cheap early warning — fires before labels arrive |
| Upstream/schema drift | The data contract | Not a distribution change at all | Renamed field, changed unit, new enum value, null-rate jump | Null rate, range, cardinality, type and enum checks | Fix the pipeline. Retraining on corrupted data makes it permanent |
The delivery example is the fourth row masquerading as the first. This distinction matters enormously in practice, because the automated response to "data drift detected" is often "retrain on recent data" — and retraining on miles-labelled-as-kilometres bakes the bug into the weights, where it will survive the pipeline fix.
Data drift means your inputs changed. Concept drift means the truth changed. Only the second one is a reason to retrain, and telling them apart requires labels — which is why label acquisition, not detection, is usually the bottleneck.
Measuring how far a distribution has moved: PSI
The Population Stability Index is the workhorse. Bin the feature, compare the proportion of mass in each bin now against a reference, and sum a symmetric divergence term:
where ei is the reference (expected) proportion in bin i and ai is the current (actual) proportion. Both are proportions, so they sum to 1.
Worked, on prompt token length binned into five quantile buckets of the reference:
| Bin (tokens) | Reference ei | Current ai | ai−ei | ln(ai/ei) | Contribution |
|---|---|---|---|---|---|
| 0–200 | 0.10 | 0.05 | −0.05 | −0.6931 | 0.03466 |
| 201–500 | 0.25 | 0.15 | −0.10 | −0.5108 | 0.05108 |
| 501–900 | 0.30 | 0.28 | −0.02 | −0.0690 | 0.00138 |
| 901–1500 | 0.25 | 0.32 | +0.07 | +0.2469 | 0.01728 |
| 1501+ | 0.10 | 0.20 | +0.10 | +0.6931 | 0.06931 |
| PSI | 0.1737 |
The conventional bands, which come from credit-risk practice and are worth treating as starting points rather than laws:
| PSI | Reading | Action |
|---|---|---|
| < 0.10 | No meaningful shift | Nothing |
| 0.10 – 0.25 | Moderate shift | Investigate; check whether accuracy has moved |
| > 0.25 | Major shift | Treat as an incident; likely retrain or pipeline fix |
So 0.174 says: something real happened to prompt lengths — the short-prompt bucket halved and the long-prompt bucket doubled — but it is not yet catastrophic. Because PSI is a sum over bins, you also get the diagnosis for free: the two largest contributions (0.0693 and 0.0511) point straight at the 1501+ and 201–500 buckets.
Two ways PSI is computed wrongly
Bins recomputed on the current data. If you take quantile bins of today's data instead of the reference's fixed edges, each bin will contain roughly 1/B of today's mass by construction, and PSI will be near zero no matter what happened. The reference bin edges must be computed once, stored, and reused. This is the single most common implementation bug, and it produces a detector that never fires.
Empty bins. If a bin has zero current mass, ln(0)=−∞. The usual patch is to substitute a small epsilon, but watch what that does. With a reference proportion of 0.10 and an epsilon of 0.0001:
(0.0001 - 0.10) x ln(0.0001 / 0.10) = (-0.0999) x ln(0.001) = (-0.0999) x (-6.9078) = 0.6901One empty bin contributes 0.69 on its own — nearly three times the "major shift" threshold — and if you had picked an epsilon of 0.00001 instead it would have contributed 0.92. The answer is dominated by a number you chose arbitrarily. The fix is not a better epsilon; it is to use 10 quantile bins from the reference so every bin starts with about 10% of the mass, and to require a minimum sample size (a few hundred at least) before computing PSI at all.
1import numpy as np23def fit_bins(reference: np.ndarray, n_bins: int = 10):4 """Compute edges ONCE on the reference. Store them alongside the model."""5 qs = np.linspace(0, 100, n_bins + 1)6 edges = np.unique(np.percentile(reference, qs))7 edges[0], edges[-1] = -np.inf, np.inf8 expected = np.histogram(reference, bins=edges)[0] / len(reference)9 return edges, expected1011def psi(current: np.ndarray, edges, expected, floor: float = 0.005):12 if len(current) < 200:13 return None # too few samples to be meaningful14 actual = np.histogram(current, bins=edges)[0] / len(current)15 a = np.clip(actual, floor, None) # floor, not epsilon: 0.5% not 0.01%16 e = np.clip(expected, floor, None)17 a, e = a / a.sum(), e / e.sum()18 return float(np.sum((a - e) * np.log(a / e)))The p-value trap, and why effect sizes win
The instinct of anyone with statistical training is to run a hypothesis test. For a continuous feature that means the two-sample Kolmogorov–Smirnov test, whose statistic is the largest vertical gap between the two empirical cumulative distributions:
and whose critical value at the 5% level is c(α)(n+m)/(nm) with c(0.05)=1.358. Work out what that means at two sample sizes:
n = m = 1,000 sqrt(2,000 / 1,000,000) = 0.04472 x 1.358 = 0.0607n = m = 50,000 sqrt(100,000 / 2.5e9) = 0.006325 x 1.358 = 0.0086At 50,000 samples per side, a D of 0.01 is "statistically significant". D = 0.01 means the two cumulative distributions never differ by more than one percentage point anywhere. Nobody in the world would call that a problem — and your alert just fired.
The same thing happens with categorical features and chi-squared, and there the algebra makes the reason explicit. Take a four-category feature over 2,000 current samples:
| Category | Reference share | Expected count | Observed count | (O−E)2/E |
|---|---|---|---|---|
| A | 0.40 | 800 | 700 | 12.500 |
| B | 0.30 | 600 | 560 | 2.667 |
| C | 0.20 | 400 | 460 | 9.000 |
| D | 0.10 | 200 | 280 | 32.000 |
| χ² (3 d.f.) | 56.17 |
The critical value for 3 degrees of freedom at the 5% level is 7.815. Our 56.17 blows past it; the p-value is roughly 4×10−12. Screaming significance.
Now compute PSI on exactly the same data. The observed shares are 0.35, 0.28, 0.23, 0.14:
(0.35-0.40) x ln(0.35/0.40) = (-0.05) x (-0.13353) = 0.006677(0.28-0.30) x ln(0.28/0.30) = (-0.02) x (-0.06899) = 0.001380(0.23-0.20) x ln(0.23/0.20) = (+0.03) x (+0.13976) = 0.004193(0.14-0.10) x ln(0.14/0.10) = (+0.04) x (+0.33647) = 0.013459 PSI = 0.0257PSI = 0.026, comfortably inside "no meaningful shift". Chi-squared says the drift is one of the most certain facts in the universe; PSI says it barely registers. Both are correct, because they answer different questions.
The reason falls out of the algebra. For small relative differences, ln(a/e)≈(a−e)/e, so PSI ≈∑(ai−ei)2/ei — which is exactly χ2/n. Check it: 56.17/2000=0.0281, against a PSI of 0.0257. Same quantity, up to the approximation.
Chi-squared is n times an effect size. Double your traffic and chi-squared doubles while nothing about your data changed. A drift monitor built on p-values does not measure drift — it measures how much data you have.
The practical rule: alert on effect sizes with magnitude thresholds, not on p-values. Use PSI, or the KS statistic D itself with a threshold like 0.1, or Jensen–Shannon distance with a threshold like 0.1. Keep the p-value if you like, as a tiebreaker for small samples — but never as the trigger.
| Measure | Data type | Range | Sample-size sensitive? | Sensible threshold |
|---|---|---|---|---|
| PSI | Binned numeric or categorical | 0 to ∞ | No | 0.1 warn / 0.25 act |
| KS statistic D | Continuous, univariate | 0 to 1 | No (the p-value is) | 0.1 warn / 0.2 act |
| Jensen–Shannon distance | Any binned distribution | 0 to 1, symmetric, bounded | No | 0.1 warn / 0.2 act |
| Wasserstein distance | Continuous; respects ordering | 0 to ∞, in feature units | No | Set relative to feature s.d. |
| Chi-squared / KS p-value | Any | 0 to 1 | Yes, severely | Do not alert on it |
Choosing the reference window
Every drift number is a comparison, so it is only as good as the thing you compare against. Two choices, and the wrong one silently disables the whole system.
| Fixed reference | Rolling reference | |
|---|---|---|
| What it is | Training data, or a frozen snapshot of known-good production traffic | The trailing 7 or 30 days |
| Detects | Total divergence from the world the model was fitted to | Sudden changes |
| Misses | Nothing, but flags benign long-term change too | Slow degradation — it chases the drift |
| Use for | Deciding whether the model is still valid | Incident detection: did something break today? |
The rolling-reference failure is worth naming because it is seductive. If the feature moves 1% a day and your reference is the last 7 days, then each day's comparison sees a 1% difference and reports PSI near zero — every single day, forever, while the feature drifts 30% away from training over a month. The detector is working exactly as specified and is completely blind. This is the boiling-frog failure, and the fix is to run both: a rolling reference for "did something break in the last hour", and a fixed training-data reference for "is this model still operating in the world it was built for".
Multivariate drift: what per-feature checks cannot see
Check each feature independently and you will miss any change that preserves the marginals but breaks the joint distribution. A toy version: suppose age and product_tier each keep their exact historical distributions, but the correlation between them inverts — young users used to buy the cheap tier and now buy the premium one. Every univariate test passes. The model, which learned the old relationship, is now systematically wrong.
The most practical detector is the domain classifier (also called adversarial validation). Label reference rows 0 and current rows 1, mix them, and train a gradient-boosted tree to tell them apart, scoring by cross-validated AUC.
1import numpy as np2from sklearn.ensemble import HistGradientBoostingClassifier3from sklearn.model_selection import cross_val_score4from sklearn.inspection import permutation_importance56def drift_auc(reference: np.ndarray, current: np.ndarray, seed: int = 0):7 n = min(len(reference), len(current)) # balance, or AUC is confounded8 rng = np.random.default_rng(seed)9 ref = reference[rng.choice(len(reference), n, replace=False)]10 cur = current[rng.choice(len(current), n, replace=False)]11 X = np.vstack([ref, cur])12 y = np.r_[np.zeros(n), np.ones(n)]13 clf = HistGradientBoostingClassifier(max_depth=4, max_iter=150)14 auc = cross_val_score(clf, X, y, cv=5, scoring="roc_auc").mean()15 clf.fit(X, y)16 imp = permutation_importance(clf, X, y, n_repeats=5, random_state=seed)17 return auc, imp.importances_meanReading the result:
- AUC ≈ 0.50 — the classifier cannot distinguish the two samples. The joint distributions are, as far as a strong learner can tell, the same.
- AUC 0.55 – 0.65 — mild separability. Common and often benign.
- AUC > 0.70 — the two periods are materially different data. Take it seriously.
- AUC > 0.90 — usually a pipeline bug or a leaked timestamp-like feature, not organic drift.
The permutation importances are the reason this technique earns its keep: they rank which features the classifier used to separate the periods, which is a direct list of what changed, including interactions no univariate test would surface.
Apply the effect-size discipline here too. Under the null hypothesis of no drift, the variance of AUC is approximately (n+m+1)/(12nm). With 5,000 rows a side:
Var = 10,001 / (12 x 5,000 x 5,000) = 10,001 / 300,000,000 = 3.334e-5sd = 0.00577an AUC of 0.52 sits (0.52 - 0.50) / 0.00577 = 3.5 standard errors above chanceStatistically significant, operationally meaningless. Threshold on the AUC value, not on its significance.
Drift in text and embeddings
For LLM inputs there are no tabular features to bin, so the same ideas move into embedding space:
- Cheap scalar proxies first. Prompt length in tokens, detected language mix, share of prompts containing code fences or URLs, share matching known templates, mean cosine similarity to the reference centroid. These are one-dimensional and go straight into PSI. They catch a surprising fraction of real incidents.
- Cluster-share drift. Cluster the reference embeddings once into, say, 50 clusters. Assign each new prompt to its nearest centroid and compute PSI over the cluster-share vector. This gives you an interpretable answer: "cluster 17 — password-reset questions — went from 3% of traffic to 19%".
- Maximum Mean Discrepancy for a proper multivariate test on raw embeddings, with a permutation test for the threshold. More rigorous, and much harder to explain to anyone when it fires.
Anomaly detection: spikes, not shifts
Drift is a slow change in shape. An anomaly is a single point that does not belong. The default tool is the z-score, (x−μ)/σ, and its default implementation has a fatal flaw: the mean and standard deviation are themselves destroyed by the anomalies you are trying to find.
Concretely. You monitor mean output length per hour. Take a 24-hour baseline where 23 hours scatter around 340 tokens with a standard deviation of 25, and one hour — an incident last Tuesday — hit 1,200.
mean = (23 x 340 + 1200) / 24 = 9,020 / 24 = 375.83sum of squared deviations from the 23 normal hours : 22 x 25^2 = 13,750 their offset from the new mean : 23 x (340 - 375.83)^2 = 29,533 the outlier : (1200 - 375.83)^2 = 679,251 total = 722,534sample variance = 722,534 / 23 = 31,415 sd = 177.2Now a genuinely anomalous hour arrives at 448 tokens:
classical z = (448 - 375.83) / 177.2 = 0.41 -> silentA z of 0.41 is utterly ordinary. One prior incident sitting in the baseline inflated the standard deviation by a factor of seven and blinded the detector.
The robust version replaces the mean with the median and the standard deviation with the median absolute deviation, MAD =median(∣xi−median(x)∣), rescaled by 1.4826 so that it estimates σ for normally distributed data. The median has a breakdown point of 50%: one outlier in 24 moves it essentially not at all. Here the median stays 340 and the MAD is about 17.
sigma_hat = 1.4826 x 17 = 25.20robust z = (448 - 340) / 25.20 = 4.29 -> fires clearlySame data, same anomaly: 0.41 versus 4.29. Use the median and MAD by default for any detector whose baseline is computed from production data, because production data contains incidents.
1import numpy as np23def robust_z(x: float, baseline: np.ndarray) -> float:4 med = np.median(baseline)5 mad = np.median(np.abs(baseline - med))6 if mad == 0: # degenerate: near-constant series7 mad = np.mean(np.abs(baseline - med)) or 1e-98 return float((x - med) / (1.4826 * mad))Seasonality will fake every anomaly you have
Traffic at 03:00 on a Sunday is not comparable to traffic at 14:00 on a Tuesday. A detector that pools all hours will flag every weekday morning ramp as an anomaly and miss a genuine Sunday outage entirely, because the Sunday baseline was diluted by weekday volume. Two fixes, in increasing order of effort:
- Same-slot comparison. Compare Tuesday 14:00 against the last eight Tuesdays at 14:00. Crude, cheap, and removes both daily and weekly seasonality at once.
- Decompose, then detect on the residual. Fit trend plus daily plus weekly components (STL, or a seasonal-naive baseline), and run the robust z-score on what is left over. Better resolution, but needs several weeks of history before it works at all.
Performance drift, and the label-lag problem
Everything above is a proxy. The metric you actually care about is whether the model is still right, and that requires ground truth — which arrives late, if at all.
| System | Label source | Typical lag | Coverage |
|---|---|---|---|
| Ad click prediction | The click, or its absence after a timeout | Minutes | Effectively 100% |
| Ticket triage | Agent re-assigns the ticket | Hours | Only wrong ones are visible — biased |
| Credit default | Missed payment | 60–90 days | Only for approved applicants — biased |
| LLM answer quality | Human review or an LLM judge | Days | Whatever you sample |
Two consequences follow. First, with a 90-day lag your accuracy dashboard describes a model that was serving traffic three months ago; it is an audit record, not a monitor. Second, several of these label channels are systematically biased — the credit model only ever gets labels for applicants it approved, so its measured performance is conditioned on its own decisions.
So you build a ladder of signals, ordered by how fast they arrive and how much you can trust them:
- Immediate, weak: input drift (PSI), prediction-mix drift, mean confidence, output length, refusal rate, retry rate. Minutes. Correlated with quality, not equal to it.
- Fast, indirect: user behaviour the model did not choose — thumbs-down rate, edit rate on generated text, escalation rate, session abandonment, "talk to a human" clicks. Hours.
- Slow, trustworthy: a stratified human-labelled sample. A few hundred examples a week, sampled independently of the model's own outputs, is enough to detect a several-point accuracy move and is the only number in the stack that cannot be gamed by a feedback loop.
- Slowest, definitive: real outcome labels, when they exist.
Every fast quality signal is a proxy, and every proxy eventually decouples from the thing it proxies. Keep one slow, expensive, human-labelled channel alive permanently — it is the only thing that tells you when your fast signals started lying.
What this means when you build one
A drift monitor is a scheduled job, a stored artefact, and a small number of thresholds. The parts that decide whether it works are unglamorous.
Version and store the reference alongside the model. Bin edges, expected proportions, per-feature medians and MADs, prediction-mix baseline, embedding centroids — all of it, in the model registry, tagged with the model version. If your reference lives in a notebook, your drift numbers are not reproducible and nobody will trust them during an incident.
Count your checks before you set thresholds. Suppose you run a KS test on 30 features every hour at the 5% level. On perfectly stable data:
30 features x 24 hours = 720 tests/day720 x 0.05 = 36 false positives per day = 252 per weekTwo hundred and fifty-two alerts a week from a system where nothing is wrong. A Bonferroni correction to α/30=0.00167 brings it to 1.2 a day, which is still 8 a week and comes at a real cost in detection power. The better answer is to abandon per-test significance entirely: use magnitude thresholds on effect sizes, require the breach to persist across two or more consecutive check intervals, and monitor a curated set of features that actually matter rather than all 300 columns in the table.
Route drift signals to a human, not to a retraining job. The wiring "PSI > 0.25 triggers retrain" is the most dangerous pattern in this whole area, because the most common cause of a large PSI is an upstream data bug — and retraining on corrupted data converts a reversible pipeline problem into a model that has learned the corruption. Drift opens a ticket with the top contributing features attached. A human decides whether the correct response is a retrain, a pipeline fix, or nothing at all.
Had the delivery platform stored a fixed reference for distance_km and computed PSI daily against it, the miles-for-kilometres switch would have produced a PSI well above 0.25 on day one, with the diagnosis attached: the two lowest bins gained enormous mass and the top bins emptied. Six weeks of short delivery estimates would have been six hours.