Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you monitor data drift in LLM systems over time?
What you need to know
The four kinds of drift
| Drift | What changes | Detector |
|---|---|---|
| Input | What users ask, in what language, how long | Topic-mix PSI, embedding-distance to baseline, share of queries matching no route |
| Corpus | Source documents move on; index goes stale | Retrieval score distribution, hit rate, age of newest document |
| Model | Provider updates an alias, or you upgrade | Pinned snapshots; scheduled golden-set runs |
| Output | Answers get longer, refuse more, use tools differently | Length, refusal, citation and tool-mix distributions |
Population Stability Index (PSI)
PSI compares two distributions of categories (for example the share of each topic this week against the baseline). A common rule of thumb: under 0.1 is stable, 0.1–0.25 is a moderate shift, over 0.25 is a major shift.
1import math23def psi(expected: dict, actual: dict, eps=1e-4) -> float:4 """Population Stability Index between two category distributions."""5 total = 0.06 for k in expected.keys() | actual.keys():7 e = max(expected.get(k, 0.0), eps)8 a = max(actual.get(k, 0.0), eps)9 total += (a - e) * math.log(a / e)10 return total1112baseline = {"status": 0.38, "eligibility": 0.30, "grievance": 0.12, "other": 0.20}13this_week = {"status": 0.25, "eligibility": 0.22, "grievance": 0.10, "other": 0.43}14print(round(psi(baseline, this_week), 3)) # 0.259The topics come from your router or a topic classifier. For raw text, you can instead cluster query embeddings and compute PSI over the clusters, or track the average distance of new queries to the baseline centroid.
From "changed" to "worse"
Drift detection is a smoke alarm, not a diagnosis. The practical setup is a nightly job that (1) computes drift scores, (2) runs the golden eval set against the pinned production config, and (3) scores a sampled slice of yesterday's live traffic with an LLM judge. Alert when any of them moves beyond tolerance, then read traces from the drifted segment.
A real-life example
A state government chatbot's weekly drift report shows topic PSI of 0.26: the "other" share jumped from 20% to 43%. Nothing is broken in the system metrics.
The team clusters the "other" queries and finds a new cluster: questions about a housing scheme announced on Monday. The retrieval index has no documents for it, so the bot answers "I don't have information about this scheme" — correct, but useless. Retrieval top-score for these questions averages 0.41 against 0.78 normally, confirming corpus drift.
The fix is a process, not a model change: scheme documents are now added to the index on the day of announcement, and the drift job's "new cluster" alert goes to the content team as well as engineering. A new router route for the housing scheme follows a week later.
Follow-up questions to expect
- "How do you detect a silent provider model update?" — Pin a dated snapshot; if you must use an alias, run the golden set daily and compare output length, refusal rate and scores.
- "What baseline do you compare against?" — A rolling window (for example the last four weeks), plus a fixed launch baseline for slow drift.
- "What do you do when drift is detected?" — Read traces from the drifted segment, add examples to the eval set, and fix the part that is stale — usually the corpus or routing, sometimes the prompt.