LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you monitor data drift in LLM systems over time?


Four sources of drift, four detectorsDriftInput: topic-mix PSI 0.26Corpus: topscore 0.41, was 0.78Model: pinnedsnapshot, nightly evalOutput: length,refusals, tool mix
The housing-scheme spike was input and corpus drift at once — the fix was new documents, not a new model.

What you need to know

The four kinds of drift

DriftWhat changesDetector
InputWhat users ask, in what language, how longTopic-mix PSI, embedding-distance to baseline, share of queries matching no route
CorpusSource documents move on; index goes staleRetrieval score distribution, hit rate, age of newest document
ModelProvider updates an alias, or you upgradePinned snapshots; scheduled golden-set runs
OutputAnswers get longer, refuse more, use tools differentlyLength, refusal, citation and tool-mix distributions

Population Stability Index (PSI)

PSI compares two distributions of categories (for example the share of each topic this week against the baseline). A common rule of thumb: under 0.1 is stable, 0.1–0.25 is a moderate shift, over 0.25 is a major shift.

Python
import mathdef psi(expected: dict, actual: dict, eps=1e-4) -> float:    """Population Stability Index between two category distributions."""    total = 0.0    for k in expected.keys() | actual.keys():        e = max(expected.get(k, 0.0), eps)        a = max(actual.get(k, 0.0), eps)        total += (a - e) * math.log(a / e)    return totalbaseline = {"status": 0.38, "eligibility": 0.30, "grievance": 0.12, "other": 0.20}this_week = {"status": 0.25, "eligibility": 0.22, "grievance": 0.10, "other": 0.43}print(round(psi(baseline, this_week), 3))   # 0.259

The topics come from your router or a topic classifier. For raw text, you can instead cluster query embeddings and compute PSI over the clusters, or track the average distance of new queries to the baseline centroid.

From "changed" to "worse"

Drift detection is a smoke alarm, not a diagnosis. The practical setup is a nightly job that (1) computes drift scores, (2) runs the golden eval set against the pinned production config, and (3) scores a sampled slice of yesterday's live traffic with an LLM judge. Alert when any of them moves beyond tolerance, then read traces from the drifted segment.

A real-life example

A state government chatbot's weekly drift report shows topic PSI of 0.26: the "other" share jumped from 20% to 43%. Nothing is broken in the system metrics.

The team clusters the "other" queries and finds a new cluster: questions about a housing scheme announced on Monday. The retrieval index has no documents for it, so the bot answers "I don't have information about this scheme" — correct, but useless. Retrieval top-score for these questions averages 0.41 against 0.78 normally, confirming corpus drift.

The fix is a process, not a model change: scheme documents are now added to the index on the day of announcement, and the drift job's "new cluster" alert goes to the content team as well as engineering. A new router route for the housing scheme follows a week later.

Follow-up questions to expect

  • "How do you detect a silent provider model update?" — Pin a dated snapshot; if you must use an alias, run the golden set daily and compare output length, refusal rate and scores.
  • "What baseline do you compare against?" — A rolling window (for example the last four weeks), plus a fixed launch baseline for slow drift.
  • "What do you do when drift is detected?" — Read traces from the drifted segment, add examples to the eval set, and fix the part that is stale — usually the corpus or routing, sometimes the prompt.