Course Content
LLM Evaluation
6 sections · 50 lessons
What is concept drift, and how does it affect model performance over time?
What you need to know
Concept drift
- The correct answer changes
- Example: leave policy updated from 24 to 26 days
- Fix: update the knowledge source and labels
- Hard to see in input statistics
Data drift
- The inputs change
- Example: a new office in Pune adds Marathi questions
- Fix: extend coverage, examples, eval slices
- Visible in input distribution
Where concept drift hits LLM systems
- Retrieval corpus — old documents stay indexed next to new ones; the retriever picks the old one.
- Parametric knowledge — the model's training data ends at a cutoff; facts after it are unknown.
- Your golden dataset — the labels were right when written. After a policy change, the eval fails correct new answers and passes wrong old ones.
Detecting drift
- Input distribution — compare the mix of topics or intents with a baseline. A common summary is the Population Stability Index (PSI):
PSI = sum over bins of (actual% - expected%) * ln(actual% / expected%)intents baseline: leave 40%, payroll 30%, benefits 20%, travel 10%intents this week: leave 30%, payroll 28%, benefits 22%, travel 20%PSI ≈ 0.10A rule of thumb treats PSI below 0.1 as little shift, 0.1 to 0.25 as moderate, and above 0.25 as large. Here travel doubled — worth a look.
- Output distribution — refusal rate, answer length, retrieval similarity scores, judge score trends.
- Behaviour — thumbs-down, rephrases and escalations often move first.
- Scheduled label review — re-check a sample of golden labels each quarter, with a named owner.
Responding
Fix the source of truth first — re-index documents, retire old versions, update labels. Changing prompts or models rarely fixes concept drift.
A real-life example
The company updated its travel policy on 1 July: hotel limits rose and a new per-diem table was added. Two weeks later, travel questions doubled in share (the PSI alert above), and thumbs-down on travel answers rose from 4% to 15%. Faithfulness scores stayed high.
The cause: both the old and new travel policy PDFs were in the index, and the old one had more text matching common questions. The assistant faithfully quoted outdated limits. The team removed superseded documents, added an "effective date" filter to retrieval, and updated 18 golden-set answers that still had the old limits — the eval had been passing the wrong answers. Thumbs-down on travel returned to 5% within a week.
Follow-up questions to expect
- "How do you prevent stale documents in RAG?" — Keep document metadata with effective and expiry dates, filter on it at retrieval time, and make document owners responsible for retiring old versions.
- "Can you detect concept drift without labels?" — Only indirectly: user signals, escalations and topic shifts. Confirmation needs fresh labels on a sample.
- "Does fine-tuning help with concept drift?" — Rarely the right fix for changing facts; retrieval from an updated source is cheaper and faster to correct.