LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What is concept drift, and how does it affect model performance over time?


Share of HR questions by intent40%30%20%10%30%28%22%20%leavepayrollbenefitstravelbaselineafter 1 JulyPSI about 0.10; travel thumbs-down rose from 4% to 15%.
The input shift pointed at travel, where old and new policy PDFs were both still in the index.

What you need to know

Concept drift

  • The correct answer changes
  • Example: leave policy updated from 24 to 26 days
  • Fix: update the knowledge source and labels
  • Hard to see in input statistics

Data drift

  • The inputs change
  • Example: a new office in Pune adds Marathi questions
  • Fix: extend coverage, examples, eval slices
  • Visible in input distribution

Where concept drift hits LLM systems

  1. Retrieval corpus — old documents stay indexed next to new ones; the retriever picks the old one.
  2. Parametric knowledge — the model's training data ends at a cutoff; facts after it are unknown.
  3. Your golden dataset — the labels were right when written. After a policy change, the eval fails correct new answers and passes wrong old ones.

Detecting drift

  • Input distribution — compare the mix of topics or intents with a baseline. A common summary is the Population Stability Index (PSI):
Text
PSI = sum over bins of (actual% - expected%) * ln(actual% / expected%)intents baseline: leave 40%, payroll 30%, benefits 20%, travel 10%intents this week: leave 30%, payroll 28%, benefits 22%, travel 20%PSI ≈ 0.10

A rule of thumb treats PSI below 0.1 as little shift, 0.1 to 0.25 as moderate, and above 0.25 as large. Here travel doubled — worth a look.

  • Output distribution — refusal rate, answer length, retrieval similarity scores, judge score trends.
  • Behaviour — thumbs-down, rephrases and escalations often move first.
  • Scheduled label review — re-check a sample of golden labels each quarter, with a named owner.

Responding

Fix the source of truth first — re-index documents, retire old versions, update labels. Changing prompts or models rarely fixes concept drift.

A real-life example

The company updated its travel policy on 1 July: hotel limits rose and a new per-diem table was added. Two weeks later, travel questions doubled in share (the PSI alert above), and thumbs-down on travel answers rose from 4% to 15%. Faithfulness scores stayed high.

The cause: both the old and new travel policy PDFs were in the index, and the old one had more text matching common questions. The assistant faithfully quoted outdated limits. The team removed superseded documents, added an "effective date" filter to retrieval, and updated 18 golden-set answers that still had the old limits — the eval had been passing the wrong answers. Thumbs-down on travel returned to 5% within a week.

Follow-up questions to expect

  • "How do you prevent stale documents in RAG?" — Keep document metadata with effective and expiry dates, filter on it at retrieval time, and make document owners responsible for retiring old versions.
  • "Can you detect concept drift without labels?" — Only indirectly: user signals, escalations and topic shifts. Confirmation needs fresh labels on a sample.
  • "Does fine-tuning help with concept drift?" — Rarely the right fix for changing facts; retrieval from an updated source is cheaper and faster to correct.