Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
You distill GPT-4 into a 7B model to save costs. It works 3 months, then accuracy silently degrades as behavior shifts. How do you monitor and refresh distilled models in production?
What you need to know
Why the decay is silent
The student is fine on the kinds of inputs it was trained on. Three months later, users ask about new products, a prompt template has changed, and the retriever returns different documents. The student meets inputs it never saw and guesses. Nothing throws an error, because a wrong answer looks like a right one. (Today's teacher would be a current frontier model; GPT-4 itself is retired. Check the teacher provider's terms on using outputs for training.)
Four monitors
| Monitor | What it catches | Speed |
|---|---|---|
| Student-teacher agreement on a 1–2% sample | Quality drop on real traffic | Days |
| Input drift: share of queries far from training clusters | New kinds of input, before accuracy falls | Hours |
| Frozen golden set of ~500 labelled cases, nightly | Regressions from your own changes | Daily |
| Implicit signals: thumbs-down, retries, escalations | User pain on 100% of traffic | Noisy, but broad |
A sketch of agreement sampling that never slows the user:
1import random23async def answer(request):4 out = await student.generate(request)5 if random.random() < 0.02: # 2% shadow sample6 queue.put_nowait({"input": request, "student": out})7 return out89async def shadow_worker():10 while True:11 item = await queue.get()12 ref = await teacher.generate(item["input"])13 agree = judge_agreement(item["student"], ref) # exact match for labels, judge for text14 metrics.record("student_teacher_agreement", agree)15 if not agree:16 store.save_disagreement(item["input"], ref) # future training dataThe teacher call happens in the background, so latency is unchanged, and the cost is roughly 2% of what running the teacher on everything would cost.
The refresh loop
- Collect — log disagreements and out-of-distribution inputs, with the teacher's answer as the label.
- Review — spot-check a sample by hand, because the teacher is also sometimes wrong.
- Retrain — add them to the training mix (keep old data too, so old skills are not lost), monthly or when agreement drops about 3 points.
- Canary — serve the new student on 5% of traffic, compare on the golden set and agreement, then promote.
- Keep rollback ready — the previous version stays one config flag away.
A real-life example
Scenario, numbers made up. A fintech distils a frontier model into a 7B open-weights model to route support tickets into 30 queues. At launch, the student agrees with the teacher 94% of the time. Three months later the company launches credit on UPI, and tickets about it get routed to the wrong teams. Nobody notices for weeks.
After adding monitoring, a drop like that shows up within a week: agreement falls to 81%, and the drift monitor shows 9% of tickets far from any training cluster. The team collects 6,000 disagreement cases, has two analysts review 500 of them, retrains, and canaries the new student. Agreement returns to 93%.
Follow-up questions to expect
- "What if the teacher is also wrong on new inputs?" — That is why a human-labelled golden set exists and why a reviewed sample of teacher labels goes into training; the teacher is a strong labeller, not ground truth.
- "How do you measure agreement for free text?" — Use an LLM judge with a rubric or a claim-level comparison; embedding similarity alone misses answers that sound alike but differ on a key fact.
- "Why not just use the teacher everywhere?" — Cost and latency. At 2% sampling, you keep most of the saving while buying early warning.