Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
You migrate from GPT-4-turbo to a newer model and 15% of your existing prompts regress immediately. How do you migrate models safely in production?
What you need to know
GPT-4-turbo is a legacy model now, so this is a very real scenario: every team eventually has to move off a model that is being retired. The skill being tested is running a migration as an engineering process rather than a big-bang switch.
The process
- Freeze a regression suite — a few hundred real requests per prompt, stratified by intent, with reference outputs and a scorer: exact match or schema checks where possible, a calibrated LLM judge where not.
- Shadow — mirror live traffic to the new model and diff the outputs. This catches long inputs, messy user text and rare tool calls a static suite misses.
- Triage by cause — sort each failure into a bucket before fixing anything.
- Fix per prompt — each prompt is versioned and pinned to a model, so one prompt can move while another stays on the old model.
- Ramp — 5%, 25%, 50%, 100%, watching online metrics at each step, with rollback by config.
Typical failure buckets
| Bucket | Example | Fix |
|---|---|---|
| Format drift | New model wraps JSON in a code fence | Structured-output mode; adjust instructions |
| Verbosity change | Answers twice as long, breaking a UI limit | Explicit length guidance; max_tokens |
| Refusal differences | Declines a benign medical-billing question | Clarify context and allowed scope in the system prompt |
| Instruction style | Follows instructions more literally, drops an implied step | State the step explicitly |
| Genuinely worse reasoning | Fails multi-step calculations it used to pass | Keep that prompt on a stronger model, or add a tool |
Only the last bucket is a real model problem. In most migrations it is the smallest.
Migration as 30 small moves
1# routing config: prompt -> pinned model2refund_reply: { prompt: 3.1.0, model: new-model-2026-04 }3ticket_classifier: { prompt: 2.0.4, model: new-model-2026-04 }4contract_summary: { prompt: 5.2.0, model: old-model } # rewrite still in evalEach prompt moves on its own when its suite passes. The migration never blocks on the hardest prompt.
What to watch while ramping
Thumbs-down rate, escalations to humans, latency, cost per request, and parse-failure rate, compared between the old and new arms at each stage.
The organisational point
Never migrate under deprecation pressure. Start when the new model appears, keep both wired up, and the eventual retirement date becomes a non-event.
A real-life example
Scenario (illustrative numbers). An insurance-tech company has 28 prompts on an older model with a retirement date four months away. On a frozen suite of 6,400 real requests, the new model regresses on 15% of them.
Triage shows 41% format drift, 27% verbosity, 18% refusals on claims involving injuries, and 14% genuine reasoning failures, concentrated in two premium-calculation prompts. Format and verbosity are fixed in a week with structured outputs and length instructions; refusals are fixed by adding the insurer's context to the system prompt. The two calculation prompts get a calculator tool. After a two-week ramp, 26 prompts run on the new model; the last two follow a month later, well before the retirement date.
Follow-up questions to expect
- "How big should the regression suite be?" — Enough per prompt to cover its intents, often 100 to 300 examples, drawn from real traffic including edge cases.
- "How do you compare free-text outputs?" — A pairwise LLM judge (old versus new) with a rubric, calibrated against human preference on a sample.
- "What if the new model is better overall but worse on one prompt?" — Pin that prompt to the old model, or a different model, while you fix it.