Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

You migrate from GPT-4-turbo to a newer model and 15% of your existing prompts regress immediately. How do you migrate models safely in production?


Why 15% of 6,400 requests regressed41%format27%verbosity18%refusals14%reasoning0123prompt fixprompt fixprompt fixrealmodel gap
Only the smallest bucket is a model problem, which is why triage comes before any decision to roll back.

What you need to know

GPT-4-turbo is a legacy model now, so this is a very real scenario: every team eventually has to move off a model that is being retired. The skill being tested is running a migration as an engineering process rather than a big-bang switch.

The process

  1. Freeze a regression suite — a few hundred real requests per prompt, stratified by intent, with reference outputs and a scorer: exact match or schema checks where possible, a calibrated LLM judge where not.
  2. Shadow — mirror live traffic to the new model and diff the outputs. This catches long inputs, messy user text and rare tool calls a static suite misses.
  3. Triage by cause — sort each failure into a bucket before fixing anything.
  4. Fix per prompt — each prompt is versioned and pinned to a model, so one prompt can move while another stays on the old model.
  5. Ramp — 5%, 25%, 50%, 100%, watching online metrics at each step, with rollback by config.

Typical failure buckets

BucketExampleFix
Format driftNew model wraps JSON in a code fenceStructured-output mode; adjust instructions
Verbosity changeAnswers twice as long, breaking a UI limitExplicit length guidance; max_tokens
Refusal differencesDeclines a benign medical-billing questionClarify context and allowed scope in the system prompt
Instruction styleFollows instructions more literally, drops an implied stepState the step explicitly
Genuinely worse reasoningFails multi-step calculations it used to passKeep that prompt on a stronger model, or add a tool

Only the last bucket is a real model problem. In most migrations it is the smallest.

Migration as 30 small moves

YAML
# routing config: prompt -> pinned modelrefund_reply:      { prompt: 3.1.0, model: new-model-2026-04 }ticket_classifier: { prompt: 2.0.4, model: new-model-2026-04 }contract_summary:  { prompt: 5.2.0, model: old-model }        # rewrite still in eval

Each prompt moves on its own when its suite passes. The migration never blocks on the hardest prompt.

What to watch while ramping

Thumbs-down rate, escalations to humans, latency, cost per request, and parse-failure rate, compared between the old and new arms at each stage.

The organisational point

Never migrate under deprecation pressure. Start when the new model appears, keep both wired up, and the eventual retirement date becomes a non-event.

A real-life example

Scenario (illustrative numbers). An insurance-tech company has 28 prompts on an older model with a retirement date four months away. On a frozen suite of 6,400 real requests, the new model regresses on 15% of them.

Triage shows 41% format drift, 27% verbosity, 18% refusals on claims involving injuries, and 14% genuine reasoning failures, concentrated in two premium-calculation prompts. Format and verbosity are fixed in a week with structured outputs and length instructions; refusals are fixed by adding the insurer's context to the system prompt. The two calculation prompts get a calculator tool. After a two-week ramp, 26 prompts run on the new model; the last two follow a month later, well before the retirement date.

Follow-up questions to expect

  • "How big should the regression suite be?" — Enough per prompt to cover its intents, often 100 to 300 examples, drawn from real traffic including edge cases.
  • "How do you compare free-text outputs?" — A pairwise LLM judge (old versus new) with a rubric, calibrated against human preference on a sample.
  • "What if the new model is better overall but worse on one prompt?" — Pin that prompt to the old model, or a different model, while you fix it.