Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

You ship a prompt update Monday. By Friday, complaints triple — no errors, no alerts, just slowly worse answers. How do you monitor LLM output quality continuously in production?


Three tiers of quality monitoringProxies on100% of trafficLLM judge ona 1-5% sampleGolden setgates each changetopbottomTag every request with prompt_version so each tier splits by release.
Refusals rose from 2 to 6 percent on day one, but nobody saw it because nothing split the metrics by prompt version.

What you need to know

Why quality needs its own monitoring

Classic monitoring watches errors, latency and uptime. A worse prompt produces none of those: every call returns 200 OK with a fluent answer. The only way to see it is to measure properties of the answers themselves and compare them over time and across versions.

The three tiers

TierCoverageExamplesCost
Proxy metrics100% of trafficRefusal rate, length, retrieval top score, answers without citations, JSON validity, rephrase-within-2-minutes, thumbs-down, escalationsAlmost free
LLM judge1–5% sampleGroundedness, relevance, instruction-following on a rubricA small share of model spend
Golden regression setEvery change, in CI200–500 real cases with expected behaviourMinutes per run

The "user rephrased immediately" signal is one of the best early warnings: people rarely ask the same thing twice when the first answer was good.

The judge must be checked before it is trusted. Have humans label a few hundred answers, measure agreement (for example Cohen's kappa of about 0.7 or higher), pin the judge model's version, and re-check each quarter.

Version everything

Python
def call_llm(messages, feature):    span = tracer.start_span("llm", attributes={        "prompt_version": PROMPTS[feature].version,   # e.g. "support-v14"        "model": MODEL_ID,        "feature": feature,        "canary": is_canary_user(),    })    reply = llm.generate(messages)    span.set_attribute("refused", looks_like_refusal(reply))    span.set_attribute("answer_tokens", count_tokens(reply))    span.end()    return reply

With these tags, every dashboard can be split by prompt version, and Monday's change becomes a visible group of requests instead of a mystery.

The release process

  1. Run the golden set in CI — block the change if it drops below the current version.
  2. Canary at 5% — for 24 hours, tagged by version.
  3. Compare arms — proxy metrics and judge scores, new against old.
  4. Roll forward or back — a flag flip, not a redeploy.

A real-life example

Scenario, numbers made up. A travel-booking assistant changes its system prompt to make answers "more concise". No errors appear. By Friday, complaints are three times higher.

After adding version tags and dashboards, the team replays the week. The new prompt's refusal rate rose from 2% to 6%, the rephrase rate from 8% to 15%, and answers about cancellation fees dropped the fee table the old prompt used to include. The judge's groundedness score fell 9 points on the sample. With a canary, all of this would have been visible on 5% of users within a day. They roll back, add 40 cancellation-fee cases to the golden set, and require a canary for every prompt change.

Follow-up questions to expect

  • "Which single metric would you add first?" — Rephrase or retry rate split by prompt version: it is free, it moves within hours, and it reflects what users actually feel.
  • "How do you trust an LLM judge?" — Calibrate it against human labels, pin its version, and use a different model family from the one being judged where possible to reduce self-preference.
  • "How big should the canary be?" — Big enough to see your key metric move in a day. At low traffic, run longer rather than bigger.