Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
You ship a prompt update Monday. By Friday, complaints triple — no errors, no alerts, just slowly worse answers. How do you monitor LLM output quality continuously in production?
What you need to know
Why quality needs its own monitoring
Classic monitoring watches errors, latency and uptime. A worse prompt produces none of those: every call returns 200 OK with a fluent answer. The only way to see it is to measure properties of the answers themselves and compare them over time and across versions.
The three tiers
| Tier | Coverage | Examples | Cost |
|---|---|---|---|
| Proxy metrics | 100% of traffic | Refusal rate, length, retrieval top score, answers without citations, JSON validity, rephrase-within-2-minutes, thumbs-down, escalations | Almost free |
| LLM judge | 1–5% sample | Groundedness, relevance, instruction-following on a rubric | A small share of model spend |
| Golden regression set | Every change, in CI | 200–500 real cases with expected behaviour | Minutes per run |
The "user rephrased immediately" signal is one of the best early warnings: people rarely ask the same thing twice when the first answer was good.
The judge must be checked before it is trusted. Have humans label a few hundred answers, measure agreement (for example Cohen's kappa of about 0.7 or higher), pin the judge model's version, and re-check each quarter.
Version everything
1def call_llm(messages, feature):2 span = tracer.start_span("llm", attributes={3 "prompt_version": PROMPTS[feature].version, # e.g. "support-v14"4 "model": MODEL_ID,5 "feature": feature,6 "canary": is_canary_user(),7 })8 reply = llm.generate(messages)9 span.set_attribute("refused", looks_like_refusal(reply))10 span.set_attribute("answer_tokens", count_tokens(reply))11 span.end()12 return replyWith these tags, every dashboard can be split by prompt version, and Monday's change becomes a visible group of requests instead of a mystery.
The release process
- Run the golden set in CI — block the change if it drops below the current version.
- Canary at 5% — for 24 hours, tagged by version.
- Compare arms — proxy metrics and judge scores, new against old.
- Roll forward or back — a flag flip, not a redeploy.
A real-life example
Scenario, numbers made up. A travel-booking assistant changes its system prompt to make answers "more concise". No errors appear. By Friday, complaints are three times higher.
After adding version tags and dashboards, the team replays the week. The new prompt's refusal rate rose from 2% to 6%, the rephrase rate from 8% to 15%, and answers about cancellation fees dropped the fee table the old prompt used to include. The judge's groundedness score fell 9 points on the sample. With a canary, all of this would have been visible on 5% of users within a day. They roll back, add 40 cancellation-fee cases to the golden set, and require a canary for every prompt change.
Follow-up questions to expect
- "Which single metric would you add first?" — Rephrase or retry rate split by prompt version: it is free, it moves within hours, and it reflects what users actually feel.
- "How do you trust an LLM judge?" — Calibrate it against human labels, pin its version, and use a different model family from the one being judged where possible to reduce self-preference.
- "How big should the canary be?" — Big enough to see your key metric move in a day. At low traffic, run longer rather than bigger.