Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your team has 30+ prompts in production. Engineers tweak them ad-hoc, nobody knows which version is live, quality silently degrades. How do you version, test, and govern prompts at scale?


A prompt change, treated like codeEditrefund_replyon a branchCI runs its 64labelled casesPublish 2.3.0to the registry10% of trafficbehind a flagPromote or rollback by configEvery request logs prompt name and version.
A reviewer cannot predict quality from a prompt diff, so the score change has to sit next to the text change.

What you need to know

A prompt change is a behaviour change, just like a code change. But a person reading a prompt diff cannot predict its effect on quality. That is why prompts need the same machinery as code, plus an eval gate.

The four controls

ControlWhat it looks likeWhat it stops
Versioningrefund_reply@2.3.0 in a registry or repo; the app pins a version in config"Which prompt is live?"
Eval gateEach prompt has 30 to 100 labelled cases; CI fails on a regression beyond a thresholdSilent quality drops
Gradual rolloutNew version to 5 to 10% of traffic behind a flag, compared with controlA bad change reaching everyone
ProvenanceEvery request logs prompt_name, prompt_version and model versionBeing unable to explain a quality change

The change flow

  1. Edit — change the prompt file on a branch.
  2. Pull request — CI runs that prompt's eval set; the reviewer sees the score change, not only the text diff.
  3. Merge — the new version is published to the registry, not yet live.
  4. Roll out — a config flag sends 10% of traffic to the new version; online metrics are compared.
  5. Promote or roll back — a config change either way, no deploy.
YAML
# prompts/refund_reply.yamlname: refund_replyversion: 2.3.0owner: support-ai-teammodel: pinned-model-2026-05template: |  You are a support agent for {{brand}}. Use only the policy text below...evals: evals/refund_reply.jsonl     # 64 labelled casesmin_score: 0.88

The application calls registry.get("refund_reply", version=config.REFUND_REPLY_VERSION). Nobody edits a string in a console.

Keep the system honest

  • Grow the eval set. Every production failure becomes a new case. A suite that never grows stops catching real problems.
  • Reduce sprawl. Thirty prompts are often eight real prompts and 22 near-copies. Pull shared instructions (tone, safety rules, output format) into reusable partials so one fix applies everywhere.
  • Pin the model too. A prompt version only means something alongside a fixed model version.

A real-life example

Scenario (illustrative numbers). A food-delivery company has 34 prompts across support, restaurant onboarding and menu extraction. Complaints about "rude" support replies rise 3x in a month. Nobody can say what changed, because prompts are edited in an admin console.

The team moves all prompts into the repository with owners and versions, and adds request-level logging of prompt name and version. Within a day the logs show a support prompt was edited 11 times in three weeks; version 7 removed a line about empathy when refusing refunds. They add 50-case eval sets for the 10 highest-traffic prompts, merge 14 near-duplicates into 3 partials, and roll out changes at 10%. Over the next quarter, two regressions are caught in CI before merge.

Follow-up questions to expect

  • "Who can change a prompt?" — Anyone can open a pull request; the prompt's owner reviews, and the eval gate must pass. Product managers can edit too, through the same flow.
  • "What if there's no labelled data for a prompt?" — Start with 20 real inputs and write expected outputs by hand, or define rule checks (format, required facts). Something is better than nothing.
  • "How do you test prompts that produce free text?" — A calibrated LLM judge with a rubric, plus deterministic checks for format and must-include facts.