Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your team has 30+ prompts in production. Engineers tweak them ad-hoc, nobody knows which version is live, quality silently degrades. How do you version, test, and govern prompts at scale?
What you need to know
A prompt change is a behaviour change, just like a code change. But a person reading a prompt diff cannot predict its effect on quality. That is why prompts need the same machinery as code, plus an eval gate.
The four controls
| Control | What it looks like | What it stops |
|---|---|---|
| Versioning | refund_reply@2.3.0 in a registry or repo; the app pins a version in config | "Which prompt is live?" |
| Eval gate | Each prompt has 30 to 100 labelled cases; CI fails on a regression beyond a threshold | Silent quality drops |
| Gradual rollout | New version to 5 to 10% of traffic behind a flag, compared with control | A bad change reaching everyone |
| Provenance | Every request logs prompt_name, prompt_version and model version | Being unable to explain a quality change |
The change flow
- Edit — change the prompt file on a branch.
- Pull request — CI runs that prompt's eval set; the reviewer sees the score change, not only the text diff.
- Merge — the new version is published to the registry, not yet live.
- Roll out — a config flag sends 10% of traffic to the new version; online metrics are compared.
- Promote or roll back — a config change either way, no deploy.
1# prompts/refund_reply.yaml2name: refund_reply3version: 2.3.04owner: support-ai-team5model: pinned-model-2026-056template: |7 You are a support agent for {{brand}}. Use only the policy text below...8evals: evals/refund_reply.jsonl # 64 labelled cases9min_score: 0.88The application calls registry.get("refund_reply", version=config.REFUND_REPLY_VERSION). Nobody edits a string in a console.
Keep the system honest
- Grow the eval set. Every production failure becomes a new case. A suite that never grows stops catching real problems.
- Reduce sprawl. Thirty prompts are often eight real prompts and 22 near-copies. Pull shared instructions (tone, safety rules, output format) into reusable partials so one fix applies everywhere.
- Pin the model too. A prompt version only means something alongside a fixed model version.
A real-life example
Scenario (illustrative numbers). A food-delivery company has 34 prompts across support, restaurant onboarding and menu extraction. Complaints about "rude" support replies rise 3x in a month. Nobody can say what changed, because prompts are edited in an admin console.
The team moves all prompts into the repository with owners and versions, and adds request-level logging of prompt name and version. Within a day the logs show a support prompt was edited 11 times in three weeks; version 7 removed a line about empathy when refusing refunds. They add 50-case eval sets for the 10 highest-traffic prompts, merge 14 near-duplicates into 3 partials, and roll out changes at 10%. Over the next quarter, two regressions are caught in CI before merge.
Follow-up questions to expect
- "Who can change a prompt?" — Anyone can open a pull request; the prompt's owner reviews, and the eval gate must pass. Product managers can edit too, through the same flow.
- "What if there's no labelled data for a prompt?" — Start with 20 real inputs and write expected outputs by hand, or define rule checks (format, required facts). Something is better than nothing.
- "How do you test prompts that produce free text?" — A calibrated LLM judge with a rubric, plus deterministic checks for format and must-include facts.