Course Content
Agentic AI Patterns
9 sections · 50 lessons
How do you ensure prompt robustness and avoid prompt drift in production systems?
What you need to know
Where drift comes from
| Cause | Example |
|---|---|
| Model update | A floating alias like "latest" moves to a new snapshot |
| Tool or schema change | An API renames amount to total_amount |
| Data shift | New vendor invoice formats; a new product line |
| Untested edits | Five small "improvements" over a month |
| Context changes | A new tool added, changing how the model picks tools |
The controls
- Prompts are versioned code. In the repo, reviewed, with an ID stamped into every trace.
- Golden set per prompt. 50 to 300 real cases, weighted towards past failures, scored automatically, run in CI.
- Pin model versions. Use dated snapshots, not floating aliases. Re-run evals before switching, then canary on a small share of traffic.
- Structured output with validation. Format drift then fails loudly, with a repair or fallback path.
- Testable instructions. "Reply in under 120 words with one citation per claim" survives a model change. "Be concise and helpful" does not.
- Online monitoring. Schema-failure rate, refusal rate, tool-error rate, escalation rate and user feedback, with alerts on step changes.
Robustness testing
Before shipping, perturb the inputs: typos, different languages (Hindi-English mix is common in India), very long inputs, missing fields, and adversarial text. A prompt that only works on clean English examples will drift the moment real traffic arrives.
A real-life example
A procurement agent extracts quote details from vendor emails. The team moved from one model snapshot to a newer one after a price cut.
Their CI run on 220 golden cases showed overall accuracy up from 94% to 95%, but the delivery_date field dropped from 97% to 88%. The new model wrote dates like "12/03" without a year and sometimes read them as month-first. Averages hid it; the per-field report caught it.
Fixes: the schema now requires ISO dates (YYYY-MM-DD), a validator rejects anything else, and the prompt gives two examples with Indian day-first dates. After the change, delivery_date was at 98% and the upgrade shipped to 10% of traffic for a week before full rollout.
Follow-up questions to expect
- "How big should the golden set be?" — Big enough to cover each important case type several times; 100 to 300 is common. Add every production failure to it.
- "How do you compare two prompt versions fairly?" — Same cases, same model settings, scored per field or per category, with several runs if outputs vary.
- "What if the provider retires your pinned model?" — Plan migrations as projects: run the golden set on candidates early, fix regressions, and canary.