Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

How do you ensure prompt robustness and avoid prompt drift in production systems?


What you need to know

Where drift comes from

CauseExample
Model updateA floating alias like "latest" moves to a new snapshot
Tool or schema changeAn API renames amount to total_amount
Data shiftNew vendor invoice formats; a new product line
Untested editsFive small "improvements" over a month
Context changesA new tool added, changing how the model picks tools

The controls

  • Prompts are versioned code. In the repo, reviewed, with an ID stamped into every trace.
  • Golden set per prompt. 50 to 300 real cases, weighted towards past failures, scored automatically, run in CI.
  • Pin model versions. Use dated snapshots, not floating aliases. Re-run evals before switching, then canary on a small share of traffic.
  • Structured output with validation. Format drift then fails loudly, with a repair or fallback path.
  • Testable instructions. "Reply in under 120 words with one citation per claim" survives a model change. "Be concise and helpful" does not.
  • Online monitoring. Schema-failure rate, refusal rate, tool-error rate, escalation rate and user feedback, with alerts on step changes.

Robustness testing

Before shipping, perturb the inputs: typos, different languages (Hindi-English mix is common in India), very long inputs, missing fields, and adversarial text. A prompt that only works on clean English examples will drift the moment real traffic arrives.

A real-life example

A procurement agent extracts quote details from vendor emails. The team moved from one model snapshot to a newer one after a price cut.

Their CI run on 220 golden cases showed overall accuracy up from 94% to 95%, but the delivery_date field dropped from 97% to 88%. The new model wrote dates like "12/03" without a year and sometimes read them as month-first. Averages hid it; the per-field report caught it.

Fixes: the schema now requires ISO dates (YYYY-MM-DD), a validator rejects anything else, and the prompt gives two examples with Indian day-first dates. After the change, delivery_date was at 98% and the upgrade shipped to 10% of traffic for a week before full rollout.

Follow-up questions to expect

  • "How big should the golden set be?" — Big enough to cover each important case type several times; 100 to 300 is common. Add every production failure to it.
  • "How do you compare two prompt versions fairly?" — Same cases, same model settings, scored per field or per category, with several runs if outputs vary.
  • "What if the provider retires your pinned model?" — Plan migrations as projects: run the golden set on candidates early, fix regressions, and canary.