Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Release Management for Prompts, Models and Indexes
In one week in the build phase, four things changed at Meridian. An engineer reworded the policy answer prompt. The platform team upgraded the reranker. The index was rebuilt with 22 policy changes. And the guard classifier was retrained on new examples. Each change went out on its own, through its own process, owned by a different person. By Friday, correctness on the golden set had fallen by four points.
Which change caused it? Nobody could say. Rolling back the prompt did not fix it. Rolling back the reranker needed the platform team, who were in a planning offsite. The index could not be rolled back at all, because nobody had kept the previous version. It took a week to untangle.
The lesson is simple: in an AI system, anything that changes an output is code, and it needs the release discipline of code. This lesson defines what a release is, how it is bundled, gated and rolled out, and how daily policy changes fit in. It produces part B of MER-09.
What counts as a release artifact
If it can change what a user sees, it is versioned, reviewed, tested and released. At Meridian that list is long.
- Prompt templates and their examples
- Model version and generation settings
- Retrieval settings: chunking rules, candidate count, final count, filters
- Embedding model and reranker
- Index content
- Guard models and their thresholds
- Tool schemas and descriptions
- Deterministic check rules
- The judge prompt used in evaluation
The last item surprises people. If the judge prompt changes, scores change, and a release can pass or fail because the ruler changed rather than the thing measured. So the judge is versioned too, and a judge change is evaluated against the human-graded calibration set before it is used.
The release bundle
All of these are pinned together in a release bundle: a manifest with one ID. Production runs exactly one bundle per capability. Rollback means redeploying the previous bundle, which restores every setting at once.
1bundle: policy-answer-2026.06.32capability: policy_answer3prompt: policy-answer-v154model: {route_primary: gen-a-2026-04-15, route_fallback: gen-b-2026-06-01}5generation: {max_tokens: 600, temperature: 0.2}6retrieval:7 chunker: headings-v3-tables-whole8 embed_model: embed-a-v29 reranker: rerank-a-v410 candidates: 4011 final: 612 filters: [role, effective_date, product]13index_lane: content # index content releases separately, see below14guard: {model: guard-8b-v3.1, block_threshold: 0.85}15checks: checks-v916evaluated_with: {golden: policy-golden-v22, judge: judge-v5}17classification: non-material18approved_by: service owner, 2026-06-18Meridian ships one change per bundle wherever possible. When a bundle contains a single change, a regression points straight at its cause. When changes must go together, such as a new reranker that needs a new candidate count, they go together in one bundle, evaluated together.
Two lanes: configuration and content
Index content is the exception. Policies change about 20 times a month, and freshness targets require new content within four hours and withdrawals within one. Content cannot wait for a full release cycle. So it has its own automated lane.
Configuration lane
- Prompts, models, retrieval settings, guards, checks
- Full layer 1 to 3 evaluation, then shadow and canary
- Human approval; model risk for material changes
- Days to weeks
Content lane
- Policy additions, changes and withdrawals
- Automated checks on affected golden cases and retrieval
- Automatic when checks pass; alert when they fail
- Minutes to hours
The content lane runs for each published change: convert, chunk, embed, build a new index version, run retrieval checks on the golden cases linked to the affected sections plus a fixed smoke set, then switch the live index alias to the new version. The previous seven index versions are kept, so rolling back content is one alias change. Withdrawals skip the gate and apply immediately, because removal is always the safe direction, but they still raise an alert if they fail, which is exactly the alert that went unwatched in the car finance incident.
Gates and rollout
A configuration bundle moves through five steps.
- Pull request — layers 1 to 3 of the evaluation pyramid on the changed capability, including the held-out set.
- Classify — material or non-material, using the change classification agreed with model risk in section 2.
- Shadow — the new bundle runs on a copy of live requests for two days, not shown to anyone; outputs are compared with the current bundle by checks and the judge.
- Canary — one team per centre, about 10% of staff, uses the new bundle for three days while its objectives are watched.
- Full rollout — the previous bundle stays deployable for 30 days.
Shadow runs have two costs worth stating. They use real tokens, about $150 for two days of policy answers, and they process real customer data, so they run under exactly the same controls as production. In shadow mode, every write tool is stubbed: a shadow letter draft is compared and discarded, never created in DocGen.
Rollback is automatic for an invariant violation, and manual but pre-approved for two other triggers: a deterministic check failure rate twice the previous bundle's, or sampled correctness "at risk" against the previous bundle during canary.
Part B of MER-09 is the release policy.
| Change | Lane | Gates | Rollout | Approver | Rollback |
|---|---|---|---|---|---|
| Prompt wording | Configuration | Layers 1 to 3, shadow | Canary then full | Service owner | Previous bundle |
| Retrieval setting or reranker | Configuration | Layers 1 to 3, component, shadow | Canary then full | Service owner | Previous bundle |
| Model version | Configuration, material | Full suite, shadow, model risk review | Canary then full | Model risk and service owner | Previous bundle |
| Guard threshold | Configuration | Adversarial set, false block rate | Canary then full | Security and service owner | Previous bundle |
| Policy added or changed | Content | Linked golden cases, smoke set | Automatic | Pipeline, alert on failure | Previous index version |
| Policy withdrawn | Content | None, applied at once | Automatic | Pipeline, alert on failure | Not applicable |
| New capability | New design | Full design-record review | Pilot | Review board | Kill switch |
Check your understanding
0 of 3 answered
1.Why is the evaluation judge's prompt versioned as part of the release process?
2.Why does policy content use a separate lane from prompts and model changes?
3.In shadow mode, the new bundle generates a letter draft. What should happen to it?