Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Release Management for Prompts, Models and Indexes


In one week in the build phase, four things changed at Meridian. An engineer reworded the policy answer prompt. The platform team upgraded the reranker. The index was rebuilt with 22 policy changes. And the guard classifier was retrained on new examples. Each change went out on its own, through its own process, owned by a different person. By Friday, correctness on the golden set had fallen by four points.

Which change caused it? Nobody could say. Rolling back the prompt did not fix it. Rolling back the reranker needed the platform team, who were in a planning offsite. The index could not be rolled back at all, because nobody had kept the previous version. It took a week to untangle.

The lesson is simple: in an AI system, anything that changes an output is code, and it needs the release discipline of code. This lesson defines what a release is, how it is bundled, gated and rolled out, and how daily policy changes fit in. It produces part B of MER-09.

Two release lanes, one disciplineConfiguration lane• Prompts, models,retrieval, guards, checks• Layers 1 to 3, then shadow and canary• Human approval; model risk if material• Days to weeks, one change per bundleContent lane• Policy added, changed or withdrawn• Checks on linked golden cases• Automatic when checks pass• Minutes to hours; withdrawals at once
Anything that changes an output is code, but policy content needs its own fast lane or the freshness targets cannot be met.

What counts as a release artifact

If it can change what a user sees, it is versioned, reviewed, tested and released. At Meridian that list is long.

  • Prompt templates and their examples
  • Model version and generation settings
  • Retrieval settings: chunking rules, candidate count, final count, filters
  • Embedding model and reranker
  • Index content
  • Guard models and their thresholds
  • Tool schemas and descriptions
  • Deterministic check rules
  • The judge prompt used in evaluation

The last item surprises people. If the judge prompt changes, scores change, and a release can pass or fail because the ruler changed rather than the thing measured. So the judge is versioned too, and a judge change is evaluated against the human-graded calibration set before it is used.

The release bundle

All of these are pinned together in a release bundle: a manifest with one ID. Production runs exactly one bundle per capability. Rollback means redeploying the previous bundle, which restores every setting at once.

YAML
bundle: policy-answer-2026.06.3capability: policy_answerprompt: policy-answer-v15model: {route_primary: gen-a-2026-04-15, route_fallback: gen-b-2026-06-01}generation: {max_tokens: 600, temperature: 0.2}retrieval:  chunker: headings-v3-tables-whole  embed_model: embed-a-v2  reranker: rerank-a-v4  candidates: 40  final: 6  filters: [role, effective_date, product]index_lane: content      # index content releases separately, see belowguard: {model: guard-8b-v3.1, block_threshold: 0.85}checks: checks-v9evaluated_with: {golden: policy-golden-v22, judge: judge-v5}classification: non-materialapproved_by: service owner, 2026-06-18

Meridian ships one change per bundle wherever possible. When a bundle contains a single change, a regression points straight at its cause. When changes must go together, such as a new reranker that needs a new candidate count, they go together in one bundle, evaluated together.

Two lanes: configuration and content

Index content is the exception. Policies change about 20 times a month, and freshness targets require new content within four hours and withdrawals within one. Content cannot wait for a full release cycle. So it has its own automated lane.

Configuration lane

  • Prompts, models, retrieval settings, guards, checks
  • Full layer 1 to 3 evaluation, then shadow and canary
  • Human approval; model risk for material changes
  • Days to weeks

Content lane

  • Policy additions, changes and withdrawals
  • Automated checks on affected golden cases and retrieval
  • Automatic when checks pass; alert when they fail
  • Minutes to hours

The content lane runs for each published change: convert, chunk, embed, build a new index version, run retrieval checks on the golden cases linked to the affected sections plus a fixed smoke set, then switch the live index alias to the new version. The previous seven index versions are kept, so rolling back content is one alias change. Withdrawals skip the gate and apply immediately, because removal is always the safe direction, but they still raise an alert if they fail, which is exactly the alert that went unwatched in the car finance incident.

Gates and rollout

A configuration bundle moves through five steps.

  1. Pull request — layers 1 to 3 of the evaluation pyramid on the changed capability, including the held-out set.
  2. Classify — material or non-material, using the change classification agreed with model risk in section 2.
  3. Shadow — the new bundle runs on a copy of live requests for two days, not shown to anyone; outputs are compared with the current bundle by checks and the judge.
  4. Canary — one team per centre, about 10% of staff, uses the new bundle for three days while its objectives are watched.
  5. Full rollout — the previous bundle stays deployable for 30 days.

Shadow runs have two costs worth stating. They use real tokens, about $150 for two days of policy answers, and they process real customer data, so they run under exactly the same controls as production. In shadow mode, every write tool is stubbed: a shadow letter draft is compared and discarded, never created in DocGen.

Rollback is automatic for an invariant violation, and manual but pre-approved for two other triggers: a deterministic check failure rate twice the previous bundle's, or sampled correctness "at risk" against the previous bundle during canary.

Part B of MER-09 is the release policy.

ChangeLaneGatesRolloutApproverRollback
Prompt wordingConfigurationLayers 1 to 3, shadowCanary then fullService ownerPrevious bundle
Retrieval setting or rerankerConfigurationLayers 1 to 3, component, shadowCanary then fullService ownerPrevious bundle
Model versionConfiguration, materialFull suite, shadow, model risk reviewCanary then fullModel risk and service ownerPrevious bundle
Guard thresholdConfigurationAdversarial set, false block rateCanary then fullSecurity and service ownerPrevious bundle
Policy added or changedContentLinked golden cases, smoke setAutomaticPipeline, alert on failurePrevious index version
Policy withdrawnContentNone, applied at onceAutomaticPipeline, alert on failureNot applicable
New capabilityNew designFull design-record reviewPilotReview boardKill switch

Check your understanding

0 of 3 answered

1.Why is the evaluation judge's prompt versioned as part of the release process?

2.Why does policy content use a separate lane from prompts and model changes?

3.In shadow mode, the new bundle generates a letter draft. What should happen to it?