Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Models Change Under You: Pinning, Drift and Retirement
Meridian's credit card team launched a complaint-summary tool in 2024. It was validated by model risk on one specific model version. Fourteen months later, the provider announced that version would be retired. The team needed to choose a replacement, re-run every evaluation, fix three regressions in how the new model handled refund amounts, and get model risk to re-validate. That took eleven weeks. The notice they received was shorter than eleven weeks. For a fortnight, the bank ran a tool on a model that was about to disappear, while arguing about whether an unvalidated replacement was acceptable.
Traditional software ages slowly. A database version is supported for years, and you choose when to upgrade. AI dependencies age fast, and often someone else chooses. Models are retired, aliases move to new versions, policies change underneath your index, and users change how they use the tool.
This lesson covers version pinning, the four kinds of drift, and how to treat deprecation as a scheduled project. It ends with the third part of MER-02: the model life cycle sheet.
Aliases and pinned versions
Most hosted providers offer two kinds of model identifier. An alias is a name that points to whatever the provider currently considers the best model in a family; the model behind it changes over time. A pinned version, often with a date in its name, refers to one specific snapshot that the provider promises not to change.
For a regulated system, the rule is simple: pin the version. What model risk validated is a system built on one specific model. If an alias moves, you are running something nobody validated, and you may not even know it happened.
Pinning does not make you safe forever. Pinned versions are retired, and notice periods vary by provider, from a few months to a year or more. Read your provider's deprecation policy, and try to get a minimum notice period into the contract. Open-weights models that you host yourself never get retired from under you, but you still have to patch the serving software, and you will eventually want to move to better models, so the same migration discipline applies.
Four kinds of drift
Drift is any change in a system's behaviour or context that happens without a code change. It comes in four kinds, and each needs its own detector.
| Drift | What changes | Meridian example | How you notice |
|---|---|---|---|
| Model drift | The model's behaviour | Provider changes serving, or an alias moves | Daily canary cases, answer length and refusal rate |
| Input drift | What users send | A rate rise brings a surge of new arrears questions | Topic mix of questions, retrieval scores |
| Knowledge drift | The source of truth | 20 policy changes a month make the index stale | Freshness lag from publish to index |
| Usage drift | How people use the tool | Staff start pasting whole complaint emails | Input length, new question types, feedback |
Knowledge drift is the one most specific to retrieval systems, and at Meridian it is the most dangerous. An assistant that answers confidently from last month's arrears policy is worse than no assistant, because it adds authority to an outdated answer. Section 4 sets freshness targets for exactly this reason.
Usage drift is the easiest to miss, because nothing technical changes. The system was evaluated on short policy questions; if staff start using it to interpret customer complaints, it is now doing a job nobody tested.
Deprecation is a scheduled project
A model migration is not a configuration change. It is a small project with predictable steps, and it takes weeks.
- Select — shortlist one or two replacement candidates, using the evidence ladder.
- Evaluate — run the full evaluation suite for every capability on each candidate.
- Fix — adjust prompts, and sometimes retrieval settings, for regressions the evaluation finds.
- Re-validate — model risk reviews the evidence at the depth the change classification requires.
- Canary — send a small share of real traffic to the new version and compare.
- Switch and retire — move all traffic, keep the old version configured until its retirement date.
Meridian's estimate for this sequence is eight to ten weeks. Two design choices shorten it. First, keep a warm candidate: re-run the evaluation suite on the most likely replacement every quarter, so that a retirement notice starts a migration from evidence rather than from nothing. Second, agree a change classification with model risk before launch. For example: a new model version is a material change needing re-validation of the affected capabilities; a prompt wording change that passes all evaluation gates is non-material and needs only notification. Agreeing this in advance turns every future change from a negotiation into a procedure.
The model life cycle sheet
Every model the system depends on gets an entry. This is Meridian's sheet at design time; the model identifiers are placeholders for the pinned versions the team will choose after the spike.
1id: MER-02-C2models:3 - role: answer_and_draft # policy answers, summaries, letter text4 source: hosted, provider A, in-region endpoint5 pinned_version: gen-a-2026-04-156 validated: pending # set by model risk at approval7 retirement_announced: none8 contract_min_notice_days: 1809 warm_candidate: gen-b-2026-06-0110 candidate_last_evaluated: 2026-07-0111 canary_cases_daily: 5012 owner: platform team13 - role: guard_and_rewrite # input checks, query rewrite14 source: self-hosted open weights, 8B parameters15 pinned_version: guard-8b-v3.116 retirement_announced: not applicable17 patch_owner: platform team18 - role: embeddings19 source: hosted, provider A20 pinned_version: embed-a-v221 reindex_cost: about 9,000 chunks, under 10 minutes22 on_change: re-run retrieval evaluation before switch23change_classification:24 model_version_change: material # re-validate affected capabilities25 prompt_change_passing_gates: non-material26 index_refresh_same_pipeline: non-material27 new_capability: materialThe sheet answers the question every model risk team eventually asks: "What happens when this model goes away?" If you can show the warm candidate's latest scores and the agreed change classification, the answer is a procedure, not a promise.
Check your understanding
0 of 3 answered
1.Why should a regulated system call a pinned model version rather than an alias?
2.Staff at Meridian begin pasting entire customer complaint emails into the policy assistant. What kind of drift is this, and why does it matter?
3.A retirement notice arrives for Meridian's generation model. What most shortens the migration?