Course Content
MLOps for AI
4 sections · 9 lessons
Model Registry and Rollback
A serving pod restarts at 09:14 and loads its model with the line it has always used: mlflow.pyfunc.load_model("models:/fraud-scorer/Production"). It comes back predicting differently from its three sibling pods. Same image, same code, same URI, different model.
The cause is that two versions — 11 and 14 — are both sitting in the Production stage. Someone promoted version 14 three weeks ago without archiving version 11, and the URI resolves to "the latest version in that stage", which is a definition that quietly changed meaning the moment a second version appeared. For three weeks, whichever pod restarted most recently determined what half the traffic saw.
Nobody noticed because nothing errored. The fraud rate moved by 0.4 percentage points and was attributed to seasonality.
This is the failure a model registry exists to prevent, and the failure a badly used registry causes. The difference between the two is almost entirely about whether your promotion mechanism is a set that can accidentally hold two members, or a pointer that can only point at one thing.
What a registry is, and what it is not
An experiment tracker records what you tried. A registry records what you decided. They are different databases with different lifecycles: you will produce two thousand runs and register nine models.
The structure has three levels:
| Level | What it is | Example | Mutable? |
|---|---|---|---|
| Registered model | A named slot for one prediction task; owns description, tags, permissions | fraud-scorer | Metadata yes; identity no |
| Model version | An immutable numbered snapshot pointing at one run's artefact | fraud-scorer version 14 | Bytes never; tags and description yes |
| Alias | A movable, single-valued name pointing at exactly one version | @champion → version 14 | That is its entire job |
The immutability of versions is the property everything else rests on. Version 14 always means the same bytes. If you retrain, you get version 15 — you never overwrite 14. That is what makes "roll back to 14" a meaningful instruction rather than a hope.
A registry answers one question that no filesystem, bucket or spreadsheet can answer reliably: which exact artefact is authorised to serve traffic right now, and who authorised it.
Registering a model, with the metadata that makes it useful
1import mlflow2from mlflow import MlflowClient34client = MlflowClient()56# 1. Create the slot once. Idempotent guard so pipelines can rerun safely.7try:8 client.create_registered_model(9 name="fraud-scorer",10 description="Scores card transactions for fraud risk. Output: P(fraud) in [0,1].",11 tags={"team": "risk", "owner": "priya", "criticality": "tier-1"},12 )13except mlflow.exceptions.MlflowException:14 pass # already exists1516# 2. Register a specific run's artefact as a new version.17run_id = "6f2a91c0d83f45e6a1b9c4d7e0f2a380"18mv = mlflow.register_model(19 model_uri=f"runs:/{run_id}/model",20 name="fraud-scorer",21)22print(f"created version {mv.version}")2324# 3. Attach the evidence a reviewer needs, on the VERSION not the model.25run = client.get_run(run_id)26client.update_model_version(27 name="fraud-scorer", version=mv.version,28 description=(29 "GBM, depth 3, lr 0.05. Trained on transactions 2023-10-01 to 2024-03-01. "30 "Change vs v13: added merchant_category_risk feature."31 ),32)33for key, value in {34 "git_sha": run.data.tags["git_sha"],35 "data_hash": run.data.tags["data_hash"],36 "val_auc": f"{run.data.metrics['val_auc']:.4f}",37 "val_recall_at_p30": f"{run.data.metrics['val_recall_at_p30']:.4f}",38 "worst_slice": "region_north",39 "worst_slice_auc": f"{run.data.metrics['slice.region_north.auc']:.4f}",40 "approved_by": "", # filled in at promotion time41}.items():42 client.set_model_version_tag("fraud-scorer", mv.version, key, value)The tags are not decoration. At 02:00, the person deciding whether to roll back to version 13 needs to know what 13 was and how it differed, without opening a browser and cross-referencing a run. Version tags put that on the object itself.
Stages versus aliases: why the mechanism changed
MLflow originally offered four fixed stages — None, Staging, Production, Archived — and a version could be moved between them. That design has three problems, all of which the 09:14 incident illustrates.
| Problem with stages | Consequence | How aliases fix it |
|---|---|---|
| A stage is a set: multiple versions can occupy it | models:/m/Production silently means "latest of several" | An alias points at exactly one version. Setting it elsewhere moves it |
| The four names are fixed | No way to express champion, challenger, shadow, eu-production | Aliases are arbitrary strings; define what you need |
| Stage conflates "where it runs" with "what state it is in" | Cannot mark a version "validated" without also deploying it | Aliases carry deployment intent; tags carry state |
Stages have been deprecated since MLflow 2.9 in favour of aliases plus tags. MLflow 3 still accepts the stage calls, with a warning that they will be removed in a future major release, so treat any stage-based code you inherit as a migration to do. The distinction to hold onto:
- An alias is a movable pointer that answers "which version plays this role?" —
@champion,@challenger,@shadow. One alias, one version, always. - A tag is a fact about a version that does not move —
validated=true,security_scan=passed,approved_by=nadia.
Side by side
1# DEPRECATED — the mechanism behind the 09:14 incident2client.transition_model_version_stage(3 name="fraud-scorer", version=14, stage="Production",4 archive_existing_versions=True, # forgetting this is the bug5)6model = mlflow.pyfunc.load_model("models:/fraud-scorer/Production")# CURRENT — a pointer that cannot hold two valuesclient.set_registered_model_alias("fraud-scorer", "champion", 14)model = mlflow.pyfunc.load_model("models:/fraud-scorer@champion")Look at what disappeared. There is no archive_existing_versions flag to forget, because setting the alias to 14 is removing it from 11. The failure mode is not merely less likely; it is not expressible.
The best fix for a class of bug is not making it less likely. It is making it inexpressible.
The URI syntax differs by one character and it is worth burning into memory: models:/name/14 is a version number, models:/name@champion is an alias, and models:/name/Production is the deprecated stage form.
Reading the registry without deprecated calls
1from mlflow import MlflowClient2client = MlflowClient()34# The exact version behind an alias — this is what serving resolves.5champ = client.get_model_version_by_alias("fraud-scorer", "champion")6print(champ.version, champ.run_id, champ.tags.get("val_auc"))78# All versions of one model, newest first.9versions = client.search_model_versions(10 filter_string="name = 'fraud-scorer'",11 order_by=["version_number DESC"],12 max_results=20,13)14for v in versions:15 print(v.version, v.aliases, v.tags.get("val_auc"), v.tags.get("approved_by"))1617# Every model a team owns.18for m in client.search_registered_models(filter_string="tags.team = 'risk'"):19 print(m.name, m.aliases) # a dict: {"champion": "14", ...}Note v.aliases on each version. Because an alias can only sit on one version, that list is the deployment state — no reconciliation required, no ambiguity about which of two "Production" versions wins.
A promotion workflow with gates that actually block
Promotion is not one action. It is a sequence of checks, each of which can stop it, ending in a pointer move.
1from dataclasses import dataclass2from datetime import datetime, timezone3import mlflow4from mlflow import MlflowClient56@dataclass7class Gate:8 name: str9 passed: bool10 detail: str1112def promote_to_champion(name: str, candidate_version: int,13 approver: str, min_gain: float = 0.004) -> str:14 client = MlflowClient()15 cand = client.get_model_version(name, str(candidate_version))16 gates: list[Gate] = []1718 # Gate 1: the candidate must carry its evidence.19 required = {"val_auc", "git_sha", "data_hash", "worst_slice_auc"}20 missing = required - set(cand.tags)21 gates.append(Gate("evidence", not missing, f"missing tags: {sorted(missing)}"))2223 # Gate 2: it must beat the incumbent by more than measured noise.24 try:25 champ = client.get_model_version_by_alias(name, "champion")26 gain = float(cand.tags["val_auc"]) - float(champ.tags["val_auc"])27 gates.append(Gate("headline", gain >= min_gain,28 f"gain {gain:+.4f} vs required {min_gain:+.4f}"))29 # Gate 3: no important slice may regress.30 slice_gain = (float(cand.tags["worst_slice_auc"])31 - float(champ.tags["worst_slice_auc"]))32 gates.append(Gate("worst_slice", slice_gain >= -0.01,33 f"worst-slice change {slice_gain:+.4f}"))34 except mlflow.exceptions.MlflowException:35 gates.append(Gate("headline", True, "no incumbent - first champion"))3637 # Gate 4: it must have survived shadow traffic.38 gates.append(Gate("shadow", cand.tags.get("shadow_hours", "0") != "0",39 f"shadow_hours={cand.tags.get('shadow_hours', '0')}"))4041 failed = [g for g in gates if not g.passed]42 if failed:43 raise PermissionError("blocked: " +44 "; ".join(f"{g.name} ({g.detail})" for g in failed))4546 # Record who authorised it BEFORE moving the pointer.47 client.set_model_version_tag(name, cand.version, "approved_by", approver)48 client.set_model_version_tag(name, cand.version, "approved_at",49 datetime.now(timezone.utc).isoformat())5051 # Keep a breadcrumb for rollback, then move the pointer.52 try:53 prev = client.get_model_version_by_alias(name, "champion")54 client.set_registered_model_alias(name, "previous-champion", prev.version)55 except mlflow.exceptions.MlflowException:56 prev = None5758 client.set_registered_model_alias(name, "champion", cand.version)59 return (f"{name}@champion: "60 f"{prev.version if prev else 'none'} -> {cand.version}")The previous-champion alias, set one line before the promotion, is the highest-value line in the function. It converts rollback from an investigation into a lookup. At 02:00 the instruction is set champion = previous-champion, and nobody has to reconstruct history from audit logs.
The min_gain of 0.004 should not be a guess. Train the same configuration with five seeds and measure the spread. If you get AUC values of 0.887, 0.881, 0.893, 0.884 and 0.890, the mean is 4.435 / 5 = 0.8870, the deviations are 0.000, −0.006, +0.006, −0.003, +0.003, their squares sum to 0.000090, and the sample standard deviation is 0.000090/4=0.0047. A required gain of 0.004 is roughly one standard deviation — the minimum defensible bar. Anything smaller and you are promoting noise on a schedule.
Rollback: four mechanisms, four recovery times
Rollback speed is a design property you choose in advance, not something you improvise during an incident. The options trade recovery time against running cost.
| Mechanism | How it works | Typical recovery | Cost | Fails when |
|---|---|---|---|---|
| Alias re-point + restart | Move @champion back, roll the pods | 2–5 min | None | Pods cache the model at start-up and rollouts are slow |
| Hot reload | Service polls the alias every N seconds and swaps in place | ~1 poll interval | Slightly more complex service | New model has incompatible input schema |
| Blue-green | Both versions loaded; a router switch flips traffic | Seconds | 2× memory and compute, continuously | Model is too large to hold two copies |
| Traffic weights | Serving platform splits by percentage; set challenger to 0% | Seconds | Requires a platform that supports it | Sticky sessions pin users to the bad version |
The numbers behind "2–5 minutes" are worth making concrete. A rolling restart of 4 replicas with maxUnavailable: 0 replaces them one at a time. Each new pod pulls a cached image (about 5 s), starts Python and loads a 180 MB model (about 20 s), then waits for its readiness probe to pass twice at a 10 s interval (about 20 s). That is roughly 45 s per pod, and 4 × 45 = 180 seconds. Blue-green skips all of it because both versions are already resident: the switch is a router config change, well under 5 seconds, paid for with double the memory footprint all year round.
Whether that trade is worth it is arithmetic you can do. Suppose a bad model costs you 3% extra false rejections on 1,200 requests per minute, and each false rejection loses USD 2.40 of margin. The bleed is 1,200 × 0.03 × 2.40 = USD 86.40 per minute. A 3-minute rolling rollback costs about USD 259. A 4-hour manual scramble — the kind you get with no previous-champion pointer — costs 240 × 86.40 = USD 20,736. Running a second copy of a 180 MB model costs a few dollars a month.
Rollback speed is not discovered during an incident. It is a design decision you made months earlier, whether or not you noticed making it.
1def rollback(name: str, reason: str) -> str:2 client = MlflowClient()3 bad = client.get_model_version_by_alias(name, "champion")4 prev = client.get_model_version_by_alias(name, "previous-champion")56 client.set_registered_model_alias(name, "champion", prev.version)7 client.set_model_version_tag(name, bad.version, "rolled_back_at",8 datetime.now(timezone.utc).isoformat())9 client.set_model_version_tag(name, bad.version, "rollback_reason", reason)10 client.set_model_version_tag(name, bad.version, "blocked", "true")11 return f"{name}@champion: {bad.version} -> {prev.version} ({reason})"Tagging the bad version blocked=true matters more than it looks. Without it, an automated retraining pipeline can happily re-promote the same broken version an hour later, and the incident repeats with the on-call engineer wondering whether the rollback took effect at all.
Hot reload without a restart
1import logging, threading, time2import mlflow3from mlflow import MlflowClient45log = logging.getLogger("serving")67class AliasedModel:8 """Serves the version behind an alias; swaps when the alias moves."""910 def __init__(self, name: str, alias: str = "champion", poll_seconds: int = 30):11 self.name, self.alias = name, alias12 self._client = MlflowClient()13 self._lock = threading.RLock()14 self._version = None15 self._model = None16 self._refresh()17 threading.Thread(target=self._loop, args=(poll_seconds,),18 daemon=True).start()1920 def _refresh(self) -> None:21 mv = self._client.get_model_version_by_alias(self.name, self.alias)22 if mv.version == self._version:23 return24 loaded = mlflow.pyfunc.load_model(f"models:/{self.name}@{self.alias}")25 with self._lock: # swap only after a successful load26 self._model, self._version = loaded, mv.version2728 def _loop(self, seconds: int) -> None:29 while True:30 time.sleep(seconds)31 try:32 self._refresh()33 except Exception as exc: # never let polling kill serving34 log.warning("model refresh failed, keeping v%s: %s",35 self._version, exc)3637 def predict(self, X):38 with self._lock:39 model, version = self._model, self._version40 return model.predict(X), versionTwo details make this safe. The new model is fully loaded before the lock is taken, so a slow or failing load never blocks or breaks live requests. And predict returns the version alongside the prediction, so every logged prediction carries the model that produced it — without which you cannot attribute a metric change to a deployment after the fact.
Canary: reducing the blast radius on the way in
Rollback limits how long a bad model serves. A canary limits how much traffic it serves while you find out. Together they bound the damage.
1STEPS = [(0.05, 20), (0.20, 30), (0.50, 60), (1.00, None)]23def canary(name: str, candidate_version: int) -> str:4 client = MlflowClient()5 client.set_registered_model_alias(name, "challenger", candidate_version)6 base = measure_health(alias="champion", minutes=10)78 for share, hold in STEPS:9 set_traffic_weight(challenger_share=share)10 if hold is None:11 client.set_registered_model_alias(name, "previous-champion",12 client.get_model_version_by_alias(name, "champion").version)13 client.set_registered_model_alias(name, "champion", candidate_version)14 client.delete_registered_model_alias(name, "challenger")15 return "promoted"1617 obs = measure_health(alias="challenger", minutes=hold)18 breach = (obs.error_rate > base.error_rate * 1.519 or obs.p99_ms > base.p99_ms * 1.320 or abs(obs.mean_score - base.mean_score) > 0.05)21 if breach:22 set_traffic_weight(challenger_share=0.0)23 client.set_model_version_tag(name, str(candidate_version),24 "canary_failed", str(obs))25 client.set_model_version_tag(name, str(candidate_version),26 "blocked", "true")27 return f"aborted at {share:.0%}: {obs}"28 return "promoted"The mean_score check is the ML-specific one and the reason a generic deployment canary is not enough. A model can be perfectly healthy by every operational measure — no errors, fast responses — while shifting its average predicted probability from 0.11 to 0.29. Nothing in the infrastructure notices; every downstream threshold is now wrong. Watching the prediction distribution catches within minutes what waiting for labels would catch in weeks.
Be honest about what a canary can and cannot detect. At 200 requests per second, 20 minutes at 5% traffic is 20 × 60 × 200 × 0.05 = 12,000 requests. That is plenty to see an error rate move from 0.5% to 2%, and nowhere near enough to detect a one-point drop in precision when labels arrive three days later. Canaries catch operational and distributional failures fast; accuracy regressions still need shadow evaluation and time.
Where registries get misused
| Belief | Reality |
|---|---|
| "The registry deploys the model" | It records which version is authorised. Something else has to notice and act — a poller, a GitOps agent, a serving platform. A registry alone changes nothing in production. |
| "Rollback means retraining the old model" | Retraining reproduces an approximation, takes hours, and needs the old data. Rollback means pointing at bytes that already exist. If your answer to an incident is "retrain", you do not have rollback. |
| "A version can be updated in place" | Version artefacts are immutable by design. Fixing version 14 means creating version 15. This is the property that makes the audit trail trustworthy. |
| "Aliases are just renamed stages" | A stage is a set with many members; an alias is a pointer with exactly one target. That difference is the entire bug class the 09:14 incident belongs to. |
| "Deleting an old version tidies the registry" | It destroys your rollback target and your audit trail. Tag it blocked=true and leave it. Storage is cheaper than an outage. |
| "The registry knows what is live" | Only if serving resolves through it every time. A pod that loaded a model at start-up and never re-checks can serve a version the registry no longer names. |
Designing for the 02:00 version of yourself
Every choice here should be judged against one scenario: it is 02:00, you were asleep ninety seconds ago, and something is wrong with the model. Under those conditions, three properties matter more than everything else combined.
The rollback target must already be named. Not derivable from logs, not reconstructable from the registry's history — named, as a pointer, set at promotion time. previous-champion costs one line and removes the entire investigation phase of an incident.
The rollback command must be the same command you use routinely. If the only way to revert is a procedure that has been executed twice, both times in a drill, it will not work at 02:00. Promotion and rollback should be the same operation with different arguments, exercised weekly.
Serving must resolve the alias, not a version number. A deployment manifest that hardcodes models:/fraud-scorer/14 requires a code change, a build, a review and a deploy to roll back. Resolving models:/fraud-scorer@champion with a hot reload turns rollback into a metadata write that takes effect in one poll interval.
Put those three in place and the 09:14 incident cannot happen, because there is no ambiguity about which version serves; and if a genuinely bad model does get promoted, the bleed is measured in minutes rather than hours. That is the whole return on a registry — not the tidy list of models in a web UI, but the fact that undoing a decision costs one line and thirty seconds.