Course Content
Model Deployment for AI Engineers
4 sections · 10 lessons
Model Versioning and A/B Testing
A fraud team retrains their model on Monday and copies the new weights over /models/fraud/current.pkl. On Thursday, chargebacks are up 40%. Someone asks the obvious question: is it the new model?
Nobody can answer. The old file is gone. The training run that produced the new one used a data snapshot that has since been overwritten too. The commit hash of the training code was not recorded. The only way back is to retrain from scratch and hope it lands in the same place, which takes six hours and will not reproduce exactly because the data has moved.
The failure here is not the model. It might genuinely be innocent — a new fraud pattern, a payments-provider change, a seasonal effect. The failure is that the deployment left no way to find out. Every technique in this lesson exists to make that Thursday question answerable in minutes.
Why overwriting a file fails
cp new_model.pkl current.pkl looks like deployment. It has four defects, and each one costs you something specific.
| Defect | What it costs |
|---|---|
| No rollback | The previous model no longer exists. Recovery means retraining. |
| No attribution | You cannot correlate a metric change with a deployment, because there is no deployment timestamp tied to an artefact. |
| Not atomic | A process that loads the file mid-copy gets a truncated pickle and crashes. |
| No provenance | Which data, which code, which hyperparameters? Unrecorded, so unreproducible. |
The fix is a rule that sounds trivial and changes everything: artefacts are immutable and addressed by version; deployment is a change of pointer, not a change of file.
/models/fraud/├── 2.3.0/ model.onnx metadata.json metrics.json├── 2.4.0/ model.onnx metadata.json metrics.json├── 2.4.1/ model.onnx metadata.json metrics.json└── current -> 2.4.1 # a symlink, or a row in a tableRollback is now repointing the symlink: sub-second, and it cannot fail on the artefact side because the artefact was already there.
Versioning that carries meaning
Borrow semantic versioning, but define what each position means for a model, because the software definitions do not transfer cleanly.
| Bump | Meaning for a model | Example | Client impact |
|---|---|---|---|
| MAJOR | The input or output contract changed | Added a required feature; output went from 3 classes to 5 | Callers break. Coordinate. |
| MINOR | Same contract, retrained or architecture changed | Retrained on three more months of data | Same fields, different numbers. Recalibrate thresholds. |
| PATCH | No behaviour change to predictions | Quantized for speed; fixed a label-mapping bug in postprocessing | None expected. |
The distinction that trips people up is MINOR. A retrained model with an identical interface still changes every prediction slightly, so any downstream threshold tuned against the old score distribution is now mistuned. A fraud team using "block if score above 0.85" may find that 0.85 on version 2.4.0 corresponds to 0.79 on 2.5.0. Nothing errors; the block rate just doubles.
A retrained model is a new model, not a bug fix — treat any change to the score distribution as a change your downstream consumers must be told about.
A minimal registry
A registry is a table that records what exists, what it scored, and what is deployed where. You do not need a platform to start; you need these fields.
1import json, hashlib, sqlite3, datetime as dt23SCHEMA = """4CREATE TABLE IF NOT EXISTS models (5 version TEXT PRIMARY KEY,6 artefact_path TEXT NOT NULL,7 sha256 TEXT NOT NULL,8 framework TEXT NOT NULL,9 git_commit TEXT NOT NULL,10 training_data TEXT NOT NULL, -- snapshot id or date range11 metrics TEXT NOT NULL, -- JSON12 created_at TEXT NOT NULL,13 stage TEXT NOT NULL -- staging | production | archived14);15CREATE TABLE IF NOT EXISTS deployments (16 id INTEGER PRIMARY KEY,17 version TEXT NOT NULL,18 env TEXT NOT NULL,19 deployed_at TEXT NOT NULL,20 deployed_by TEXT NOT NULL,21 rolled_back_at TEXT22);23"""2425def register(db, version, path, framework, git_commit, data_snapshot, metrics):26 digest = hashlib.sha256(open(path, "rb").read()).hexdigest()27 db.execute(28 "INSERT INTO models VALUES (?,?,?,?,?,?,?,?,?)",29 (version, path, digest, framework, git_commit, data_snapshot,30 json.dumps(metrics), dt.datetime.utcnow().isoformat(), "staging"),31 )32 db.commit()33 return digestThe sha256 is not bureaucracy. When production behaviour does not match your offline evaluation, the first thing to rule out is that the file on the serving node is the file you think it is. A digest turns that from a debate into a one-line check.
Equally important: git_commit and training_data together are what make a model reproducible. Recording only the metrics tells you what happened; recording the inputs tells you how to make it happen again.
The deployments table is what answers the Thursday question. Overlay deployed_at on your chargeback graph and you either see the step change line up with a deployment or you do not — and either answer saves you a day.
A/B testing: comparing two models on live traffic
Offline evaluation tells you which model scores better on a held-out set. It does not tell you which model produces better outcomes, because the held-out set is historical and the deployed model changes the world it operates in. A fraud model that blocks more transactions changes what fraudsters attempt next. A recommender changes what users see and therefore what they click. Only live traffic settles it.
The assignment rule that matters
Here is the naive version, and it is wrong:
import randommodel = model_b if random.random() < 0.5 else model_a # WRONGThe same user gets a different model on every request. A user browsing five products sees recommendations from two different models, and any metric you compute per user is a blend of both arms. Worse, users notice inconsistency and their behaviour changes in response to that rather than to either model.
Assignment must be deterministic in the user identifier:
1import hashlib23def bucket(user_id: str, experiment: str, buckets: int = 100) -> int:4 """Stable 0-99 bucket. Same user + experiment always lands identically."""5 key = f"{experiment}:{user_id}".encode()6 return int(hashlib.md5(key).hexdigest()[:8], 16) % buckets78def assign(user_id: str, experiment: str, split: dict) -> str:9 b = bucket(user_id, experiment)10 cumulative = 011 for variant, pct in split.items():12 cumulative += pct13 if b < cumulative:14 return variant15 return list(split)[-1]1617SPLIT = {"control": 50, "treatment": 50}18assign("user_8814", "reco_v25_test", SPLIT) # same answer every call, every serverThree properties fall out of this. It needs no shared state, so every replica computes the same answer independently. It survives restarts and redeploys. And salting the hash with the experiment name means a user who landed in treatment for one experiment is not systematically in treatment for the next — without the salt, the same users are the guinea pigs forever, and your results are biased toward whatever is peculiar about them.
Reading the result is a statistics question
This is where most A/B tests go wrong: someone looks at the dashboard, sees treatment ahead, and ships. Work through a real case.
After two weeks:
| Arm | Users | Conversions | Rate |
|---|---|---|---|
| Control (model 2.4.1) | 10,000 | 412 | 4.12% |
| Treatment (model 2.5.0) | 10,000 | 468 | 4.68% |
That is a 13.6% relative lift. It looks like a clear win. Test it properly with a two-proportion z-test.
Pooled rate under the null hypothesis that both arms are the same:
Standard error of the difference:
A two-sided p-value for z=1.93 is about 0.054. Above the conventional 0.05 threshold. The honest reading is: this is suggestive, and we do not yet have enough data to call it. A gap this large would turn up about one time in eighteen even if the two models were identical — ship on evidence like this as a habit and you will regularly ship pure noise as a 13.6% improvement, and then defend it in a quarterly review.
from statsmodels.stats.proportion import proportions_ztestz, p = proportions_ztest(count=[468, 412], nobs=[10_000, 10_000])print(f"z={z:.3f} p={p:.4f}") # z=1.931 p=0.0535Decide the sample size before you start
The related mistake is peeking: checking daily and stopping the moment p drops below 0.05. If you look twenty times, you will cross 0.05 by chance eventually even when the arms are identical. The false-positive rate for a test peeked at daily for two weeks is closer to 25% than 5%.
Compute the required sample first. A workable approximation for a proportion, at 80% power and 5% significance:
To detect a 10% relative lift on a 4.0% baseline — an absolute δ=0.004:
At 5,000 eligible users a day split evenly, that is 76,800 users total, or about 15 days. Write that number down before launch, and do not read the result before you reach it.
An A/B test where you decided the stopping point after seeing the data is not an experiment; it is a search for a favourable moment to stop.
Two more practical constraints. Run for whole weeks, because weekday and weekend behaviour differ and a Tuesday-to-Friday test measures the mid-week population. And pick one primary metric before you start — testing ten metrics at 5% each means a 40% chance that at least one looks significant by luck.
Canary deployments: limiting blast radius
A/B testing answers "which is better". Canary deployment answers a narrower and more urgent question: "is the new model broken?" You send a small fraction of traffic to it, watch the operational metrics, and expand only if nothing is wrong.
Stage 1: 1% for 30 min → Stage 2: 5% for 1 hStage 3: 25% for 2 h → Stage 4: 50% for 2 h → Stage 5: 100%Any stage fails its gate → immediate rollback to 0%1STAGES = [(0.01, 30), (0.05, 60), (0.25, 120), (0.50, 120), (1.00, None)]23GATES = {4 "error_rate": lambda c, b: c <= max(b * 2, 0.01), # never worse than 2x baseline5 "p99_latency": lambda c, b: c <= b * 1.5,6 "null_rate": lambda c, b: c <= b + 0.005,7}8MIN_REQUESTS = 1000 # below this, do not act on the numbers910def evaluate_stage(canary, baseline):11 if canary["requests"] < MIN_REQUESTS:12 return "wait"13 for metric, gate in GATES.items():14 if not gate(canary[metric], baseline[metric]):15 return f"rollback: {metric} {canary[metric]:.4f} vs baseline {baseline[metric]:.4f}"16 return "promote"The MIN_REQUESTS guard is doing real work, and the arithmetic shows why. Suppose baseline error rate is 0.5%.
At 1,000 canary requests you expect 5 errors, with standard deviation 5=2.24. Observing 21 errors is (21−5)/2.24=7.1 standard deviations out — that is a real regression, roll back.
At 100 canary requests you expect 0.5 errors, standard deviation 0.5=0.71. Observing 2 errors is (2−0.5)/0.71=2.1 standard deviations — well within what noise produces routinely. Roll back on that and you will roll back healthy deployments constantly, then start ignoring the automation, which is worse than not having it.
Notice which metrics are on the gate list: error rate, latency, null-prediction rate. All operational, all measurable within minutes. Conversion rate is deliberately absent — it takes days to measure and belongs to the A/B test, not the canary.
Shadow deployment: production traffic, zero risk
Shadow mode sends every request to both models but returns only the current model's answer. The new model's predictions are logged and never used.
1@app.post("/predict")2def predict(req: PredictRequest, bg: BackgroundTasks): # plain def: the call blocks3 result = production_model.predict(req.features) # this is what the user gets4 bg.add_task(run_shadow, req.request_id, req.features, result)5 return result67def run_shadow(request_id, features, prod_result):8 try:9 with timeout(0.5): # never let shadow affect anything10 shadow_result = shadow_model.predict(features)11 log.info(json.dumps({12 "event": "shadow_comparison",13 "request_id": request_id,14 "prod": prod_result["label"], "shadow": shadow_result["label"],15 "agree": prod_result["label"] == shadow_result["label"],16 "prod_score": prod_result["score"], "shadow_score": shadow_result["score"],17 }))18 except Exception as e:19 log.warning("shadow failed: %s", e) # swallow: shadow must never pageWhat shadow mode is uniquely good at is catching problems that no offline test can see, because they come from the difference between your training data and live traffic: feature distributions that shifted, fields that are null in production but never in the training snapshot, latency under real concurrency and real payload sizes, and crashes on inputs your test set never contained.
Its hard limit is equally important: shadow mode cannot measure outcomes. The shadow model's predictions never reach a user, so nobody ever clicks or converts on them. You learn that it runs correctly and where it disagrees; you learn nothing about whether its disagreements are improvements. That question requires live traffic and therefore requires A/B.
Cost is the other consideration: you are paying for double the inference compute for the duration.
Feature flags: separating deploy from release
All of the above assumes you can change which model serves traffic without a redeploy. That is what a flag layer gives you — the code containing the new model ships to production, dormant, and a configuration change activates it.
1class ModelRouter:2 def __init__(self, config_source):3 self.config = config_source # a table, Redis key, or config service45 def select(self, user_id: str, context: dict) -> str:6 cfg = self.config.get("model_routing")78 if cfg.get("kill_switch"): # instant global revert9 return cfg["safe_version"]10 if user_id in cfg.get("force", {}): # pin internal testers11 return cfg["force"][user_id]12 if context.get("region") in cfg.get("region_overrides", {}):13 return cfg["region_overrides"][context["region"]]14 return assign(user_id, cfg["experiment"], cfg["split"])The kill switch is the part that earns its keep. Rolling back a container image takes several minutes: build, push, orchestrator rollout, health checks. Flipping a boolean in a config store takes seconds. During an incident that difference is the difference between a footnote and a postmortem.
The force map matters more than it looks too: it lets your own team hit the new model deterministically in production before any real user does.
Choosing between them
| Shadow | Canary | A/B test | Blue-green | |
|---|---|---|---|---|
| Question answered | Does it work on real inputs? | Is it broken? | Is it better? | Can we switch instantly? |
| User exposure | None | 1–50%, growing | 50% for the duration | 0% then 100% |
| Time to a verdict | Hours | Hours | 1–3 weeks | Immediate |
| Measures business outcomes | No | No | Yes | No |
| Compute cost | 2× | 1× | 1× | 2× during overlap |
| Rollback speed | N/A | Seconds | Seconds | Seconds |
They are not alternatives; they are stages. A mature release for a model that matters runs all of them in sequence: shadow for 48 hours to prove it executes correctly on live inputs, canary through 1/5/25/50% over a day to prove it is not operationally broken, then an A/B test at 50/50 for the pre-computed number of days to prove it is actually better — with a flag-based kill switch available at every stage.
What to put in place before your next deployment
The practical test of everything here is a single scenario: it is Thursday, a metric has moved, and someone asks whether the model did it. You should be able to answer within ten minutes, and answering requires exactly four things to already exist.
First, a deployment log — version, environment, timestamp, who — that you can overlay on any metric graph. Most incidents are resolved by noticing that the step change does or does not align with a deploy.
Second, the previous artefact still on disk, byte-identical, with its digest recorded. This is what makes rollback a thirty-second operation instead of a retraining project.
Third, the model version stamped on every prediction you log. Without it you cannot compute per-version metrics after the fact, which means every question about a past deployment requires a new experiment.
Fourth, a kill switch you have actually tested. An untested rollback path is not a rollback path. Exercise it in production during a quiet hour, on purpose, before you need it under pressure — the same way you would test a backup restore rather than trusting that the backups exist.
None of these require a platform, a vendor, or a migration. A table, a directory of immutable versions, one extra field in your prediction log, and a boolean in a config store cover all four. Teams that skip them are not saving work; they are deferring it to the worst possible moment.