Course Content
MLOps for AI
4 sections · 9 lessons
ML Lifecycle Management — Notebook to Production
On 3 March, Priya trains a churn model in a Jupyter notebook. Five-fold cross-validated AUC comes out at 0.871. The business likes the number, so the pickled model gets copied to a server and wired behind a Flask endpoint. Everyone moves on.
On 17 April the model needs retraining on fresher data. Priya opens the same notebook, changes nothing except the date range in the SQL cell, and runs it top to bottom. AUC: 0.794. She reruns it. 0.802. She reverts the date range to the original window — the exact same data as March — and gets 0.811, not 0.871.
It takes three days to find the three separate causes:
- Some time in late March she ran
pip install -U scikit-learnfor an unrelated project. Version 1.3.2 became 1.5.0, and the default handling of a categorical encoder changed. - The "same" SQL was not the same. It filtered on
signup_date >= '2024-01-01'with no upper bound, so the March run and the April run pulled different row counts from a table that is still being written to. - No random seed was set anywhere. The train/test split, the model's internal bootstrap sampling, and the shuffling in cross-validation were all different every run.
None of these is a modelling mistake. Priya's feature engineering was fine, her validation strategy was fine, her choice of algorithm was fine. What failed was everything around the model: the code that decides which library version, which rows, which random draw. That surrounding machinery is what lifecycle management is about, and it is where most of the pain in production machine learning actually lives.
A model is not an artefact. It is the output of a function whose inputs are data, code, configuration, environment and randomness. If you do not version all five, you cannot reproduce the output.
The four phases, and what each one owes the next
It is tempting to draw the ML lifecycle as a straight line: get data, train, deploy, done. It is more useful to think of four phases, each of which produces artefacts the next phase depends on. When a phase produces sloppy artefacts, the failure shows up two phases later and looks like something else entirely.
Phase 1 — Development
This is where you turn a business question into a learning problem. What is the label? Churn within 30 days of the observation date, or churn ever? What is the unit of prediction — one row per customer, or one row per customer per month? Which features can be computed at prediction time without leaking the future?
That last question is the one that ruins the most projects. A feature like total_support_tickets computed over the customer's whole history is fine in a training table built retrospectively and catastrophic in production, because at scoring time you only have the tickets up to now. Models trained on leaked features look brilliant offline and useless live. The offline AUC of 0.94 becomes 0.61 in the first week and no amount of infrastructure will save it.
Development ends when you have a training script that runs unattended and a written definition of the label and the feature cut-off.
Phase 2 — Validation and testing
Validation asks two different questions that people routinely conflate:
- Is the model good enough? Held-out metrics, calibration, performance on the slices that matter (new customers, enterprise accounts, each region).
- Does the code work? Does the preprocessing pipeline handle a null in a column that was never null in training? Does the serving code produce the same prediction as the training code for the same input?
The second question is a software testing problem and deserves ordinary unit tests. The single highest-value test in all of ML engineering is the train/serve consistency test: take ten rows, run them through the training preprocessing and the serving preprocessing, and assert the predictions match to within a tiny tolerance. It catches the entire category of bugs where the training pipeline scales a feature and the serving path forgets to.
Phase 3 — Deployment
Deployment is the act of making the model reachable and, crucially, replaceable. A deployment that cannot be reverted in under five minutes is not a deployment, it is a commitment. The artefacts here are a container image, a model file with a version identifier, and a manifest that says which image and which model version are currently live.
Phase 4 — Monitoring and maintenance
Software either works or throws an exception. A model can be silently, gradually wrong. It returns a well-formed probability of 0.31 for every request while the world has shifted underneath it. Monitoring therefore has to watch three separate things: that the service is up, that the inputs still look like training data, and that the predictions are still accurate once labels arrive.
| Phase | Question it answers | Artefacts it must hand on | Classic failure if done badly |
|---|---|---|---|
| Development | What are we predicting, from what? | Training script, label definition, feature cut-off rules | Target leakage — brilliant offline, useless live |
| Validation | Is it good enough, and does the code work? | Metrics per slice, test suite, train/serve consistency check | Preprocessing mismatch between training and serving |
| Deployment | How do users reach it, and how do we undo it? | Container image digest, model version, deployment manifest | Cannot roll back; a bad model stays live for hours |
| Monitoring | Is it still right? | Input distribution baselines, prediction logs, alert thresholds | Silent degradation discovered by a customer, not a dashboard |
The loop closes: monitoring detects degradation, which sends you back to development with new data. That is why "lifecycle" and not "pipeline".
Structuring a project so it can actually be operated
A notebook is a fine place to think and a terrible place to keep a system. The problem is not aesthetic. A notebook has hidden state — cell 12 might depend on a variable defined in cell 3 that you have since edited — and it cannot be imported, tested or run from a scheduler without heroics.
The layout below is not sacred, but the separation it encodes is:
churn/├── configs/│ ├── base.yaml # shared defaults│ ├── train_xgb.yaml # one file per experiment│ └── serve.yaml├── src/churn/│ ├── data.py # loading + splitting only│ ├── features.py # transformations, importable by BOTH│ ├── train.py # fit, evaluate, save│ ├── predict.py # load, transform, score│ └── schema.py # expected columns, dtypes, ranges├── tests/│ ├── test_features.py│ └── test_train_serve_parity.py├── pipelines/ # orchestration DAG definitions├── docker/│ ├── train.Dockerfile│ └── serve.Dockerfile├── requirements.in # what you asked for├── requirements.txt # what you got, pinned + hashed└── notebooks/ # exploration only, never importedThree rules make this structure earn its keep:
features.pyis imported by bothtrain.pyandpredict.py. Not copied. Not reimplemented "more efficiently" for serving. One implementation, one behaviour.- No path, hostname, threshold or hyperparameter is written in a
.pyfile. They all live inconfigs/. Code that contains/Users/priya/data/churn_v3_final.csvworks on exactly one machine. notebooks/is a leaf. Nothing imports from it. It may import fromsrc/. The moment a notebook contains logic that production needs, that logic moves intosrc/.
Dependency management: pinning is not optional
Priya's scikit-learn upgrade is the most common reproducibility failure in the field, and it has a precise fix. The key idea is separating what you asked for from what you got.
requirements.in holds your intent, loosely:
scikit-learn>=1.5,<1.6pandas>=2.2,<3xgboost>=2.1,<3requirements.txt is generated from it and holds the exact resolved graph, including transitive dependencies you never named:
1# with uv (fast) or pip-tools2uv pip compile requirements.in -o requirements.txt --generate-hashes34# in CI / Docker, install ONLY from the lock5uv pip sync requirements.txtThe resulting file pins every package to an exact version and its SHA-256 hash. If PyPI ever serves a different artefact for the same version number, installation fails loudly instead of silently changing your model.
| Approach | Reproducible? | Catches transitive drift? | Honest verdict |
|---|---|---|---|
pip install scikit-learn | No | No | Guarantees the Priya failure eventually |
scikit-learn==1.5.0 in requirements.txt, hand-written | Partly | No — numpy, scipy, joblib still float | Better; still breaks on a numpy major release |
| Compiled lock file with hashes | Yes, on the same OS/arch | Yes | The minimum bar for anything scheduled |
| Lock file inside a container pinned by digest | Yes, everywhere | Yes | What production should use |
That last row matters more than people expect. FROM python:3.11-slim is not a fixed base image — the tag is repointed at new builds regularly, so the same Dockerfile produces a different image next month. Pinning by digest fixes it:
FROM python:3.11-slim@sha256:2f4c...bd91A version tag is a promise about intent; a hash is a statement of fact. Reproducibility is built on hashes.
Configuration: one source of truth, overridable
Configuration should be data, validated on load, and reachable from both the CLI and the scheduler. A layered YAML file plus a typed loader gives you that in about forty lines.
# configs/train_xgb.yamldata: source_table: analytics.customer_features as_of_date: "2024-04-15" # explicit upper bound, never "now" test_size: 0.2model: type: xgboost params: n_estimators: 400 max_depth: 6 learning_rate: 0.05training: seed: 42 cv_folds: 51from dataclasses import dataclass, field2import hashlib, json, yaml34@dataclass(frozen=True)5class DataCfg:6 source_table: str7 as_of_date: str8 test_size: float = 0.2910@dataclass(frozen=True)11class TrainCfg:12 seed: int = 4213 cv_folds: int = 51415@dataclass(frozen=True)16class Config:17 data: DataCfg18 model: dict19 training: TrainCfg2021 @property22 def fingerprint(self) -> str:23 """Stable hash of the whole config — goes into the run name."""24 blob = json.dumps(self.__dict__, sort_keys=True, default=vars)25 return hashlib.sha256(blob.encode()).hexdigest()[:10]2627def load(path: str) -> Config:28 raw = yaml.safe_load(open(path))29 cfg = Config(30 data=DataCfg(**raw["data"]),31 model=raw["model"],32 training=TrainCfg(**raw["training"]),33 )34 if not 0.05 <= cfg.data.test_size <= 0.5:35 raise ValueError(f"test_size {cfg.data.test_size} outside sane range")36 return cfgTwo details are doing real work here. frozen=True means no part of the code can quietly mutate the config halfway through a run, so what you logged at the start is what actually ran. And fingerprint gives every distinct configuration a short stable identity — xgb_a91f3c2b04 — which you attach to the run, the model file and the metrics row. Two runs with the same fingerprint should produce the same model. If they do not, you have found a reproducibility leak.
Note also as_of_date: "2024-04-15". Never now(). A query that means something different depending on when you run it makes reruns meaningless — this was Priya's second bug.
Taming randomness — and knowing how much is left
Setting a seed is the easy half. Knowing what seeds do not control is the half that matters.
1import os, random2import numpy as np34def set_seeds(seed: int = 42) -> None:5 os.environ["PYTHONHASHSEED"] = str(seed) # must be set before interpreter work6 random.seed(seed)7 np.random.seed(seed)8 try:9 import torch10 torch.manual_seed(seed)11 torch.cuda.manual_seed_all(seed)12 torch.backends.cudnn.deterministic = True13 torch.backends.cudnn.benchmark = False # benchmark picks kernels non-deterministically14 torch.use_deterministic_algorithms(True, warn_only=True)15 except ImportError:16 passThings this still does not fix:
- Parallel floating-point reduction. Summing a million floats in a different order gives a slightly different answer, because floating-point addition is not associative. Multi-threaded XGBoost and GPU reductions both do this. The difference is around the 15th decimal place per operation, but a decision boundary sitting exactly on a threshold can flip.
- DataLoader worker ordering. With
num_workers > 0you need aworker_init_fnand a seededgenerator, or each worker reseeds from entropy. - The data itself. If the source table is still receiving writes, no seed on earth helps. Snapshot it, hash the snapshot, and record the hash.
- Hardware. A model trained on an A100 and one trained on a T4 can differ in the last bits. Same code, same seed, different silicon.
How much variation is normal? Measure it.
Run the identical configuration with five different seeds and record test AUC:
seed 0 : 0.871seed 1 : 0.864seed 2 : 0.879seed 3 : 0.868seed 4 : 0.873The mean is (0.871 + 0.864 + 0.879 + 0.868 + 0.873) / 5 = 4.355 / 5 = 0.8710. The deviations from the mean are 0.000, −0.007, +0.008, −0.003, +0.002; their squares sum to 0.000126, so the sample standard deviation is 0.000126/4=0.0056.
That single number changes how you read every experiment afterwards. A colleague's new feature set scores 0.876 against the 0.871 baseline. The improvement is +0.005 — less than one standard deviation of pure seed noise. It is not evidence of anything. To detect a genuine +0.005 improvement you would need the standard error of the mean to be well under that, and with five seeds the standard error is 0.0056/5=0.0025, so the two means are barely two standard errors apart even before accounting for the baseline's own uncertainty.
Until you know your seed-to-seed standard deviation, you cannot tell an improvement from a coincidence — and teams routinely ship coincidences.
Putting it together: a run that can be reconstructed
The pieces combine into a training entry point where every input is captured alongside the output.
1import hashlib, json, subprocess, time2from pathlib import Path3import pandas as pd45def git_sha() -> str:6 return subprocess.check_output(7 ["git", "rev-parse", "HEAD"], text=True).strip()89def git_dirty() -> bool:10 return bool(subprocess.check_output(11 ["git", "status", "--porcelain"], text=True).strip())1213def frame_hash(df: pd.DataFrame) -> str:14 return hashlib.sha256(15 pd.util.hash_pandas_object(df, index=True).values.tobytes()16 ).hexdigest()[:16]1718def train(cfg_path: str, out_dir: str) -> dict:19 cfg = load(cfg_path)20 if git_dirty():21 raise RuntimeError("Refusing to train from a dirty working tree")2223 set_seeds(cfg.training.seed)24 df = load_snapshot(cfg.data.source_table, cfg.data.as_of_date)2526 X_tr, X_te, y_tr, y_te = split(df, cfg.data.test_size, cfg.training.seed)27 model = build(cfg.model)28 model.fit(X_tr, y_tr)29 auc = evaluate(model, X_te, y_te)3031 run_id = f"{cfg.model['type']}_{cfg.fingerprint}_{int(time.time())}"32 out = Path(out_dir) / run_id33 out.mkdir(parents=True)3435 save_model(model, out / "model.joblib")36 manifest = {37 "run_id": run_id,38 "git_sha": git_sha(),39 "config_fingerprint": cfg.fingerprint,40 "config": json.loads(json.dumps(cfg, default=vars)),41 "data_hash": frame_hash(df),42 "rows": len(df),43 "seed": cfg.training.seed,44 "metrics": {"test_auc": round(auc, 4)},45 "model_sha256": sha256_file(out / "model.joblib"),46 }47 (out / "manifest.json").write_text(json.dumps(manifest, indent=2))48 return manifestThe git_dirty() check looks pedantic and is the most valuable four lines in the file. A run trained from uncommitted changes has a git_sha that is a lie: the commit it names does not contain the code that ran. Every reproduction attempt from that manifest will fail confusingly.
The manifest is the thing that would have saved Priya three days. Given git_sha, config_fingerprint, data_hash and a locked environment, an April rerun that produces a different AUC tells you immediately which of the four changed.
Named failure modes and how they present
| Symptom | Usual cause | Fix |
|---|---|---|
| Same code, same data, different metric on rerun | Unseeded randomness, or floating-point non-determinism | Central set_seeds(); measure and report seed variance |
| Reproduces on your laptop, not in CI | Unpinned transitive dependency, or unpinned base image tag | Hash-pinned lock file; base image by digest |
| Offline AUC 0.94, live AUC 0.61 | Target leakage — a feature that encodes the future | Enforce a feature cut-off time; audit every feature's availability at scoring time |
| Predictions differ between batch and API for identical input | Two implementations of the same preprocessing | Shared features.py; a parity test in CI |
| Metric drops but nothing was deployed | Upstream data change — a schema edit, a new NULL, a unit change | Schema and range assertions on load; alert on input distribution shift |
| Cannot say which model is live | No deployment manifest linking image digest to model version | Declare the live version in version control |
What this changes about how you work
The practical test of whether a project has lifecycle management is not whether it uses any particular tool. It is this: pick a model that is currently serving traffic and try to rebuild it from scratch. Not approximately — bit-for-bit, or close enough that the metrics match to three decimal places. If you can do that in an afternoon, you have it. If you cannot, everything downstream is guesswork, because you cannot attribute a change in behaviour to any specific cause.
Concretely, that means adopting a small number of habits before you adopt any platform. Commit before you train, and refuse to train from a dirty tree. Give every query an explicit upper date bound. Compile a lock file and install only from it. Put every knob in a config file and hash that file into the run identity. Set seeds in one place, then measure how much variance survives so you know what an honest improvement looks like. Write a manifest next to every model.
None of that requires infrastructure. It requires about a hundred lines of unglamorous Python, written before you need it. The teams that add it after their first Priya incident spend the same effort — they just spend three days of confusion first.