MLOps for AI

ML Lifecycle Management — Notebook to Production


On 3 March, Priya trains a churn model in a Jupyter notebook. Five-fold cross-validated AUC comes out at 0.871. The business likes the number, so the pickled model gets copied to a server and wired behind a Flask endpoint. Everyone moves on.

On 17 April the model needs retraining on fresher data. Priya opens the same notebook, changes nothing except the date range in the SQL cell, and runs it top to bottom. AUC: 0.794. She reruns it. 0.802. She reverts the date range to the original window — the exact same data as March — and gets 0.811, not 0.871.

It takes three days to find the three separate causes:

  • Some time in late March she ran pip install -U scikit-learn for an unrelated project. Version 1.3.2 became 1.5.0, and the default handling of a categorical encoder changed.
  • The "same" SQL was not the same. It filtered on signup_date >= '2024-01-01' with no upper bound, so the March run and the April run pulled different row counts from a table that is still being written to.
  • No random seed was set anywhere. The train/test split, the model's internal bootstrap sampling, and the shuffling in cross-validation were all different every run.

None of these is a modelling mistake. Priya's feature engineering was fine, her validation strategy was fine, her choice of algorithm was fine. What failed was everything around the model: the code that decides which library version, which rows, which random draw. That surrounding machinery is what lifecycle management is about, and it is where most of the pain in production machine learning actually lives.

A model is not an artefact. It is the output of a function whose inputs are data, code, configuration, environment and randomness. If you do not version all five, you cannot reproduce the output.

What each phase owes the one after itDevelopment — a run someone else can reconstructValidation — data tests,model tests, behaviour testsDeployment — a pinned image and a way backMonitoring — drift, performance and the label lag
The pickled notebook model skipped three of these, so nobody could say what data produced AUC 0.871.

The four phases, and what each one owes the next

It is tempting to draw the ML lifecycle as a straight line: get data, train, deploy, done. It is more useful to think of four phases, each of which produces artefacts the next phase depends on. When a phase produces sloppy artefacts, the failure shows up two phases later and looks like something else entirely.

Phase 1 — Development

This is where you turn a business question into a learning problem. What is the label? Churn within 30 days of the observation date, or churn ever? What is the unit of prediction — one row per customer, or one row per customer per month? Which features can be computed at prediction time without leaking the future?

That last question is the one that ruins the most projects. A feature like total_support_tickets computed over the customer's whole history is fine in a training table built retrospectively and catastrophic in production, because at scoring time you only have the tickets up to now. Models trained on leaked features look brilliant offline and useless live. The offline AUC of 0.94 becomes 0.61 in the first week and no amount of infrastructure will save it.

Development ends when you have a training script that runs unattended and a written definition of the label and the feature cut-off.

Phase 2 — Validation and testing

Validation asks two different questions that people routinely conflate:

  • Is the model good enough? Held-out metrics, calibration, performance on the slices that matter (new customers, enterprise accounts, each region).
  • Does the code work? Does the preprocessing pipeline handle a null in a column that was never null in training? Does the serving code produce the same prediction as the training code for the same input?

The second question is a software testing problem and deserves ordinary unit tests. The single highest-value test in all of ML engineering is the train/serve consistency test: take ten rows, run them through the training preprocessing and the serving preprocessing, and assert the predictions match to within a tiny tolerance. It catches the entire category of bugs where the training pipeline scales a feature and the serving path forgets to.

Phase 3 — Deployment

Deployment is the act of making the model reachable and, crucially, replaceable. A deployment that cannot be reverted in under five minutes is not a deployment, it is a commitment. The artefacts here are a container image, a model file with a version identifier, and a manifest that says which image and which model version are currently live.

Phase 4 — Monitoring and maintenance

Software either works or throws an exception. A model can be silently, gradually wrong. It returns a well-formed probability of 0.31 for every request while the world has shifted underneath it. Monitoring therefore has to watch three separate things: that the service is up, that the inputs still look like training data, and that the predictions are still accurate once labels arrive.

PhaseQuestion it answersArtefacts it must hand onClassic failure if done badly
DevelopmentWhat are we predicting, from what?Training script, label definition, feature cut-off rulesTarget leakage — brilliant offline, useless live
ValidationIs it good enough, and does the code work?Metrics per slice, test suite, train/serve consistency checkPreprocessing mismatch between training and serving
DeploymentHow do users reach it, and how do we undo it?Container image digest, model version, deployment manifestCannot roll back; a bad model stays live for hours
MonitoringIs it still right?Input distribution baselines, prediction logs, alert thresholdsSilent degradation discovered by a customer, not a dashboard

The loop closes: monitoring detects degradation, which sends you back to development with new data. That is why "lifecycle" and not "pipeline".

Structuring a project so it can actually be operated

A notebook is a fine place to think and a terrible place to keep a system. The problem is not aesthetic. A notebook has hidden state — cell 12 might depend on a variable defined in cell 3 that you have since edited — and it cannot be imported, tested or run from a scheduler without heroics.

The layout below is not sacred, but the separation it encodes is:

Text
churn/├── configs/│   ├── base.yaml              # shared defaults│   ├── train_xgb.yaml         # one file per experiment│   └── serve.yaml├── src/churn/│   ├── data.py                # loading + splitting only│   ├── features.py            # transformations, importable by BOTH│   ├── train.py               # fit, evaluate, save│   ├── predict.py             # load, transform, score│   └── schema.py              # expected columns, dtypes, ranges├── tests/│   ├── test_features.py│   └── test_train_serve_parity.py├── pipelines/                 # orchestration DAG definitions├── docker/│   ├── train.Dockerfile│   └── serve.Dockerfile├── requirements.in            # what you asked for├── requirements.txt           # what you got, pinned + hashed└── notebooks/                 # exploration only, never imported

Three rules make this structure earn its keep:

  1. features.py is imported by both train.py and predict.py. Not copied. Not reimplemented "more efficiently" for serving. One implementation, one behaviour.
  2. No path, hostname, threshold or hyperparameter is written in a .py file. They all live in configs/. Code that contains /Users/priya/data/churn_v3_final.csv works on exactly one machine.
  3. notebooks/ is a leaf. Nothing imports from it. It may import from src/. The moment a notebook contains logic that production needs, that logic moves into src/.

Dependency management: pinning is not optional

Priya's scikit-learn upgrade is the most common reproducibility failure in the field, and it has a precise fix. The key idea is separating what you asked for from what you got.

requirements.in holds your intent, loosely:

Text
scikit-learn>=1.5,<1.6pandas>=2.2,<3xgboost>=2.1,<3

requirements.txt is generated from it and holds the exact resolved graph, including transitive dependencies you never named:

Bash
# with uv (fast) or pip-toolsuv pip compile requirements.in -o requirements.txt --generate-hashes# in CI / Docker, install ONLY from the lockuv pip sync requirements.txt

The resulting file pins every package to an exact version and its SHA-256 hash. If PyPI ever serves a different artefact for the same version number, installation fails loudly instead of silently changing your model.

ApproachReproducible?Catches transitive drift?Honest verdict
pip install scikit-learnNoNoGuarantees the Priya failure eventually
scikit-learn==1.5.0 in requirements.txt, hand-writtenPartlyNo — numpy, scipy, joblib still floatBetter; still breaks on a numpy major release
Compiled lock file with hashesYes, on the same OS/archYesThe minimum bar for anything scheduled
Lock file inside a container pinned by digestYes, everywhereYesWhat production should use

That last row matters more than people expect. FROM python:3.11-slim is not a fixed base image — the tag is repointed at new builds regularly, so the same Dockerfile produces a different image next month. Pinning by digest fixes it:

Text
FROM python:3.11-slim@sha256:2f4c...bd91

A version tag is a promise about intent; a hash is a statement of fact. Reproducibility is built on hashes.

Configuration: one source of truth, overridable

Configuration should be data, validated on load, and reachable from both the CLI and the scheduler. A layered YAML file plus a typed loader gives you that in about forty lines.

Text
# configs/train_xgb.yamldata:  source_table: analytics.customer_features  as_of_date: "2024-04-15"      # explicit upper bound, never "now"  test_size: 0.2model:  type: xgboost  params:    n_estimators: 400    max_depth: 6    learning_rate: 0.05training:  seed: 42  cv_folds: 5
Python
from dataclasses import dataclass, fieldimport hashlib, json, yaml@dataclass(frozen=True)class DataCfg:    source_table: str    as_of_date: str    test_size: float = 0.2@dataclass(frozen=True)class TrainCfg:    seed: int = 42    cv_folds: int = 5@dataclass(frozen=True)class Config:    data: DataCfg    model: dict    training: TrainCfg    @property    def fingerprint(self) -> str:        """Stable hash of the whole config — goes into the run name."""        blob = json.dumps(self.__dict__, sort_keys=True, default=vars)        return hashlib.sha256(blob.encode()).hexdigest()[:10]def load(path: str) -> Config:    raw = yaml.safe_load(open(path))    cfg = Config(        data=DataCfg(**raw["data"]),        model=raw["model"],        training=TrainCfg(**raw["training"]),    )    if not 0.05 <= cfg.data.test_size <= 0.5:        raise ValueError(f"test_size {cfg.data.test_size} outside sane range")    return cfg

Two details are doing real work here. frozen=True means no part of the code can quietly mutate the config halfway through a run, so what you logged at the start is what actually ran. And fingerprint gives every distinct configuration a short stable identity — xgb_a91f3c2b04 — which you attach to the run, the model file and the metrics row. Two runs with the same fingerprint should produce the same model. If they do not, you have found a reproducibility leak.

Note also as_of_date: "2024-04-15". Never now(). A query that means something different depending on when you run it makes reruns meaningless — this was Priya's second bug.

Taming randomness — and knowing how much is left

Setting a seed is the easy half. Knowing what seeds do not control is the half that matters.

Python
import os, randomimport numpy as npdef set_seeds(seed: int = 42) -> None:    os.environ["PYTHONHASHSEED"] = str(seed)   # must be set before interpreter work    random.seed(seed)    np.random.seed(seed)    try:        import torch        torch.manual_seed(seed)        torch.cuda.manual_seed_all(seed)        torch.backends.cudnn.deterministic = True        torch.backends.cudnn.benchmark = False   # benchmark picks kernels non-deterministically        torch.use_deterministic_algorithms(True, warn_only=True)    except ImportError:        pass

Things this still does not fix:

  • Parallel floating-point reduction. Summing a million floats in a different order gives a slightly different answer, because floating-point addition is not associative. Multi-threaded XGBoost and GPU reductions both do this. The difference is around the 15th decimal place per operation, but a decision boundary sitting exactly on a threshold can flip.
  • DataLoader worker ordering. With num_workers > 0 you need a worker_init_fn and a seeded generator, or each worker reseeds from entropy.
  • The data itself. If the source table is still receiving writes, no seed on earth helps. Snapshot it, hash the snapshot, and record the hash.
  • Hardware. A model trained on an A100 and one trained on a T4 can differ in the last bits. Same code, same seed, different silicon.

How much variation is normal? Measure it.

Run the identical configuration with five different seeds and record test AUC:

Text
seed 0 : 0.871seed 1 : 0.864seed 2 : 0.879seed 3 : 0.868seed 4 : 0.873

The mean is (0.871 + 0.864 + 0.879 + 0.868 + 0.873) / 5 = 4.355 / 5 = 0.8710. The deviations from the mean are 0.000, −0.007, +0.008, −0.003, +0.002; their squares sum to 0.000126, so the sample standard deviation is 0.000126/4=0.0056\sqrt{0.000126 / 4} = 0.0056.

That single number changes how you read every experiment afterwards. A colleague's new feature set scores 0.876 against the 0.871 baseline. The improvement is +0.005 — less than one standard deviation of pure seed noise. It is not evidence of anything. To detect a genuine +0.005 improvement you would need the standard error of the mean to be well under that, and with five seeds the standard error is 0.0056/5=0.00250.0056/\sqrt{5} = 0.0025, so the two means are barely two standard errors apart even before accounting for the baseline's own uncertainty.

Until you know your seed-to-seed standard deviation, you cannot tell an improvement from a coincidence — and teams routinely ship coincidences.

Putting it together: a run that can be reconstructed

The pieces combine into a training entry point where every input is captured alongside the output.

Python
import hashlib, json, subprocess, timefrom pathlib import Pathimport pandas as pddef git_sha() -> str:    return subprocess.check_output(        ["git", "rev-parse", "HEAD"], text=True).strip()def git_dirty() -> bool:    return bool(subprocess.check_output(        ["git", "status", "--porcelain"], text=True).strip())def frame_hash(df: pd.DataFrame) -> str:    return hashlib.sha256(        pd.util.hash_pandas_object(df, index=True).values.tobytes()    ).hexdigest()[:16]def train(cfg_path: str, out_dir: str) -> dict:    cfg = load(cfg_path)    if git_dirty():        raise RuntimeError("Refusing to train from a dirty working tree")    set_seeds(cfg.training.seed)    df = load_snapshot(cfg.data.source_table, cfg.data.as_of_date)    X_tr, X_te, y_tr, y_te = split(df, cfg.data.test_size, cfg.training.seed)    model = build(cfg.model)    model.fit(X_tr, y_tr)    auc = evaluate(model, X_te, y_te)    run_id = f"{cfg.model['type']}_{cfg.fingerprint}_{int(time.time())}"    out = Path(out_dir) / run_id    out.mkdir(parents=True)    save_model(model, out / "model.joblib")    manifest = {        "run_id": run_id,        "git_sha": git_sha(),        "config_fingerprint": cfg.fingerprint,        "config": json.loads(json.dumps(cfg, default=vars)),        "data_hash": frame_hash(df),        "rows": len(df),        "seed": cfg.training.seed,        "metrics": {"test_auc": round(auc, 4)},        "model_sha256": sha256_file(out / "model.joblib"),    }    (out / "manifest.json").write_text(json.dumps(manifest, indent=2))    return manifest

The git_dirty() check looks pedantic and is the most valuable four lines in the file. A run trained from uncommitted changes has a git_sha that is a lie: the commit it names does not contain the code that ran. Every reproduction attempt from that manifest will fail confusingly.

The manifest is the thing that would have saved Priya three days. Given git_sha, config_fingerprint, data_hash and a locked environment, an April rerun that produces a different AUC tells you immediately which of the four changed.

Named failure modes and how they present

SymptomUsual causeFix
Same code, same data, different metric on rerunUnseeded randomness, or floating-point non-determinismCentral set_seeds(); measure and report seed variance
Reproduces on your laptop, not in CIUnpinned transitive dependency, or unpinned base image tagHash-pinned lock file; base image by digest
Offline AUC 0.94, live AUC 0.61Target leakage — a feature that encodes the futureEnforce a feature cut-off time; audit every feature's availability at scoring time
Predictions differ between batch and API for identical inputTwo implementations of the same preprocessingShared features.py; a parity test in CI
Metric drops but nothing was deployedUpstream data change — a schema edit, a new NULL, a unit changeSchema and range assertions on load; alert on input distribution shift
Cannot say which model is liveNo deployment manifest linking image digest to model versionDeclare the live version in version control

What this changes about how you work

The practical test of whether a project has lifecycle management is not whether it uses any particular tool. It is this: pick a model that is currently serving traffic and try to rebuild it from scratch. Not approximately — bit-for-bit, or close enough that the metrics match to three decimal places. If you can do that in an afternoon, you have it. If you cannot, everything downstream is guesswork, because you cannot attribute a change in behaviour to any specific cause.

Concretely, that means adopting a small number of habits before you adopt any platform. Commit before you train, and refuse to train from a dirty tree. Give every query an explicit upper date bound. Compile a lock file and install only from it. Put every knob in a config file and hash that file into the run identity. Set seeds in one place, then measure how much variance survives so you know what an honest improvement looks like. Write a manifest next to every model.

None of that requires infrastructure. It requires about a hundred lines of unglamorous Python, written before you need it. The teams that add it after their first Priya incident spend the same effort — they just spend three days of confusion first.