Course Content
MLOps for AI
4 sections · 9 lessons
GitOps, Docker, and Model Versioning
It is 02:40 and the fraud model is rejecting 38% of legitimate transactions. The on-call engineer needs to roll back. She asks the obvious question in the incident channel: what was live before?
Nobody knows. The model file on the production box is called model_v2_final_FIXED.pkl, modified eleven days ago. There is a model_v2_final.pkl next to it, and a model_v2.pkl, and a model_backup.pkl. None of them has a matching entry anywhere in version control. The deploy was done by someone who is asleep, by copying a file over SSH. The container running the service was built from a Dockerfile that says pip install scikit-learn, so rebuilding it today would not produce the same image anyway.
The rollback takes four hours and is done by guessing. That outage is not a modelling problem or a Kubernetes problem. It is a problem of state that exists nowhere except on a machine — and the entire discipline covered here exists to eliminate that.
Three ideas do the work. GitOps makes the desired state of production a file in a repository. Docker makes the environment a reproducible artefact rather than a machine's accumulated history. Model versioning gives every trained model an identity you can name in a rollback command at 02:40.
GitOps: the repository is the source of truth
The core claim of GitOps is deceptively simple: the desired state of your system is described declaratively in Git, and an automated process continuously makes reality match it. Nobody changes production by typing commands at it.
The four principles
| Principle | What it means in practice | What breaks without it |
|---|---|---|
| Declarative | You describe the end state ("model v3.2.0, 4 replicas, 2 GB memory"), not the steps to reach it | Scripts that only work from a specific starting state; half-applied deploys |
| Versioned and immutable | Every desired state is a commit. History is append-only | The 02:40 question: nobody can say what was live yesterday |
| Pulled automatically | An agent inside the cluster watches the repo and applies changes | Credentials for production handed to CI; drift when someone applies by hand |
| Continuously reconciled | The agent re-checks constantly and corrects any manual change | An "emergency fix" typed into production survives for months, undocumented |
Reconciliation is the principle people underrate. If an engineer runs kubectl scale deployment fraud-api --replicas=12 during an incident, a reconciling agent notices within a minute that reality disagrees with the repository and scales back to 4. That is annoying exactly once, and then people learn to change the file — which means the change is reviewed, attributed and revertible.
Rollback in a GitOps system is
git revert. That is the whole point: the recovery procedure is the same procedure you use every day, so it works under stress.
Why this matters more for ML than for ordinary services
A web service has two moving parts: code and configuration. A machine learning service has five.
| Moving part | Ordinary service | ML service | Consequence |
|---|---|---|---|
| Application code | Yes | Yes | Same as usual |
| Configuration | Yes | Yes | Same as usual |
| Model weights | — | Large binary, changes on its own schedule | Too big for Git; needs a pointer + object store |
| Training data | — | Changes continuously, upstream, without warning | Same code produces a different model tomorrow |
| Decision thresholds | — | Tuned per business objective, changed by non-engineers | Behaviour changes with no code change at all |
That third row is the practical snag. A 400 MB model file must not go into Git — repositories become unusable, clones take minutes, and Git's delta compression is useless on binary blobs. The standard resolution is to put the bytes in object storage and the reference in Git:
deploy/production/fraud-api.yaml---model: uri: s3://ml-models/fraud/v3.2.0/model.joblib sha256: 9c1f0b7e5a4d3c28ff61aa07b93e5d18c4a2f7b0e9d6c3a1 registry_version: 17image: 123456789.dkr.ecr.eu-west-1.amazonaws.com/fraud-api@sha256:2f4c9d...bd91replicas: 4threshold: 0.63Everything that determines behaviour is now text in a reviewed file. Note both hashes: the image is pinned by digest, not by a tag like :latest or even :v3.2.0, because tags can be repointed at new bytes while a digest cannot.
A repository layout that supports this
fraud-ml/├── src/fraud/ # library code: features, train, predict├── configs/ # training configs (hyperparameters, data windows)├── docker/│ ├── train.Dockerfile│ └── serve.Dockerfile├── deploy/│ ├── base/ # shared manifests│ ├── staging/ # overrides: 1 replica, small instance│ └── production/ # overrides: 4 replicas, model URI, threshold├── .github/workflows/│ ├── ci.yaml # lint, test, parity check│ └── build-push.yaml # build image, push, open PR to bump digest└── tests/The separation that matters is src/ versus deploy/. A change to src/ requires a rebuild and a new image; a change to deploy/production/ takes effect on the next reconciliation with no build at all. Dropping the fraud threshold from 0.63 to 0.58 during an incident is a one-line, reviewable, revertible commit — not an SSH session.
Docker: making the environment an artefact
A container image is a filesystem plus a start command, built by executing a Dockerfile. Each instruction produces a layer, and layers are cached: if nothing an instruction depends on has changed, Docker reuses the previous result.
That caching rule dictates the single most important ordering decision in any ML Dockerfile.
The ordering that saves you six minutes per build
# WRONG — every code change reinstalls every dependencyFROM python:3.11-slimWORKDIR /appCOPY . .RUN pip install -r requirements.txtCMD ["python", "-m", "fraud.train"]# RIGHT — dependency layer is cached until requirements.txt changesFROM python:3.11-slim@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d5b6f01c2a9e4d7b3WORKDIR /appRUN apt-get update && apt-get install -y --no-install-recommends \ build-essential libgomp1 \ && rm -rf /var/lib/apt/lists/*COPY requirements.txt .RUN pip install --no-cache-dir --require-hashes -r requirements.txtCOPY src/ ./src/COPY configs/ ./configs/ENV PYTHONUNBUFFERED=1 PYTHONDONTWRITEBYTECODE=1CMD ["python", "-m", "fraud.train", "--config", "configs/train_xgb.yaml"]With the wrong ordering, changing one line in train.py invalidates the COPY . . layer and therefore the pip install layer beneath it. Installing pandas, scikit-learn, xgboost and their dependencies takes roughly 6 minutes. With the right ordering the same change rebuilds only the two small COPY layers — about 8 seconds. Over 20 iterations in an afternoon that is the difference between 2 hours of waiting and 3 minutes.
Three other details in that file earn their place. --no-install-recommends plus deleting the apt lists keeps a few hundred megabytes of documentation and suggested packages out. --require-hashes refuses to install anything not pinned by hash. PYTHONUNBUFFERED=1 makes log lines appear immediately instead of sitting in a buffer, which matters enormously when you are watching a training job's output in a log stream.
Serving needs a different image from training
Training needs compilers, plotting libraries, a full pandas stack and possibly CUDA development tools. Serving needs almost none of that. Building one image for both means shipping the training toolchain to every serving replica.
A multi-stage build compiles in one stage and copies only the results into a clean final stage:
FROM python:3.11-slim@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d5b6f01c2a9e4d7b3 AS builderWORKDIR /buildRUN apt-get update && apt-get install -y --no-install-recommends build-essential \ && rm -rf /var/lib/apt/lists/*COPY requirements-serve.txt .RUN pip install --no-cache-dir --prefix=/install -r requirements-serve.txtFROM python:3.11-slim@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d5b6f01c2a9e4d7b3WORKDIR /appCOPY --from=builder /install /usr/localCOPY src/fraud/ ./fraud/COPY --chown=nobody:nogroup models/model.joblib ./models/RUN useradd --create-home --shell /bin/false appuserUSER appuserENV PYTHONUNBUFFERED=1 MODEL_PATH=/app/models/model.joblibEXPOSE 8000HEALTHCHECK --interval=30s --timeout=3s --start-period=40s \ CMD python -c "import urllib.request,sys; \ sys.exit(0 if urllib.request.urlopen('http://localhost:8000/health').status==200 else 1)"CMD ["gunicorn", "-w", "2", "-k", "uvicorn.workers.UvicornWorker", \ "-b", "0.0.0.0:8000", "fraud.api:app"]Real numbers from a typical scikit-learn service: the single-stage image with build tools comes to about 1.24 GB; the multi-stage one lands at about 480 MB. At a 200 Mbit/s effective pull rate that is 9,920 Mbit ÷ 200 = 49.6 seconds versus 3,840 Mbit ÷ 200 = 19.2 seconds. Thirty seconds sounds trivial until a traffic spike causes an autoscaler to start 30 new pods on nodes with a cold cache; then it is 30 seconds of the incident spent waiting for bytes.
The USER appuser line is not decoration. A container running as root that suffers a remote code execution gives the attacker root inside the container, which is a much shorter hop to the host than an unprivileged account.
.dockerignore: the file everyone forgets
Before building, Docker sends the entire build context to the daemon. A repository with a data/ directory, a .git history and a few mlruns/ experiment folders can easily send 3 GB for a build that needs 4 MB of source.
.git.venv__pycache__/*.pycdata/mlruns/notebooks/*.ipynb.pytest_cache/*.csv*.parquet!models/model.joblibThe final line is an un-ignore: everything else large is excluded, but the one model file you genuinely need is let back in. Without a .dockerignore, a COPY . . can also silently bake your .env file, complete with credentials, into a layer that persists even if a later instruction deletes it.
Multi-container setups for local development
Training rarely runs alone — it wants a tracking server, a database, an object store. Compose describes the whole set so a new joiner runs one command:
services: mlflow: image: ghcr.io/mlflow/mlflow:v3.16.1-full # -full includes psycopg2 and boto3 command: > mlflow server --host 0.0.0.0 --port 5000 --backend-store-uri postgresql://mlflow:mlflow@db:5432/mlflow --artifacts-destination s3://mlflow-artifacts environment: MLFLOW_S3_ENDPOINT_URL: http://minio:9000 # S3 calls go to MinIO, not AWS AWS_ACCESS_KEY_ID: minioadmin # MinIO's default local credentials AWS_SECRET_ACCESS_KEY: minioadmin ports: ["5000:5000"] depends_on: [db, minio] db: image: postgres:16 environment: POSTGRES_USER: mlflow POSTGRES_PASSWORD: mlflow volumes: ["pgdata:/var/lib/postgresql/data"] minio: image: quay.io/minio/minio command: server /data --console-address ":9001" ports: ["9000:9000", "9001:9001"] trainer: build: context: . dockerfile: docker/train.Dockerfile environment: MLFLOW_TRACKING_URI: http://mlflow:5000 volumes: ["./data:/app/data:ro"] depends_on: [mlflow]volumes: pgdata:Two things worth noticing: services address each other by service name (http://mlflow:5000), not localhost, and the data volume is mounted read-only with :ro so a buggy training script cannot corrupt the source data. Two practical details also bite first-time users: the plain ghcr.io/mlflow/mlflow image ships without the Postgres and S3 client libraries, which is why the -full tag is used here, and the mlflow-artifacts bucket has to be created once in the MinIO console (port 9001) before the first run can store anything.
Model versioning: giving a model a name you can trust
A model needs an identity that survives being copied, that tells a human something useful, and that can be verified. Three schemes are in common use and they solve different problems.
| Scheme | Example | Communicates | Strength | Weakness |
|---|---|---|---|---|
| Semantic | v3.2.0 | Compatibility: major = breaking input/output change, minor = retrain with new features, patch = same model, fixed bug | Consumers instantly know if an upgrade is safe | Assigned by a human, so it can lie |
| Date-based | 2024-04-15.02 | Recency — which data window this saw | Perfect for scheduled retraining; sorts naturally | Says nothing about compatibility or quality |
| Content hash | a91f3c2b04 | Exact identity of the bytes | Cannot be faked or duplicated; deduplicates automatically | Meaningless to humans; unordered |
The mistake is choosing one. Mature systems use all three at once, because they answer different questions: is this safe to upgrade to?, how stale is it?, and is this the exact artefact that passed the tests?
1import hashlib, json, re2from dataclasses import dataclass, asdict3from datetime import datetime, timezone4from pathlib import Path56@dataclass(frozen=True)7class ModelVersion:8 name: str9 semver: str # "3.2.0"10 trained_at: str # ISO-8601 UTC11 content_sha: str # first 12 hex chars of the file's SHA-25612 git_sha: str13 metrics: dict1415 @property16 def uri_path(self) -> str:17 return f"{self.name}/v{self.semver}/{self.content_sha}"1819 def bump(self, level: str, **changes) -> "ModelVersion":20 major, minor, patch = (int(p) for p in self.semver.split("."))21 if level == "major":22 major, minor, patch = major + 1, 0, 023 elif level == "minor":24 minor, patch = minor + 1, 025 elif level == "patch":26 patch += 127 else:28 raise ValueError(f"unknown bump level: {level}")29 return ModelVersion(**{**asdict(self),30 "semver": f"{major}.{minor}.{patch}",31 **changes})3233def sha256_file(path: Path, chunk: int = 1 << 20) -> str:34 h = hashlib.sha256()35 with path.open("rb") as f:36 for block in iter(lambda: f.read(chunk), b""):37 h.update(block)38 return h.hexdigest()3940def register(path: Path, name: str, semver: str,41 git_sha: str, metrics: dict) -> ModelVersion:42 if not re.fullmatch(r"\d+\.\d+\.\d+", semver):43 raise ValueError(f"{semver!r} is not semantic version")44 mv = ModelVersion(45 name=name,46 semver=semver,47 trained_at=datetime.now(timezone.utc).isoformat(timespec="seconds"),48 content_sha=sha256_file(path)[:12],49 git_sha=git_sha,50 metrics=metrics,51 )52 (path.parent / "version.json").write_text(json.dumps(asdict(mv), indent=2))53 return mvThe rule that makes semantic versioning useful for models — and that teams get wrong — is what counts as a major bump. It is not "the model got much better". It is any change that breaks a consumer: a new required input feature, a changed output shape, a switch from returning a probability to returning a class label, or a recalibration that shifts the meaning of the score so existing thresholds are wrong. That last one is the sneaky case. Going from a model whose scores average 0.12 to one averaging 0.31 does not change the API signature at all, yet every downstream threshold is now wrong. It is a breaking change, and it deserves a major bump.
Version numbers exist to answer one question for someone who did not train the model: can I upgrade without changing my code? Anything that makes the answer "no" is a major version.
Wiring it together with CI/CD
The pipeline should build one image, test that image, push it, and then propose a change to the deployment manifest. It should never push directly to production.
name: build-and-proposeon: push: branches: [main]jobs: test: runs-on: ubuntu-latest steps: - uses: actions/checkout@v5 - uses: actions/setup-python@v6 with: python-version: "3.11" cache: pip - run: pip install --require-hashes -r requirements.txt - run: ruff check src tests - run: pytest tests -q --cov=src --cov-fail-under=80 - name: Train/serve parity run: pytest tests/test_train_serve_parity.py -q build: needs: test runs-on: ubuntu-latest permissions: contents: write packages: write pull-requests: write steps: - uses: actions/checkout@v5 - uses: docker/setup-buildx-action@v4 - uses: docker/login-action@v4 with: registry: ghcr.io username: ${{ github.actor }} password: ${{ secrets.GITHUB_TOKEN }} - uses: docker/build-push-action@v7 id: push with: context: . file: docker/serve.Dockerfile push: true tags: ghcr.io/${{ github.repository }}/fraud-api:${{ github.sha }} cache-from: type=gha cache-to: type=gha,mode=max - name: Smoke test the built image run: | docker run -d --name svc -p 8000:8000 \ ghcr.io/${{ github.repository }}/fraud-api@${{ steps.push.outputs.digest }} for i in $(seq 30); do curl -sf http://localhost:8000/health && break || sleep 2 done curl -sf -X POST http://localhost:8000/predict \ -H 'Content-Type: application/json' \ -d '{"tenure":14,"monthly_charges":79.2,"contract":"month-to-month"}' - name: Open PR bumping the staging digest env: GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} run: | git switch -c "bump-${{ github.sha }}" python scripts/set_image_digest.py \ deploy/staging/fraud-api.yaml "${{ steps.push.outputs.digest }}" git -c user.name=ci-bot -c user.email=ci-bot@users.noreply.github.com \ commit -am "chore: staging image ${{ github.sha }}" git push origin "bump-${{ github.sha }}" gh pr create --base main --head "bump-${{ github.sha }}" \ --title "chore: staging image ${{ github.sha }}" --fillTwo design choices here matter more than the YAML.
First, the smoke test runs against the digest that was just pushed, not against the tag. Between the push and the test, a concurrent build could repoint the tag; testing the digest tests exactly the bytes that will be deployed.
Second, the job's final act is to open a pull request, not to deploy. The deployment happens when a human merges that PR and the in-cluster agent reconciles. This keeps production credentials out of CI entirely — CI can write to a container registry and to the repository, and has no access to the cluster at all. If the CI system is compromised, the attacker still has to get a malicious commit through review.
The habits this buys you
Return to 02:40. In a system built this way, the on-call engineer runs git log deploy/production/, sees that the previous commit pointed at model.uri: s3://ml-models/fraud/v3.1.4/ and image digest sha256:8b2e..., and runs git revert. The agent reconciles within a minute. Total time: about four minutes, with a written record of exactly what changed and why.
Getting there does not require adopting everything at once. If you do only three things, do these:
- Never deploy an artefact a human built locally. If it was not built by CI from a commit, it does not go to production. This one rule eliminates
model_v2_final_FIXED.pklpermanently. - Pin by digest, everywhere. Base images, deployed images, model files. A tag is a mutable pointer; treating it as an identity is how "the same" deploy produces different behaviour.
- Make the desired state of production a reviewed file. Not a wiki page, not a runbook step, not tribal knowledge — a file whose history answers "what was live on Tuesday" without anyone having to remember.
Everything else — the multi-stage builds, the version manager, the Compose file — is optimisation on top of those three. They save minutes. The three above save four-hour outages.