MLOps for AI

GitOps, Docker, and Model Versioning


It is 02:40 and the fraud model is rejecting 38% of legitimate transactions. The on-call engineer needs to roll back. She asks the obvious question in the incident channel: what was live before?

Nobody knows. The model file on the production box is called model_v2_final_FIXED.pkl, modified eleven days ago. There is a model_v2_final.pkl next to it, and a model_v2.pkl, and a model_backup.pkl. None of them has a matching entry anywhere in version control. The deploy was done by someone who is asleep, by copying a file over SSH. The container running the service was built from a Dockerfile that says pip install scikit-learn, so rebuilding it today would not produce the same image anyway.

The rollback takes four hours and is done by guessing. That outage is not a modelling problem or a Kubernetes problem. It is a problem of state that exists nowhere except on a machine — and the entire discipline covered here exists to eliminate that.

Three ideas do the work. GitOps makes the desired state of production a file in a repository. Docker makes the environment a reproducible artefact rather than a machine's accumulated history. Model versioning gives every trained model an identity you can name in a rollback command at 02:40.

The only path a change is allowed to takeMerge codeand configCI builds atagged imagePush image,register modelBump the versionin the manifestReconcilerapplies itto the clusterAt 02:40 the answer to what was live before is a git log, not somebody's recollection.
If the cluster only ever changes through a merged commit, every rollback is a revert you can do half asleep.

GitOps: the repository is the source of truth

The core claim of GitOps is deceptively simple: the desired state of your system is described declaratively in Git, and an automated process continuously makes reality match it. Nobody changes production by typing commands at it.

The four principles

PrincipleWhat it means in practiceWhat breaks without it
DeclarativeYou describe the end state ("model v3.2.0, 4 replicas, 2 GB memory"), not the steps to reach itScripts that only work from a specific starting state; half-applied deploys
Versioned and immutableEvery desired state is a commit. History is append-onlyThe 02:40 question: nobody can say what was live yesterday
Pulled automaticallyAn agent inside the cluster watches the repo and applies changesCredentials for production handed to CI; drift when someone applies by hand
Continuously reconciledThe agent re-checks constantly and corrects any manual changeAn "emergency fix" typed into production survives for months, undocumented

Reconciliation is the principle people underrate. If an engineer runs kubectl scale deployment fraud-api --replicas=12 during an incident, a reconciling agent notices within a minute that reality disagrees with the repository and scales back to 4. That is annoying exactly once, and then people learn to change the file — which means the change is reviewed, attributed and revertible.

Rollback in a GitOps system is git revert. That is the whole point: the recovery procedure is the same procedure you use every day, so it works under stress.

Why this matters more for ML than for ordinary services

A web service has two moving parts: code and configuration. A machine learning service has five.

Moving partOrdinary serviceML serviceConsequence
Application codeYesYesSame as usual
ConfigurationYesYesSame as usual
Model weights—Large binary, changes on its own scheduleToo big for Git; needs a pointer + object store
Training data—Changes continuously, upstream, without warningSame code produces a different model tomorrow
Decision thresholds—Tuned per business objective, changed by non-engineersBehaviour changes with no code change at all

That third row is the practical snag. A 400 MB model file must not go into Git — repositories become unusable, clones take minutes, and Git's delta compression is useless on binary blobs. The standard resolution is to put the bytes in object storage and the reference in Git:

Text
deploy/production/fraud-api.yaml---model:  uri: s3://ml-models/fraud/v3.2.0/model.joblib  sha256: 9c1f0b7e5a4d3c28ff61aa07b93e5d18c4a2f7b0e9d6c3a1  registry_version: 17image: 123456789.dkr.ecr.eu-west-1.amazonaws.com/fraud-api@sha256:2f4c9d...bd91replicas: 4threshold: 0.63

Everything that determines behaviour is now text in a reviewed file. Note both hashes: the image is pinned by digest, not by a tag like :latest or even :v3.2.0, because tags can be repointed at new bytes while a digest cannot.

A repository layout that supports this

Text
fraud-ml/├── src/fraud/              # library code: features, train, predict├── configs/                # training configs (hyperparameters, data windows)├── docker/│   ├── train.Dockerfile│   └── serve.Dockerfile├── deploy/│   ├── base/               # shared manifests│   ├── staging/            # overrides: 1 replica, small instance│   └── production/         # overrides: 4 replicas, model URI, threshold├── .github/workflows/│   ├── ci.yaml             # lint, test, parity check│   └── build-push.yaml     # build image, push, open PR to bump digest└── tests/

The separation that matters is src/ versus deploy/. A change to src/ requires a rebuild and a new image; a change to deploy/production/ takes effect on the next reconciliation with no build at all. Dropping the fraud threshold from 0.63 to 0.58 during an incident is a one-line, reviewable, revertible commit — not an SSH session.

Docker: making the environment an artefact

A container image is a filesystem plus a start command, built by executing a Dockerfile. Each instruction produces a layer, and layers are cached: if nothing an instruction depends on has changed, Docker reuses the previous result.

That caching rule dictates the single most important ordering decision in any ML Dockerfile.

The ordering that saves you six minutes per build

Text
# WRONG — every code change reinstalls every dependencyFROM python:3.11-slimWORKDIR /appCOPY . .RUN pip install -r requirements.txtCMD ["python", "-m", "fraud.train"]
Text
# RIGHT — dependency layer is cached until requirements.txt changesFROM python:3.11-slim@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d5b6f01c2a9e4d7b3WORKDIR /appRUN apt-get update && apt-get install -y --no-install-recommends \        build-essential libgomp1 \    && rm -rf /var/lib/apt/lists/*COPY requirements.txt .RUN pip install --no-cache-dir --require-hashes -r requirements.txtCOPY src/ ./src/COPY configs/ ./configs/ENV PYTHONUNBUFFERED=1 PYTHONDONTWRITEBYTECODE=1CMD ["python", "-m", "fraud.train", "--config", "configs/train_xgb.yaml"]

With the wrong ordering, changing one line in train.py invalidates the COPY . . layer and therefore the pip install layer beneath it. Installing pandas, scikit-learn, xgboost and their dependencies takes roughly 6 minutes. With the right ordering the same change rebuilds only the two small COPY layers — about 8 seconds. Over 20 iterations in an afternoon that is the difference between 2 hours of waiting and 3 minutes.

Three other details in that file earn their place. --no-install-recommends plus deleting the apt lists keeps a few hundred megabytes of documentation and suggested packages out. --require-hashes refuses to install anything not pinned by hash. PYTHONUNBUFFERED=1 makes log lines appear immediately instead of sitting in a buffer, which matters enormously when you are watching a training job's output in a log stream.

Serving needs a different image from training

Training needs compilers, plotting libraries, a full pandas stack and possibly CUDA development tools. Serving needs almost none of that. Building one image for both means shipping the training toolchain to every serving replica.

A multi-stage build compiles in one stage and copies only the results into a clean final stage:

Text
FROM python:3.11-slim@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d5b6f01c2a9e4d7b3 AS builderWORKDIR /buildRUN apt-get update && apt-get install -y --no-install-recommends build-essential \    && rm -rf /var/lib/apt/lists/*COPY requirements-serve.txt .RUN pip install --no-cache-dir --prefix=/install -r requirements-serve.txtFROM python:3.11-slim@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d5b6f01c2a9e4d7b3WORKDIR /appCOPY --from=builder /install /usr/localCOPY src/fraud/ ./fraud/COPY --chown=nobody:nogroup models/model.joblib ./models/RUN useradd --create-home --shell /bin/false appuserUSER appuserENV PYTHONUNBUFFERED=1 MODEL_PATH=/app/models/model.joblibEXPOSE 8000HEALTHCHECK --interval=30s --timeout=3s --start-period=40s \  CMD python -c "import urllib.request,sys; \      sys.exit(0 if urllib.request.urlopen('http://localhost:8000/health').status==200 else 1)"CMD ["gunicorn", "-w", "2", "-k", "uvicorn.workers.UvicornWorker", \     "-b", "0.0.0.0:8000", "fraud.api:app"]

Real numbers from a typical scikit-learn service: the single-stage image with build tools comes to about 1.24 GB; the multi-stage one lands at about 480 MB. At a 200 Mbit/s effective pull rate that is 9,920 Mbit ÷ 200 = 49.6 seconds versus 3,840 Mbit ÷ 200 = 19.2 seconds. Thirty seconds sounds trivial until a traffic spike causes an autoscaler to start 30 new pods on nodes with a cold cache; then it is 30 seconds of the incident spent waiting for bytes.

The USER appuser line is not decoration. A container running as root that suffers a remote code execution gives the attacker root inside the container, which is a much shorter hop to the host than an unprivileged account.

.dockerignore: the file everyone forgets

Before building, Docker sends the entire build context to the daemon. A repository with a data/ directory, a .git history and a few mlruns/ experiment folders can easily send 3 GB for a build that needs 4 MB of source.

Text
.git.venv__pycache__/*.pycdata/mlruns/notebooks/*.ipynb.pytest_cache/*.csv*.parquet!models/model.joblib

The final line is an un-ignore: everything else large is excluded, but the one model file you genuinely need is let back in. Without a .dockerignore, a COPY . . can also silently bake your .env file, complete with credentials, into a layer that persists even if a later instruction deletes it.

Multi-container setups for local development

Training rarely runs alone — it wants a tracking server, a database, an object store. Compose describes the whole set so a new joiner runs one command:

Text
services:  mlflow:    image: ghcr.io/mlflow/mlflow:v3.16.1-full   # -full includes psycopg2 and boto3    command: >      mlflow server --host 0.0.0.0 --port 5000      --backend-store-uri postgresql://mlflow:mlflow@db:5432/mlflow      --artifacts-destination s3://mlflow-artifacts    environment:      MLFLOW_S3_ENDPOINT_URL: http://minio:9000   # S3 calls go to MinIO, not AWS      AWS_ACCESS_KEY_ID: minioadmin               # MinIO's default local credentials      AWS_SECRET_ACCESS_KEY: minioadmin    ports: ["5000:5000"]    depends_on: [db, minio]  db:    image: postgres:16    environment:      POSTGRES_USER: mlflow      POSTGRES_PASSWORD: mlflow    volumes: ["pgdata:/var/lib/postgresql/data"]  minio:    image: quay.io/minio/minio    command: server /data --console-address ":9001"    ports: ["9000:9000", "9001:9001"]  trainer:    build:      context: .      dockerfile: docker/train.Dockerfile    environment:      MLFLOW_TRACKING_URI: http://mlflow:5000    volumes: ["./data:/app/data:ro"]    depends_on: [mlflow]volumes:  pgdata:

Two things worth noticing: services address each other by service name (http://mlflow:5000), not localhost, and the data volume is mounted read-only with :ro so a buggy training script cannot corrupt the source data. Two practical details also bite first-time users: the plain ghcr.io/mlflow/mlflow image ships without the Postgres and S3 client libraries, which is why the -full tag is used here, and the mlflow-artifacts bucket has to be created once in the MinIO console (port 9001) before the first run can store anything.

Model versioning: giving a model a name you can trust

A model needs an identity that survives being copied, that tells a human something useful, and that can be verified. Three schemes are in common use and they solve different problems.

SchemeExampleCommunicatesStrengthWeakness
Semanticv3.2.0Compatibility: major = breaking input/output change, minor = retrain with new features, patch = same model, fixed bugConsumers instantly know if an upgrade is safeAssigned by a human, so it can lie
Date-based2024-04-15.02Recency — which data window this sawPerfect for scheduled retraining; sorts naturallySays nothing about compatibility or quality
Content hasha91f3c2b04Exact identity of the bytesCannot be faked or duplicated; deduplicates automaticallyMeaningless to humans; unordered

The mistake is choosing one. Mature systems use all three at once, because they answer different questions: is this safe to upgrade to?, how stale is it?, and is this the exact artefact that passed the tests?

Python
import hashlib, json, refrom dataclasses import dataclass, asdictfrom datetime import datetime, timezonefrom pathlib import Path@dataclass(frozen=True)class ModelVersion:    name: str    semver: str          # "3.2.0"    trained_at: str      # ISO-8601 UTC    content_sha: str     # first 12 hex chars of the file's SHA-256    git_sha: str    metrics: dict    @property    def uri_path(self) -> str:        return f"{self.name}/v{self.semver}/{self.content_sha}"    def bump(self, level: str, **changes) -> "ModelVersion":        major, minor, patch = (int(p) for p in self.semver.split("."))        if level == "major":            major, minor, patch = major + 1, 0, 0        elif level == "minor":            minor, patch = minor + 1, 0        elif level == "patch":            patch += 1        else:            raise ValueError(f"unknown bump level: {level}")        return ModelVersion(**{**asdict(self),                              "semver": f"{major}.{minor}.{patch}",                              **changes})def sha256_file(path: Path, chunk: int = 1 << 20) -> str:    h = hashlib.sha256()    with path.open("rb") as f:        for block in iter(lambda: f.read(chunk), b""):            h.update(block)    return h.hexdigest()def register(path: Path, name: str, semver: str,             git_sha: str, metrics: dict) -> ModelVersion:    if not re.fullmatch(r"\d+\.\d+\.\d+", semver):        raise ValueError(f"{semver!r} is not semantic version")    mv = ModelVersion(        name=name,        semver=semver,        trained_at=datetime.now(timezone.utc).isoformat(timespec="seconds"),        content_sha=sha256_file(path)[:12],        git_sha=git_sha,        metrics=metrics,    )    (path.parent / "version.json").write_text(json.dumps(asdict(mv), indent=2))    return mv

The rule that makes semantic versioning useful for models — and that teams get wrong — is what counts as a major bump. It is not "the model got much better". It is any change that breaks a consumer: a new required input feature, a changed output shape, a switch from returning a probability to returning a class label, or a recalibration that shifts the meaning of the score so existing thresholds are wrong. That last one is the sneaky case. Going from a model whose scores average 0.12 to one averaging 0.31 does not change the API signature at all, yet every downstream threshold is now wrong. It is a breaking change, and it deserves a major bump.

Version numbers exist to answer one question for someone who did not train the model: can I upgrade without changing my code? Anything that makes the answer "no" is a major version.

Wiring it together with CI/CD

The pipeline should build one image, test that image, push it, and then propose a change to the deployment manifest. It should never push directly to production.

Text
name: build-and-proposeon:  push:    branches: [main]jobs:  test:    runs-on: ubuntu-latest    steps:      - uses: actions/checkout@v5      - uses: actions/setup-python@v6        with:          python-version: "3.11"          cache: pip      - run: pip install --require-hashes -r requirements.txt      - run: ruff check src tests      - run: pytest tests -q --cov=src --cov-fail-under=80      - name: Train/serve parity        run: pytest tests/test_train_serve_parity.py -q  build:    needs: test    runs-on: ubuntu-latest    permissions:      contents: write      packages: write      pull-requests: write    steps:      - uses: actions/checkout@v5      - uses: docker/setup-buildx-action@v4      - uses: docker/login-action@v4        with:          registry: ghcr.io          username: ${{ github.actor }}          password: ${{ secrets.GITHUB_TOKEN }}      - uses: docker/build-push-action@v7        id: push        with:          context: .          file: docker/serve.Dockerfile          push: true          tags: ghcr.io/${{ github.repository }}/fraud-api:${{ github.sha }}          cache-from: type=gha          cache-to: type=gha,mode=max      - name: Smoke test the built image        run: |          docker run -d --name svc -p 8000:8000 \            ghcr.io/${{ github.repository }}/fraud-api@${{ steps.push.outputs.digest }}          for i in $(seq 30); do            curl -sf http://localhost:8000/health && break || sleep 2          done          curl -sf -X POST http://localhost:8000/predict \            -H 'Content-Type: application/json' \            -d '{"tenure":14,"monthly_charges":79.2,"contract":"month-to-month"}'      - name: Open PR bumping the staging digest        env:          GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}        run: |          git switch -c "bump-${{ github.sha }}"          python scripts/set_image_digest.py \            deploy/staging/fraud-api.yaml "${{ steps.push.outputs.digest }}"          git -c user.name=ci-bot -c user.email=ci-bot@users.noreply.github.com \            commit -am "chore: staging image ${{ github.sha }}"          git push origin "bump-${{ github.sha }}"          gh pr create --base main --head "bump-${{ github.sha }}" \            --title "chore: staging image ${{ github.sha }}" --fill

Two design choices here matter more than the YAML.

First, the smoke test runs against the digest that was just pushed, not against the tag. Between the push and the test, a concurrent build could repoint the tag; testing the digest tests exactly the bytes that will be deployed.

Second, the job's final act is to open a pull request, not to deploy. The deployment happens when a human merges that PR and the in-cluster agent reconciles. This keeps production credentials out of CI entirely — CI can write to a container registry and to the repository, and has no access to the cluster at all. If the CI system is compromised, the attacker still has to get a malicious commit through review.

The habits this buys you

Return to 02:40. In a system built this way, the on-call engineer runs git log deploy/production/, sees that the previous commit pointed at model.uri: s3://ml-models/fraud/v3.1.4/ and image digest sha256:8b2e..., and runs git revert. The agent reconciles within a minute. Total time: about four minutes, with a written record of exactly what changed and why.

Getting there does not require adopting everything at once. If you do only three things, do these:

  • Never deploy an artefact a human built locally. If it was not built by CI from a commit, it does not go to production. This one rule eliminates model_v2_final_FIXED.pkl permanently.
  • Pin by digest, everywhere. Base images, deployed images, model files. A tag is a mutable pointer; treating it as an identity is how "the same" deploy produces different behaviour.
  • Make the desired state of production a reviewed file. Not a wiki page, not a runbook step, not tribal knowledge — a file whose history answers "what was live on Tuesday" without anyone having to remember.

Everything else — the multi-stage builds, the version manager, the Compose file — is optimisation on top of those three. They save minutes. The three above save four-hour outages.