Model Deployment for AI Engineers

CI/CD for Model Pipelines


A team has a solid CI pipeline. Every pull request runs 340 unit tests, a linter, and a type checker. All green. On Tuesday they merge a change that swaps the model artefact from resnet18_v4.onnx to resnet50_v1.onnx — a one-line diff in a config file. Every test passes, because the tests mock the model.

In production, p95 latency goes from 82 ms to 340 ms. The image grew from 98 MB to 312 MB, so autoscaling events now take four minutes instead of one. And the new model's label ordering differs from the old one, so class 2 and class 3 are silently swapped — accuracy on those two classes falls to near zero while overall accuracy only drops a few points, which is small enough that nobody notices for three days.

Every one of those three failures is mechanically detectable. None of them is detectable by a test suite designed for code, because the artefact that changed was not code. That gap is what model CI/CD exists to close.

From commit to production, one gate at a timeLint andunit testsValidate theartefact itselfBuild andpush the imageIntegrationtest thecontainerDeploy tostagingCanary toproduction340 green tests missed a swapped .onnx file because every one of them mocked the model.
The stage that catches a model change is the one that loads the real artefact and scores a fixed eval set.

What is different about shipping a model

Code CI/CDModel CI/CD
What changesSource files in gitSource, plus a binary artefact, plus training data
CorrectnessDeterministic: the test passes or failsStatistical: accuracy 0.918 versus 0.921 — is that a failure?
Artefact sizeKilobytesTens of megabytes to tens of gigabytes
Non-functional regressionsRare and usually obviousRoutine: size, latency, and memory all move with the model
ReproducibilitySame input, same outputDepends on seeds, hardware, library versions
Failure visibilityExceptions and stack tracesOften silent: slightly worse predictions

The row that matters most is the last one. A code bug throws. A model regression returns a confident, well-formed, wrong answer with a 200 status code. Your existing monitoring will not catch it. The pipeline has to.

Code CI asks "does it run?"; model CI must also ask "does it still predict what it used to, at the speed and size it used to?"

The shape of the pipeline

Text
push / PR  │  ├─ 1. Code quality      lint, type-check, unit tests            ~2 min  ├─ 2. Model validation   loads? shapes? size? latency? accuracy? ~5 min  ├─ 3. Build              container image, tagged by git sha      ~4 min  ├─ 4. Integration        run the real container, hit the real API ~3 min  ├─ 5. Staging deploy     + smoke tests against staging           ~3 min  └─ 6. Production deploy  canary → full, with automated rollback  ~15 min

Stages 1, 3, and 5 exist in any pipeline. Stages 2 and 4 are the model-specific ones, and they are where the three Tuesday failures would have been caught.

GateCatchesCost of not having it
Export equivalenceA converted artefact that no longer matches the trained modelSilently wrong predictions with no error anywhere
Golden predictionsLabel reordering, preprocessing drift, wrong artefact in the imageDays of degraded accuracy before anyone notices
SizeAn unquantized model or a larger architecture shipped by accidentSlow image pulls; autoscaling stops keeping up with traffic
LatencyA model that is correct but too slowLatency budget breached in production, discovered by users
Accuracy, per classA regression concentrated in one or two classesAggregate metrics look fine while a segment is broken
IntegrationPackaging faults: missing libraries, wrong paths, stale weightsThe build is green and the container does not work

The pipeline, stage by stage

Stage 1 — code quality

YAML
name: model-pipelineon:  push: { branches: [main] }  pull_request: { branches: [main] }env:  REGISTRY: ghcr.io  IMAGE: ghcr.io/acme/defect-apijobs:  code-quality:    runs-on: ubuntu-latest    steps:      - uses: actions/checkout@v5        with: { lfs: true }              # model artefacts live in git-lfs      - uses: actions/setup-python@v6        with: { python-version: "3.11", cache: pip }      - run: pip install -r requirements.txt -r requirements-dev.txt      - run: ruff check app/ tests/      - run: mypy app/ --strict      - run: pytest tests/unit -q --cov=app --cov-fail-under=80

lfs: true is easy to forget and produces a baffling failure: without it, git-lfs files check out as 130-byte text pointers, and your model load fails with an unpickling error that suggests the artefact is corrupt rather than absent.

Stage 2 — validate the model itself

This is the stage that does not exist in ordinary CI, and it is the important one. Four gates, each catching a different class of regression.

YAML
  model-validation:    runs-on: ubuntu-latest    needs: code-quality    steps:      - uses: actions/checkout@v5        with: { lfs: true }      - uses: actions/setup-python@v6        with: { python-version: "3.11", cache: pip }      - run: pip install -r requirements.txt      - name: Model loads and exports match        run: python scripts/validate_artefact.py --model models/current/      - name: Size regression        run: python scripts/check_size.py --model models/current/model.onnx --max-growth 1.20      - name: Latency regression        run: python scripts/check_latency.py --model models/current/model.onnx --max-p95-ms 100      - name: Accuracy on the held-out set        run: python scripts/check_accuracy.py --model models/current/ --min-f1 0.915      - uses: actions/upload-artifact@v6        if: always()        with: { name: validation-report, path: reports/ }

Gate one: it loads, and the export matches the original

Python
# scripts/validate_artefact.pyimport numpy as np, onnx, onnxruntime as ort, torch, sys, jsondef main(model_dir):    onnx_path = f"{model_dir}/model.onnx"    onnx.checker.check_model(onnx.load(onnx_path))          # structural only    sess = ort.InferenceSession(onnx_path, providers=["CPUExecutionProvider"])    inp = sess.get_inputs()[0]    print("input:", inp.name, inp.shape, inp.type)    torch_model = torch.jit.load(f"{model_dir}/model.pt", map_location="cpu").eval()    rng = np.random.default_rng(0)    worst = 0.0    for batch in (1, 4, 16):                                # exercise the dynamic axis        x = rng.standard_normal((batch, 3, 224, 224), dtype=np.float32)        with torch.no_grad():            expected = torch_model(torch.from_numpy(x)).numpy()        actual = sess.run(None, {inp.name: x})[0]        worst = max(worst, float(np.abs(expected - actual).max()))    print(f"max abs diff across batch sizes: {worst:.3e}")    if worst > 1e-4:        sys.exit(f"FAIL: ONNX export diverges from PyTorch by {worst:.3e}")    # Golden predictions: twelve real samples with recorded expected labels.    golden = np.load("tests/golden.npz")    preds = sess.run(None, {inp.name: golden["inputs"]})[0].argmax(axis=1)    mismatches = int((preds != golden["labels"]).sum())    if mismatches:        sys.exit(f"FAIL: {mismatches}/{len(preds)} golden predictions changed")    print("golden set: all match")if __name__ == "__main__":    main(sys.argv[sys.argv.index("--model") + 1])

The golden set is what catches the swapped-label bug from the opening. Twelve real samples, their expected labels recorded when the pipeline was known good, checked on every build. If the label ordering changes, class 2 and class 3 flip in the golden set and the build fails immediately — instead of silently degrading for three days.

A dozen real inputs with recorded expected outputs is the highest-value test in a model pipeline, because it fails on exactly the changes that produce no error message.

Gate two: size

Python
# scripts/check_size.pyimport os, json, sys, argparsep = argparse.ArgumentParser()p.add_argument("--model"); p.add_argument("--max-growth", type=float, default=1.20)a = p.parse_args()new_mb = os.path.getsize(a.model) / 1e6baseline = json.load(open("baselines/model.json"))["size_mb"]ratio = new_mb / baselineprint(f"size {new_mb:.1f} MB vs baseline {baseline:.1f} MB  ({ratio:.2f}x)")if ratio > a.max_growth:    sys.exit(f"FAIL: model grew {ratio:.2f}x, limit is {a.max_growth:.2f}x")

On the Tuesday change this reports 312.0 MB vs baseline 98.4 MB (3.17x) and fails. A 3.17× jump is never accidental in a healthy pipeline — it means either the architecture changed or someone shipped an unquantized artefact by mistake.

Gate three: latency

Python
# scripts/check_latency.pyimport time, numpy as np, onnxruntime as ort, sys, json, argparsep = argparse.ArgumentParser()p.add_argument("--model"); p.add_argument("--max-p95-ms", type=float)a = p.parse_args()opts = ort.SessionOptions()opts.intra_op_num_threads = 1            # pin threads: CI runners vary wildlysess = ort.InferenceSession(a.model, opts, providers=["CPUExecutionProvider"])name = sess.get_inputs()[0].namex = np.random.randn(1, 3, 224, 224).astype(np.float32)for _ in range(20):                       # warm up: allocations, kernel selection    sess.run(None, {name: x})samples = []for _ in range(200):    t0 = time.perf_counter()    sess.run(None, {name: x})    samples.append((time.perf_counter() - t0) * 1000)p50, p95 = np.percentile(samples, [50, 95])print(f"p50 {p50:.1f} ms   p95 {p95:.1f} ms")json.dump({"p50_ms": p50, "p95_ms": p95}, open("reports/latency.json", "w"))if p95 > a.max_p95_ms:    sys.exit(f"FAIL: p95 {p95:.1f} ms exceeds budget {a.max_p95_ms:.1f} ms")

Three details make this measurement trustworthy rather than noise. Warm up first — the first several runs include memory allocation and kernel selection and can be five times slower than steady state. Pin thread count, because CI runners have between two and sixteen cores depending on the day and an unpinned run varies by 3× for reasons unrelated to your model. Compare p95, not the mean, and use an absolute budget rather than a ratio to the previous run — ratios on a noisy shared runner produce flaky failures, which train people to hit "re-run" without reading the output.

Gate four: accuracy

This one needs a rule for what counts as a regression, and the rule must account for the fact that a fixed held-out set gives you a noisy estimate. On a 2,000-sample test set, a measured F1 of 0.921 has a standard error of roughly 0.006. A drop to 0.918 is well within noise; a drop to 0.895 is not.

Python
baseline_f1, new_f1 = 0.9210, 0.9130tolerance = 0.0050print(f"F1 {new_f1:.4f} vs baseline {baseline_f1:.4f} (delta {new_f1-baseline_f1:+.4f})")if baseline_f1 - new_f1 > tolerance:    sys.exit("FAIL: accuracy regression beyond tolerance")# Per-class: an aggregate that holds can hide a class that collapsed.for cls, (base, now) in per_class.items():    if base - now > 0.03:        sys.exit(f"FAIL: class '{cls}' fell from {base:.3f} to {now:.3f}")

The per-class check is not optional. Swapping two of four class labels moves aggregate F1 by only a few points if those classes are rare, but drives their individual F1 to near zero. Aggregate metrics are exactly the wrong instrument for detecting that.

Stage 3 — build and push

YAML
  build:    runs-on: ubuntu-latest    needs: model-validation    permissions: { contents: read, packages: write }    outputs:      digest: ${{ steps.push.outputs.digest }}    steps:      - uses: actions/checkout@v5        with: { lfs: true }      - uses: docker/setup-buildx-action@v4      - uses: docker/login-action@v4        with:          registry: ghcr.io          username: ${{ github.actor }}          password: ${{ secrets.GITHUB_TOKEN }}      - id: push        uses: docker/build-push-action@v7        with:          context: .                     # use the checkout, so LFS files are real files          push: true          tags: |            ghcr.io/acme/defect-api:${{ github.sha }}            ghcr.io/acme/defect-api:latest          cache-from: type=gha          cache-to: type=gha,mode=max      - name: Scan for vulnerabilities        uses: aquasecurity/trivy-action@v0.36.0   # pin a release (better: its commit SHA), never @master        with:          image-ref: ghcr.io/acme/defect-api:${{ github.sha }}          severity: HIGH,CRITICAL          exit-code: "1"

Tagging by commit SHA rather than only latest is what makes rollback possible: every deployment points at an immutable digest, and rolling back is deploying a previous SHA. A pipeline that only produces latest has no rollback target at all.

The context: . line matters for this repository in particular. Without it, the build action fetches the source itself with plain Git, ignoring the checkout step, so the git-lfs model files arrive in the image as 130-byte pointer files. The GitHub Actions cache (type=gha) matters more for ML images than for ordinary ones. Without it, every build reinstalls PyTorch — around four minutes. With it and a correctly ordered Dockerfile, a code-only change rebuilds in under thirty seconds.

Stage 4 — integration test against the real container

Unit tests mock the model. This stage does not. It starts the image you just built and talks to it over HTTP, which is the only way to catch failures that live in the gap between your code and its packaging: a missing system library, a wrong working directory, a model path that exists on your laptop and not in the image.

YAML
  integration:    runs-on: ubuntu-latest    needs: build    steps:      - uses: actions/checkout@v5      - name: Start the container        run: |          docker run -d --name api -p 8000:8000 \            ghcr.io/acme/defect-api:${{ github.sha }}          for i in $(seq 1 60); do            curl -fsS http://localhost:8000/readyz && break            sleep 2          done      - run: pytest tests/integration -q --base-url http://localhost:8000      - if: failure()        run: docker logs api

The readiness loop rather than a fixed sleep 30 is the difference between a pipeline that is reliable and one that is flaky: model load time varies with runner load, and a fixed sleep is either too short (random failures) or too long (wasted minutes on every build). Dumping container logs on failure saves the round trip of re-running the job with debugging added.

Python
# tests/integration/test_api.pyimport requests, base64, pathlib, pytestdef test_health(base_url):    assert requests.get(f"{base_url}/readyz", timeout=5).json()["ready"] is Truedef test_known_image(base_url):    img = base64.b64encode(pathlib.Path("tests/fixtures/scratch_01.jpg")                           .read_bytes()).decode()    r = requests.post(f"{base_url}/predict", json={"image_b64": img}, timeout=30)    assert r.status_code == 200    body = r.json()    assert body["label"] == "scratch"    assert body["confidence"] > 0.80    assert body["model_version"] == "2.6.0"     # the right artefact is insidedef test_rejects_garbage(base_url):    r = requests.post(f"{base_url}/predict", json={"image_b64": "notbase64"}, timeout=10)    assert r.status_code == 422                  # not 500@pytest.mark.parametrize("n", [1, 8, 32])def test_batch_sizes(base_url, n):    r = requests.post(f"{base_url}/predict/batch",                      json={"items": [SAMPLE] * n}, timeout=60)    assert len(r.json()["predictions"]) == n

Asserting on model_version is a small thing that catches a real and embarrassing failure: the image built correctly but baked in a stale model because a path in the Dockerfile pointed at the wrong directory.

Stages 5 and 6 — staging, then production

YAML
  deploy-staging:    runs-on: ubuntu-latest    needs: integration    if: github.ref == 'refs/heads/main'    environment: staging    steps:      - run: |          kubectl set image deployment/defect-api \            api=ghcr.io/acme/defect-api:${{ github.sha }} -n staging          kubectl rollout status deployment/defect-api -n staging --timeout=300s      - run: python scripts/smoke.py --url https://staging.internal/predict  deploy-production:    runs-on: ubuntu-latest    needs: deploy-staging    environment: production          # requires a human approval in repo settings    steps:      - name: Canary at 10%        run: |          kubectl apply -f k8s/canary-10.yaml          kubectl set image deployment/defect-api-canary \            api=ghcr.io/acme/defect-api:${{ github.sha }} -n prod          kubectl rollout status deployment/defect-api-canary -n prod --timeout=300s      - name: Watch canary metrics for 10 minutes        run: python scripts/watch_canary.py --window 600 --error-budget 0.02      - name: Promote to 100%        id: promote        run: |          kubectl set image deployment/defect-api \            api=ghcr.io/acme/defect-api:${{ github.sha }} -n prod          kubectl rollout status deployment/defect-api -n prod --timeout=600s          kubectl scale deployment/defect-api-canary --replicas=0 -n prod

The environment: production line hooks into GitHub's protected environments, which can require a named reviewer to approve before the job runs. Full automation to production is a goal, not a starting point — earn it by first accumulating evidence that stages 2 and 4 catch what they claim to.

GitLab CI: the same ideas, different nouns

YAML
stages: [quality, validate, build, integration, deploy]variables:  IMAGE: $CI_REGISTRY_IMAGE:$CI_COMMIT_SHAcode-quality:  stage: quality  image: python:3.11-slim  script:    - pip install -r requirements.txt -r requirements-dev.txt    - ruff check app/ tests/    - pytest tests/unit -qmodel-validation:  stage: validate  image: python:3.11-slim  script:    - python scripts/validate_artefact.py --model models/current/    - python scripts/check_size.py --model models/current/model.onnx --max-growth 1.20    - python scripts/check_latency.py --model models/current/model.onnx --max-p95-ms 100  artifacts:    when: always    paths: [reports/]    expire_in: 30 daysbuild:  stage: build  image: docker:27  services: [docker:27-dind]  script:    - echo "$CI_REGISTRY_PASSWORD" | docker login -u "$CI_REGISTRY_USER" --password-stdin "$CI_REGISTRY"    - docker build -t "$IMAGE" .    - docker push "$IMAGE"deploy-production:  stage: deploy  when: manual                       # equivalent to a protected environment  environment: { name: production, url: https://api.acme.com }  script:    - kubectl set image deployment/defect-api api="$IMAGE" -n prod    - kubectl rollout status deployment/defect-api -n prod --timeout=600s

The mapping is direct: jobs become stages plus jobs, needs becomes stage ordering, upload-artifact becomes artifacts, and protected environments become when: manual with an environment block.

Automated rollback

Deploying is only half of it. The pipeline must also be able to undo itself without a human, because the window where a bad model is doing damage is the window where a human is still reading the alert.

Python
# scripts/watch_canary.pyimport time, sys, statistics, argparse, requestsdef sample(url):    m = requests.get(f"{url}/stats", timeout=5).json()    return m["error_rate"], m["p99_ms"], m["requests"]def main(window, error_budget):    canary, baseline = "http://canary.prod.internal", "http://stable.prod.internal"    deadline = time.time() + window    while time.time() < deadline:        c_err, c_p99, c_n = sample(canary)        b_err, b_p99, _   = sample(baseline)        print(f"canary n={c_n} err={c_err:.4f} p99={c_p99:.0f}ms | "              f"stable err={b_err:.4f} p99={b_p99:.0f}ms")        if c_n < 1000:                       # too little data to judge            time.sleep(30); continue        if c_err > max(b_err * 2, error_budget):            sys.exit(f"ROLLBACK: canary error rate {c_err:.4f}")        if c_p99 > b_p99 * 1.5:            sys.exit(f"ROLLBACK: canary p99 {c_p99:.0f}ms vs stable {b_p99:.0f}ms")        time.sleep(30)    print("canary healthy")if __name__ == "__main__":    p = argparse.ArgumentParser()    p.add_argument("--window", type=int); p.add_argument("--error-budget", type=float)    a = p.parse_args(); main(a.window, a.error_budget)
YAML
      - name: Pull the canary on failure        if: failure()        run: kubectl scale deployment/defect-api-canary --replicas=0 -n prod      - name: Undo the main rollout if promotion had started        if: failure() && steps.promote.conclusion == 'failure'        run: |          kubectl rollout undo deployment/defect-api -n prod          kubectl rollout status deployment/defect-api -n prod --timeout=300s      - name: Tell the team        if: failure()        run: |          curl -X POST "$SLACK_WEBHOOK" \            -d '{"text":"defect-api release failed and was rolled back"}'

These steps go at the end of the deploy-production job. The order of failure matters: if the canary watch fails, the main deployment has not been touched yet, so the fix is to take the canary away — running rollout undo on the main deployment then would roll it back past the version that was healthy. Only if the promotion step itself fails is the main deployment undone.

The c_n < 1000 guard is the same statistical point that governs canaries generally: at 100 requests with a 0.5% baseline error rate you expect 0.5 errors, so observing 2 is ordinary noise. Automation that rolls back on noise gets disabled by frustrated engineers within a fortnight, which leaves you worse off than having no automation.

A rollback that depends on a human noticing, deciding, and typing is not a safety mechanism; it is a hope that someone is awake.

Mistakes that make a model pipeline useless

Mocking the model in every test. Mocks make unit tests fast, which is right — but if nothing in the pipeline loads the real artefact, then a corrupted, mismatched, or missing model passes CI cleanly. At least one stage must execute the actual file.

Committing model binaries to git without LFS. A 300 MB file in git history is permanent, and every clone thereafter pays for it. Use git-lfs or, better, keep artefacts in object storage and commit only a version identifier plus a checksum.

Secrets in the workflow file. A registry password or cloud key committed to a public repository is scraped within minutes. Use the platform's secret store, and prefer short-lived credentials such as OIDC federation over long-lived keys.

Latency gates on unpinned runners. Covered above, and worth restating because the failure mode is social rather than technical: flaky gates get ignored, then bypassed, then removed.

No baseline file under version control. If the numbers you compare against live in someone's spreadsheet, the comparison is not reproducible and nobody can tell when the baseline moved. Keep baselines/model.json in the repository and update it in a reviewed pull request, so a change to the standard is as visible as a change to the code.

Where to start if you have none of this

Building all six stages at once is how these projects stall. Build them in the order of what has already gone wrong for you, and stop at the point where the next stage costs more than the failures it prevents.

If you have nothing, the highest-value single addition is the golden-prediction test: a dozen real inputs, their expected outputs recorded from a known-good model, checked on every build. It takes an hour to write, needs no infrastructure, and catches label reordering, preprocessing drift, corrupted artefacts, and wrong-model-in-the-image — four distinct failure modes with one test.

Second: the integration stage. Starting the real container and hitting the real endpoint catches everything that lives in the gap between code and packaging, which is where a surprising share of production incidents originate.

Third: size and latency gates, because they are cheap and they catch regressions that no accuracy metric will ever show. Only then invest in canary automation and automated rollback, which are the most work and matter only once deployments are frequent enough that a human cannot watch each one.

The measure of whether it is working is not the number of stages. It is whether you would be comfortable letting a colleague merge a model change on a Friday afternoon. If the honest answer is no, the pipeline is telling you which gate is still missing.