Course Content
Model Deployment for AI Engineers
4 sections · 10 lessons
CI/CD for Model Pipelines
A team has a solid CI pipeline. Every pull request runs 340 unit tests, a linter, and a type checker. All green. On Tuesday they merge a change that swaps the model artefact from resnet18_v4.onnx to resnet50_v1.onnx — a one-line diff in a config file. Every test passes, because the tests mock the model.
In production, p95 latency goes from 82 ms to 340 ms. The image grew from 98 MB to 312 MB, so autoscaling events now take four minutes instead of one. And the new model's label ordering differs from the old one, so class 2 and class 3 are silently swapped — accuracy on those two classes falls to near zero while overall accuracy only drops a few points, which is small enough that nobody notices for three days.
Every one of those three failures is mechanically detectable. None of them is detectable by a test suite designed for code, because the artefact that changed was not code. That gap is what model CI/CD exists to close.
What is different about shipping a model
| Code CI/CD | Model CI/CD | |
|---|---|---|
| What changes | Source files in git | Source, plus a binary artefact, plus training data |
| Correctness | Deterministic: the test passes or fails | Statistical: accuracy 0.918 versus 0.921 — is that a failure? |
| Artefact size | Kilobytes | Tens of megabytes to tens of gigabytes |
| Non-functional regressions | Rare and usually obvious | Routine: size, latency, and memory all move with the model |
| Reproducibility | Same input, same output | Depends on seeds, hardware, library versions |
| Failure visibility | Exceptions and stack traces | Often silent: slightly worse predictions |
The row that matters most is the last one. A code bug throws. A model regression returns a confident, well-formed, wrong answer with a 200 status code. Your existing monitoring will not catch it. The pipeline has to.
Code CI asks "does it run?"; model CI must also ask "does it still predict what it used to, at the speed and size it used to?"
The shape of the pipeline
push / PR │ ├─ 1. Code quality lint, type-check, unit tests ~2 min ├─ 2. Model validation loads? shapes? size? latency? accuracy? ~5 min ├─ 3. Build container image, tagged by git sha ~4 min ├─ 4. Integration run the real container, hit the real API ~3 min ├─ 5. Staging deploy + smoke tests against staging ~3 min └─ 6. Production deploy canary → full, with automated rollback ~15 minStages 1, 3, and 5 exist in any pipeline. Stages 2 and 4 are the model-specific ones, and they are where the three Tuesday failures would have been caught.
| Gate | Catches | Cost of not having it |
|---|---|---|
| Export equivalence | A converted artefact that no longer matches the trained model | Silently wrong predictions with no error anywhere |
| Golden predictions | Label reordering, preprocessing drift, wrong artefact in the image | Days of degraded accuracy before anyone notices |
| Size | An unquantized model or a larger architecture shipped by accident | Slow image pulls; autoscaling stops keeping up with traffic |
| Latency | A model that is correct but too slow | Latency budget breached in production, discovered by users |
| Accuracy, per class | A regression concentrated in one or two classes | Aggregate metrics look fine while a segment is broken |
| Integration | Packaging faults: missing libraries, wrong paths, stale weights | The build is green and the container does not work |
The pipeline, stage by stage
Stage 1 — code quality
1name: model-pipeline2on:3 push: { branches: [main] }4 pull_request: { branches: [main] }56env:7 REGISTRY: ghcr.io8 IMAGE: ghcr.io/acme/defect-api910jobs:11 code-quality:12 runs-on: ubuntu-latest13 steps:14 - uses: actions/checkout@v515 with: { lfs: true } # model artefacts live in git-lfs16 - uses: actions/setup-python@v617 with: { python-version: "3.11", cache: pip }18 - run: pip install -r requirements.txt -r requirements-dev.txt19 - run: ruff check app/ tests/20 - run: mypy app/ --strict21 - run: pytest tests/unit -q --cov=app --cov-fail-under=80lfs: true is easy to forget and produces a baffling failure: without it, git-lfs files check out as 130-byte text pointers, and your model load fails with an unpickling error that suggests the artefact is corrupt rather than absent.
Stage 2 — validate the model itself
This is the stage that does not exist in ordinary CI, and it is the important one. Four gates, each catching a different class of regression.
1 model-validation:2 runs-on: ubuntu-latest3 needs: code-quality4 steps:5 - uses: actions/checkout@v56 with: { lfs: true }7 - uses: actions/setup-python@v68 with: { python-version: "3.11", cache: pip }9 - run: pip install -r requirements.txt1011 - name: Model loads and exports match12 run: python scripts/validate_artefact.py --model models/current/1314 - name: Size regression15 run: python scripts/check_size.py --model models/current/model.onnx --max-growth 1.201617 - name: Latency regression18 run: python scripts/check_latency.py --model models/current/model.onnx --max-p95-ms 1001920 - name: Accuracy on the held-out set21 run: python scripts/check_accuracy.py --model models/current/ --min-f1 0.9152223 - uses: actions/upload-artifact@v624 if: always()25 with: { name: validation-report, path: reports/ }Gate one: it loads, and the export matches the original
1# scripts/validate_artefact.py2import numpy as np, onnx, onnxruntime as ort, torch, sys, json34def main(model_dir):5 onnx_path = f"{model_dir}/model.onnx"6 onnx.checker.check_model(onnx.load(onnx_path)) # structural only78 sess = ort.InferenceSession(onnx_path, providers=["CPUExecutionProvider"])9 inp = sess.get_inputs()[0]10 print("input:", inp.name, inp.shape, inp.type)1112 torch_model = torch.jit.load(f"{model_dir}/model.pt", map_location="cpu").eval()1314 rng = np.random.default_rng(0)15 worst = 0.016 for batch in (1, 4, 16): # exercise the dynamic axis17 x = rng.standard_normal((batch, 3, 224, 224), dtype=np.float32)18 with torch.no_grad():19 expected = torch_model(torch.from_numpy(x)).numpy()20 actual = sess.run(None, {inp.name: x})[0]21 worst = max(worst, float(np.abs(expected - actual).max()))2223 print(f"max abs diff across batch sizes: {worst:.3e}")24 if worst > 1e-4:25 sys.exit(f"FAIL: ONNX export diverges from PyTorch by {worst:.3e}")2627 # Golden predictions: twelve real samples with recorded expected labels.28 golden = np.load("tests/golden.npz")29 preds = sess.run(None, {inp.name: golden["inputs"]})[0].argmax(axis=1)30 mismatches = int((preds != golden["labels"]).sum())31 if mismatches:32 sys.exit(f"FAIL: {mismatches}/{len(preds)} golden predictions changed")33 print("golden set: all match")3435if __name__ == "__main__":36 main(sys.argv[sys.argv.index("--model") + 1])The golden set is what catches the swapped-label bug from the opening. Twelve real samples, their expected labels recorded when the pipeline was known good, checked on every build. If the label ordering changes, class 2 and class 3 flip in the golden set and the build fails immediately — instead of silently degrading for three days.
A dozen real inputs with recorded expected outputs is the highest-value test in a model pipeline, because it fails on exactly the changes that produce no error message.
Gate two: size
1# scripts/check_size.py2import os, json, sys, argparse34p = argparse.ArgumentParser()5p.add_argument("--model"); p.add_argument("--max-growth", type=float, default=1.20)6a = p.parse_args()78new_mb = os.path.getsize(a.model) / 1e69baseline = json.load(open("baselines/model.json"))["size_mb"]10ratio = new_mb / baseline1112print(f"size {new_mb:.1f} MB vs baseline {baseline:.1f} MB ({ratio:.2f}x)")13if ratio > a.max_growth:14 sys.exit(f"FAIL: model grew {ratio:.2f}x, limit is {a.max_growth:.2f}x")On the Tuesday change this reports 312.0 MB vs baseline 98.4 MB (3.17x) and fails. A 3.17× jump is never accidental in a healthy pipeline — it means either the architecture changed or someone shipped an unquantized artefact by mistake.
Gate three: latency
1# scripts/check_latency.py2import time, numpy as np, onnxruntime as ort, sys, json, argparse34p = argparse.ArgumentParser()5p.add_argument("--model"); p.add_argument("--max-p95-ms", type=float)6a = p.parse_args()78opts = ort.SessionOptions()9opts.intra_op_num_threads = 1 # pin threads: CI runners vary wildly10sess = ort.InferenceSession(a.model, opts, providers=["CPUExecutionProvider"])11name = sess.get_inputs()[0].name12x = np.random.randn(1, 3, 224, 224).astype(np.float32)1314for _ in range(20): # warm up: allocations, kernel selection15 sess.run(None, {name: x})1617samples = []18for _ in range(200):19 t0 = time.perf_counter()20 sess.run(None, {name: x})21 samples.append((time.perf_counter() - t0) * 1000)2223p50, p95 = np.percentile(samples, [50, 95])24print(f"p50 {p50:.1f} ms p95 {p95:.1f} ms")25json.dump({"p50_ms": p50, "p95_ms": p95}, open("reports/latency.json", "w"))26if p95 > a.max_p95_ms:27 sys.exit(f"FAIL: p95 {p95:.1f} ms exceeds budget {a.max_p95_ms:.1f} ms")Three details make this measurement trustworthy rather than noise. Warm up first — the first several runs include memory allocation and kernel selection and can be five times slower than steady state. Pin thread count, because CI runners have between two and sixteen cores depending on the day and an unpinned run varies by 3× for reasons unrelated to your model. Compare p95, not the mean, and use an absolute budget rather than a ratio to the previous run — ratios on a noisy shared runner produce flaky failures, which train people to hit "re-run" without reading the output.
Gate four: accuracy
This one needs a rule for what counts as a regression, and the rule must account for the fact that a fixed held-out set gives you a noisy estimate. On a 2,000-sample test set, a measured F1 of 0.921 has a standard error of roughly 0.006. A drop to 0.918 is well within noise; a drop to 0.895 is not.
1baseline_f1, new_f1 = 0.9210, 0.91302tolerance = 0.005034print(f"F1 {new_f1:.4f} vs baseline {baseline_f1:.4f} (delta {new_f1-baseline_f1:+.4f})")5if baseline_f1 - new_f1 > tolerance:6 sys.exit("FAIL: accuracy regression beyond tolerance")78# Per-class: an aggregate that holds can hide a class that collapsed.9for cls, (base, now) in per_class.items():10 if base - now > 0.03:11 sys.exit(f"FAIL: class '{cls}' fell from {base:.3f} to {now:.3f}")The per-class check is not optional. Swapping two of four class labels moves aggregate F1 by only a few points if those classes are rare, but drives their individual F1 to near zero. Aggregate metrics are exactly the wrong instrument for detecting that.
Stage 3 — build and push
1 build:2 runs-on: ubuntu-latest3 needs: model-validation4 permissions: { contents: read, packages: write }5 outputs:6 digest: ${{ steps.push.outputs.digest }}7 steps:8 - uses: actions/checkout@v59 with: { lfs: true }10 - uses: docker/setup-buildx-action@v411 - uses: docker/login-action@v412 with:13 registry: ghcr.io14 username: ${{ github.actor }}15 password: ${{ secrets.GITHUB_TOKEN }}16 - id: push17 uses: docker/build-push-action@v718 with:19 context: . # use the checkout, so LFS files are real files20 push: true21 tags: |22 ghcr.io/acme/defect-api:${{ github.sha }}23 ghcr.io/acme/defect-api:latest24 cache-from: type=gha25 cache-to: type=gha,mode=max26 - name: Scan for vulnerabilities27 uses: aquasecurity/trivy-action@v0.36.0 # pin a release (better: its commit SHA), never @master28 with:29 image-ref: ghcr.io/acme/defect-api:${{ github.sha }}30 severity: HIGH,CRITICAL31 exit-code: "1"Tagging by commit SHA rather than only latest is what makes rollback possible: every deployment points at an immutable digest, and rolling back is deploying a previous SHA. A pipeline that only produces latest has no rollback target at all.
The context: . line matters for this repository in particular. Without it, the build action fetches the source itself with plain Git, ignoring the checkout step, so the git-lfs model files arrive in the image as 130-byte pointer files. The GitHub Actions cache (type=gha) matters more for ML images than for ordinary ones. Without it, every build reinstalls PyTorch — around four minutes. With it and a correctly ordered Dockerfile, a code-only change rebuilds in under thirty seconds.
Stage 4 — integration test against the real container
Unit tests mock the model. This stage does not. It starts the image you just built and talks to it over HTTP, which is the only way to catch failures that live in the gap between your code and its packaging: a missing system library, a wrong working directory, a model path that exists on your laptop and not in the image.
1 integration:2 runs-on: ubuntu-latest3 needs: build4 steps:5 - uses: actions/checkout@v56 - name: Start the container7 run: |8 docker run -d --name api -p 8000:8000 \9 ghcr.io/acme/defect-api:${{ github.sha }}10 for i in $(seq 1 60); do11 curl -fsS http://localhost:8000/readyz && break12 sleep 213 done14 - run: pytest tests/integration -q --base-url http://localhost:800015 - if: failure()16 run: docker logs apiThe readiness loop rather than a fixed sleep 30 is the difference between a pipeline that is reliable and one that is flaky: model load time varies with runner load, and a fixed sleep is either too short (random failures) or too long (wasted minutes on every build). Dumping container logs on failure saves the round trip of re-running the job with debugging added.
1# tests/integration/test_api.py2import requests, base64, pathlib, pytest34def test_health(base_url):5 assert requests.get(f"{base_url}/readyz", timeout=5).json()["ready"] is True67def test_known_image(base_url):8 img = base64.b64encode(pathlib.Path("tests/fixtures/scratch_01.jpg")9 .read_bytes()).decode()10 r = requests.post(f"{base_url}/predict", json={"image_b64": img}, timeout=30)11 assert r.status_code == 20012 body = r.json()13 assert body["label"] == "scratch"14 assert body["confidence"] > 0.8015 assert body["model_version"] == "2.6.0" # the right artefact is inside1617def test_rejects_garbage(base_url):18 r = requests.post(f"{base_url}/predict", json={"image_b64": "notbase64"}, timeout=10)19 assert r.status_code == 422 # not 5002021@pytest.mark.parametrize("n", [1, 8, 32])22def test_batch_sizes(base_url, n):23 r = requests.post(f"{base_url}/predict/batch",24 json={"items": [SAMPLE] * n}, timeout=60)25 assert len(r.json()["predictions"]) == nAsserting on model_version is a small thing that catches a real and embarrassing failure: the image built correctly but baked in a stale model because a path in the Dockerfile pointed at the wrong directory.
Stages 5 and 6 — staging, then production
1 deploy-staging:2 runs-on: ubuntu-latest3 needs: integration4 if: github.ref == 'refs/heads/main'5 environment: staging6 steps:7 - run: |8 kubectl set image deployment/defect-api \9 api=ghcr.io/acme/defect-api:${{ github.sha }} -n staging10 kubectl rollout status deployment/defect-api -n staging --timeout=300s11 - run: python scripts/smoke.py --url https://staging.internal/predict1213 deploy-production:14 runs-on: ubuntu-latest15 needs: deploy-staging16 environment: production # requires a human approval in repo settings17 steps:18 - name: Canary at 10%19 run: |20 kubectl apply -f k8s/canary-10.yaml21 kubectl set image deployment/defect-api-canary \22 api=ghcr.io/acme/defect-api:${{ github.sha }} -n prod23 kubectl rollout status deployment/defect-api-canary -n prod --timeout=300s24 - name: Watch canary metrics for 10 minutes25 run: python scripts/watch_canary.py --window 600 --error-budget 0.0226 - name: Promote to 100%27 id: promote28 run: |29 kubectl set image deployment/defect-api \30 api=ghcr.io/acme/defect-api:${{ github.sha }} -n prod31 kubectl rollout status deployment/defect-api -n prod --timeout=600s32 kubectl scale deployment/defect-api-canary --replicas=0 -n prodThe environment: production line hooks into GitHub's protected environments, which can require a named reviewer to approve before the job runs. Full automation to production is a goal, not a starting point — earn it by first accumulating evidence that stages 2 and 4 catch what they claim to.
GitLab CI: the same ideas, different nouns
1stages: [quality, validate, build, integration, deploy]23variables:4 IMAGE: $CI_REGISTRY_IMAGE:$CI_COMMIT_SHA56code-quality:7 stage: quality8 image: python:3.11-slim9 script:10 - pip install -r requirements.txt -r requirements-dev.txt11 - ruff check app/ tests/12 - pytest tests/unit -q1314model-validation:15 stage: validate16 image: python:3.11-slim17 script:18 - python scripts/validate_artefact.py --model models/current/19 - python scripts/check_size.py --model models/current/model.onnx --max-growth 1.2020 - python scripts/check_latency.py --model models/current/model.onnx --max-p95-ms 10021 artifacts:22 when: always23 paths: [reports/]24 expire_in: 30 days2526build:27 stage: build28 image: docker:2729 services: [docker:27-dind]30 script:31 - echo "$CI_REGISTRY_PASSWORD" | docker login -u "$CI_REGISTRY_USER" --password-stdin "$CI_REGISTRY"32 - docker build -t "$IMAGE" .33 - docker push "$IMAGE"3435deploy-production:36 stage: deploy37 when: manual # equivalent to a protected environment38 environment: { name: production, url: https://api.acme.com }39 script:40 - kubectl set image deployment/defect-api api="$IMAGE" -n prod41 - kubectl rollout status deployment/defect-api -n prod --timeout=600sThe mapping is direct: jobs become stages plus jobs, needs becomes stage ordering, upload-artifact becomes artifacts, and protected environments become when: manual with an environment block.
Automated rollback
Deploying is only half of it. The pipeline must also be able to undo itself without a human, because the window where a bad model is doing damage is the window where a human is still reading the alert.
1# scripts/watch_canary.py2import time, sys, statistics, argparse, requests34def sample(url):5 m = requests.get(f"{url}/stats", timeout=5).json()6 return m["error_rate"], m["p99_ms"], m["requests"]78def main(window, error_budget):9 canary, baseline = "http://canary.prod.internal", "http://stable.prod.internal"10 deadline = time.time() + window11 while time.time() < deadline:12 c_err, c_p99, c_n = sample(canary)13 b_err, b_p99, _ = sample(baseline)14 print(f"canary n={c_n} err={c_err:.4f} p99={c_p99:.0f}ms | "15 f"stable err={b_err:.4f} p99={b_p99:.0f}ms")1617 if c_n < 1000: # too little data to judge18 time.sleep(30); continue19 if c_err > max(b_err * 2, error_budget):20 sys.exit(f"ROLLBACK: canary error rate {c_err:.4f}")21 if c_p99 > b_p99 * 1.5:22 sys.exit(f"ROLLBACK: canary p99 {c_p99:.0f}ms vs stable {b_p99:.0f}ms")23 time.sleep(30)24 print("canary healthy")2526if __name__ == "__main__":27 p = argparse.ArgumentParser()28 p.add_argument("--window", type=int); p.add_argument("--error-budget", type=float)29 a = p.parse_args(); main(a.window, a.error_budget)1 - name: Pull the canary on failure2 if: failure()3 run: kubectl scale deployment/defect-api-canary --replicas=0 -n prod4 - name: Undo the main rollout if promotion had started5 if: failure() && steps.promote.conclusion == 'failure'6 run: |7 kubectl rollout undo deployment/defect-api -n prod8 kubectl rollout status deployment/defect-api -n prod --timeout=300s9 - name: Tell the team10 if: failure()11 run: |12 curl -X POST "$SLACK_WEBHOOK" \13 -d '{"text":"defect-api release failed and was rolled back"}'These steps go at the end of the deploy-production job. The order of failure matters: if the canary watch fails, the main deployment has not been touched yet, so the fix is to take the canary away — running rollout undo on the main deployment then would roll it back past the version that was healthy. Only if the promotion step itself fails is the main deployment undone.
The c_n < 1000 guard is the same statistical point that governs canaries generally: at 100 requests with a 0.5% baseline error rate you expect 0.5 errors, so observing 2 is ordinary noise. Automation that rolls back on noise gets disabled by frustrated engineers within a fortnight, which leaves you worse off than having no automation.
A rollback that depends on a human noticing, deciding, and typing is not a safety mechanism; it is a hope that someone is awake.
Mistakes that make a model pipeline useless
Mocking the model in every test. Mocks make unit tests fast, which is right — but if nothing in the pipeline loads the real artefact, then a corrupted, mismatched, or missing model passes CI cleanly. At least one stage must execute the actual file.
Committing model binaries to git without LFS. A 300 MB file in git history is permanent, and every clone thereafter pays for it. Use git-lfs or, better, keep artefacts in object storage and commit only a version identifier plus a checksum.
Secrets in the workflow file. A registry password or cloud key committed to a public repository is scraped within minutes. Use the platform's secret store, and prefer short-lived credentials such as OIDC federation over long-lived keys.
Latency gates on unpinned runners. Covered above, and worth restating because the failure mode is social rather than technical: flaky gates get ignored, then bypassed, then removed.
No baseline file under version control. If the numbers you compare against live in someone's spreadsheet, the comparison is not reproducible and nobody can tell when the baseline moved. Keep baselines/model.json in the repository and update it in a reviewed pull request, so a change to the standard is as visible as a change to the code.
Where to start if you have none of this
Building all six stages at once is how these projects stall. Build them in the order of what has already gone wrong for you, and stop at the point where the next stage costs more than the failures it prevents.
If you have nothing, the highest-value single addition is the golden-prediction test: a dozen real inputs, their expected outputs recorded from a known-good model, checked on every build. It takes an hour to write, needs no infrastructure, and catches label reordering, preprocessing drift, corrupted artefacts, and wrong-model-in-the-image — four distinct failure modes with one test.
Second: the integration stage. Starting the real container and hitting the real endpoint catches everything that lives in the gap between code and packaging, which is where a surprising share of production incidents originate.
Third: size and latency gates, because they are cheap and they catch regressions that no accuracy metric will ever show. Only then invest in canary automation and automated rollback, which are the most work and matter only once deployments are frequent enough that a human cannot watch each one.
The measure of whether it is working is not the number of stages. It is whether you would be comfortable letting a colleague merge a model change on a Friday afternoon. If the honest answer is no, the pipeline is telling you which gate is still missing.