Course Content
Model Deployment for AI Engineers
4 sections · 10 lessons
Containerization with Docker
An engineer pushes a working model server to the deployment repo. CI builds it. The image is 10.6 GB. Pulling it onto a new node takes about six minutes, so every autoscale event adds six minutes of unserved traffic. Worse, each code change — a one-line fix to a log message — triggers a full rebuild that reinstalls PyTorch from scratch and takes eleven minutes.
The same application, built properly, is about 1.4 GB and rebuilds in nine seconds after a code change. Nothing about the model or the API changed. What changed was the order of four lines in a Dockerfile and the choice of one base image.
Containers are not complicated, but the difference between a naive Dockerfile and a considered one is roughly an order of magnitude on every metric that matters in production: image size, rebuild time, cold-start time, and attack surface. This lesson is about that difference.
What a container actually is
The problem containers solve is specific. A Python model server depends on far more than its Python packages: a particular glibc, libgomp for OpenMP threading, libjpeg and libpng for Pillow, sometimes CUDA runtime libraries. "It works on my machine" usually means "my machine happens to have libjpeg-turbo 2.1 and the CI runner has 1.5".
A container image bundles all of that — the application, its Python packages, and the system libraries underneath — into a single addressable artefact. The container is that image running as an isolated process on a shared kernel.
| Term | What it is | Analogy |
|---|---|---|
| Image | Read-only stack of filesystem layers plus metadata | A class |
| Container | A running image with a thin writable layer on top | An instance |
| Layer | The filesystem diff produced by one build instruction | A git commit |
| Dockerfile | The recipe that produces the layers | Source code |
| Registry | Where images are stored and pulled from | A package index |
| Volume | Storage that outlives the container | A mounted disk |
The distinction from a virtual machine is worth being precise about, because it explains the performance characteristics. A VM ships a whole guest kernel and boots it — hundreds of milliseconds to seconds, gigabytes of overhead. A container is just a Linux process with namespaces restricting what it can see (its own view of the filesystem, process table, and network stack) and cgroups restricting what it can consume (CPU shares, memory limits). Start-up is process start-up: tens of milliseconds. There is no virtualisation overhead on the inference itself.
A container gives you the reproducibility of shipping a whole machine at roughly the cost of starting a process.
Layers and the cache — the part that determines rebuild time
Each instruction in a Dockerfile produces a layer. Docker caches layers by a hash of the instruction and its inputs. On rebuild, it reuses cached layers until it hits the first changed instruction — and then every layer after that is rebuilt, whether or not it depended on the change.
That cascade is the whole game. Consider the naive ordering:
FROM python:3.11 cachedCOPY . . ← your source changed, so this layer is invalidatedRUN pip install -r requirements.txt ← rebuilt: reinstalls PyTorch, 11 minutesversus the correct ordering:
FROM python:3.11-slim cachedCOPY requirements.txt . cached (the file did not change)RUN pip install -r requirements.txt cached (960 MB of wheels, untouched)COPY . . ← rebuilt: 9 secondsThe rule generalises: order instructions from least likely to change to most likely to change. System packages, then Python dependencies, then model weights, then application code. Application code changes fifty times a day; the dependency list changes twice a month.
A production Dockerfile, line by line
1FROM python:3.11-slim-bookworm23# System libraries the Python wheels link against. Not optional for Pillow.4RUN apt-get update \5 && apt-get install -y --no-install-recommends libgomp1 libjpeg62-turbo curl \6 && rm -rf /var/lib/apt/lists/*78ENV PYTHONUNBUFFERED=1 \9 PYTHONDONTWRITEBYTECODE=1 \10 PIP_NO_CACHE_DIR=1 \11 OMP_NUM_THREADS=11213WORKDIR /app1415# Dependencies first: this layer survives every code change.16COPY requirements.txt .17RUN pip install --no-cache-dir -r requirements.txt1819# Model weights next: change per release, not per commit.20COPY models/ /app/models/2122# Application code last: changes constantly.23COPY app/ /app/app/2425# Never run as root.26RUN useradd --create-home --uid 10001 appuser \27 && chown -R appuser:appuser /app28USER appuser2930EXPOSE 80003132HEALTHCHECK --interval=30s --timeout=3s --start-period=40s --retries=3 \33 CMD curl -fsS http://localhost:8000/health || exit 13435CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]Several of those lines fix a specific failure.
python:3.11-slim-bookworm, not python:3.11. The full image is about 1.0 GB because it carries build toolchains, documentation, and a large set of development headers. Slim is around 130 MB. Both run the same Python. Pinning the Debian release (bookworm) matters too: slim alone will silently change base OS at the next Debian release, which is exactly the kind of surprise containers are supposed to prevent.
apt-get update and install in one RUN. Split across two instructions, Docker caches the update layer, and weeks later your install pulls package versions from a stale index that no longer exist. Deleting /var/lib/apt/lists in the same instruction is also required — do it in a later RUN and the files still exist in the earlier layer, so the image is just as big with an extra layer of confusion.
PYTHONUNBUFFERED=1. Without it, Python buffers stdout when it is not a terminal. Your logs appear in 8 KB chunks, and if the container is killed the last chunk is lost — which is precisely the output describing the crash. Skipping this line is why "the container died and there's nothing in the logs".
OMP_NUM_THREADS=1. PyTorch and NumPy default to one compute thread per visible core. Inside a container with a 2-core CPU limit, they often still see all 32 host cores and spawn 32 threads that thrash against a cgroup quota. Setting this explicitly, and running multiple uvicorn workers instead, gives markedly better and far more predictable throughput.
USER appuser. Containers run as root by default. Combined with any container-escape vulnerability, or a plain volume mount, root inside means root-adjacent damage outside. The line costs nothing.
--start-period=40s on the health check. Loading a PyTorch model takes 15–30 seconds. Without a grace period the health check fails three times during start-up and the orchestrator kills the container before it ever becomes ready — producing a restart loop that looks like a crash but is purely a timing misconfiguration.
CMD in exec form. The JSON-array form runs uvicorn as PID 1 directly. The shell form (CMD uvicorn ...) wraps it in /bin/sh -c, and sh does not forward SIGTERM to its child. Your container then ignores graceful shutdown and gets SIGKILLed after the timeout, dropping every in-flight request on every deploy.
The .dockerignore file
Without one, COPY . . sends your entire working directory to the build daemon — including .git (often hundreds of megabytes), virtual environments, datasets, and notebook checkpoints. It also bakes any .env file into the image, where anyone who can pull it can read your credentials.
.git.gitignore__pycache__/*.py[cod].venv/venv/.pytest_cache/.ipynb_checkpoints/notebooks/data/tests/*.md.env.env.*Dockerfiledocker-compose.ymlMulti-stage builds
Some Python packages need a C compiler to install but not to run. Installing build-essential adds roughly 300 MB that is pure dead weight at runtime. A multi-stage build compiles in one image and copies only the results into a clean final image.
1FROM python:3.11-slim-bookworm AS builder23RUN apt-get update \4 && apt-get install -y --no-install-recommends build-essential gcc g++ \5 && rm -rf /var/lib/apt/lists/*67RUN python -m venv /opt/venv8ENV PATH="/opt/venv/bin:$PATH"910COPY requirements.txt .11RUN pip install --no-cache-dir -r requirements.txt1213# ---------- runtime ----------14FROM python:3.11-slim-bookworm1516RUN apt-get update \17 && apt-get install -y --no-install-recommends libgomp1 libjpeg62-turbo curl \18 && rm -rf /var/lib/apt/lists/*1920COPY --from=builder /opt/venv /opt/venv21ENV PATH="/opt/venv/bin:$PATH" PYTHONUNBUFFERED=1 OMP_NUM_THREADS=12223WORKDIR /app24COPY models/ /app/models/25COPY app/ /app/app/2627RUN useradd --create-home --uid 10001 appuser && chown -R appuser:appuser /app28USER appuser2930EXPOSE 800031CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000"]The virtual environment is the trick. Everything pip installed lands in /opt/venv, so a single COPY --from=builder moves the whole dependency set without dragging the compilers along. The builder stage never appears in the final image or the registry.
Where the gigabytes actually are
For an ML image, the base image is not the problem — the dependencies are. Numbers for a PyTorch image classifier on Python 3.11 and x86-64 Linux, with PyTorch 2.14 (installed sizes, not download sizes):
| Choice | Contribution | Better choice | Contribution |
|---|---|---|---|
python:3.11 | 1,010 MB | python:3.11-slim-bookworm | 130 MB |
torch (default = CUDA build, with its NVIDIA libraries) | ~5,200 MB | torch from the CPU index | ~750 MB |
build-essential left in | ~300 MB | Multi-stage build | 0 MB |
| pip cache retained (the downloaded wheels, kept a second time) | ~3,000 MB | --no-cache-dir | 0 MB |
.git and datasets copied in | ~600 MB | .dockerignore | 0 MB |
| App, model weights, other deps | ~515 MB | unchanged | 515 MB |
| Total | ~10,600 MB | ~1,400 MB |
The single biggest line is one flag. If you are serving on CPU, the default Linux build of PyTorch pulls in around 4.5 GB of CUDA, cuDNN, NCCL and Triton libraries you will never execute:
# requirements.txt--extra-index-url https://download.pytorch.org/whl/cputorch==2.14.0+cputorchvision==0.29.0+cpufastapi==0.141.1uvicorn[standard]==0.53.0pillow==12.3.0Before optimising anything else, check whether you are shipping GPU kernels to a CPU-only service — it is usually the largest single item in the image.
Pin exact versions, not ranges. torch>=2.0 means today's build and next month's build are different software, which defeats the purpose of containerising in the first place.
Building, running, and inspecting
1# Build with a real tag. "latest" tells you nothing during an incident.2docker build -t defect-api:2.1.0 .3docker tag defect-api:2.1.0 registry.example.com/ml/defect-api:2.1.045# Where did the size go?6docker history defect-api:2.1.0 --human --format "{{.Size}}\t{{.CreatedBy}}"78# Run with the resource limits production will actually impose.9docker run -d --name defect-api \10 -p 8000:8000 \11 --memory 2g --cpus 2 \12 --restart unless-stopped \13 -e LOG_LEVEL=info \14 -v /srv/models:/app/models:ro \15 defect-api:2.1.01617docker logs -f defect-api18docker stats defect-api # live CPU / memory / network19docker exec -it defect-api /bin/bash20docker inspect defect-api --format '{{.State.Health.Status}}'Two of those flags matter more than they look.
--memory 2g --cpus 2 during local testing is how you discover, on your laptop rather than at 03:00, that four uvicorn workers each holding a 180 MB model need 720 MB plus overhead. Testing without limits on a 32 GB machine tells you nothing about a 2 GB pod.
-v /srv/models:/app/models:ro mounts weights from the host read-only rather than baking them into the image. This is a genuine architectural fork:
| Weights baked into the image | Weights mounted at runtime | |
|---|---|---|
| Reproducibility | Total — image digest pins the exact model | Partial — same image, different model |
| Image size | + model size, on every pull | Small image |
| Swapping a model | Rebuild and redeploy | Change the mount, restart |
| Rollback | One image tag rolls back everything | Two things to roll back independently |
| Best for | Production releases | Development, or models over a few GB |
For production, prefer baking them in. One immutable digest that fully determines behaviour is worth the extra pull time, and it removes the "which model was actually loaded?" question from every future incident.
Docker Compose for the whole stack
A model service rarely runs alone. Compose describes the set and their relationships in one file.
1services:2 api:3 build: .4 image: defect-api:2.1.05 expose: ["8000"] # reachable by the proxy, not published to the host6 environment:7 REDIS_URL: redis://cache:6379/08 LOG_LEVEL: info9 OMP_NUM_THREADS: "1"10 depends_on:11 cache:12 condition: service_healthy13 deploy:14 resources:15 limits: {cpus: "2.0", memory: 2G}16 healthcheck:17 test: ["CMD", "curl", "-fsS", "http://localhost:8000/health"]18 interval: 30s19 timeout: 3s20 start_period: 40s21 retries: 322 restart: unless-stopped2324 cache:25 image: redis:7-alpine26 command: ["redis-server", "--maxmemory", "256mb", "--maxmemory-policy", "allkeys-lru"]27 volumes: ["redis-data:/data"]28 healthcheck:29 test: ["CMD", "redis-cli", "ping"]30 interval: 10s31 retries: 53233 proxy:34 image: nginx:1.27-alpine35 ports: ["80:80"]36 volumes: ["./nginx.conf:/etc/nginx/nginx.conf:ro"]37 depends_on: [api]3839volumes:40 redis-data:The subtlety worth knowing is depends_on. The plain list form only waits for the container to start, not to be usable — so your API connects to a Redis that is still loading its dump file and crashes on the first command. The condition: service_healthy form waits for the health check to pass. Use it, and give every dependency a health check.
Note also that services address each other by service name (redis://cache:6379), resolved by Compose's internal DNS. Only proxy publishes a port to the host; api and cache are reachable only from inside the network.
1docker compose up -d --build2docker compose logs -f api3docker compose ps4docker compose up -d --scale api=3 # three API replicas behind the proxy5docker compose down -v # -v also deletes named volumesSecurity: the four things that go wrong
Secrets baked into layers
This is the most common serious mistake, and it survives deletion:
1# WRONG — the key is now permanently in layer 32COPY .env /app/.env3RUN python fetch_model.py4RUN rm /app/.env # deletes it from the final filesystem, NOT from the imageLayers are immutable and additive. Anyone who pulls the image can extract layer 3 and read the key. The correct approach is a build secret, which is mounted only for the duration of one instruction and never written to a layer:
# syntax=docker/dockerfile:1.7RUN --mount=type=secret,id=hf_token \ HF_TOKEN=$(cat /run/secrets/hf_token) python fetch_model.pydocker build --secret id=hf_token,src=./hf_token.txt -t defect-api:2.1.0 .Runtime secrets come from the environment or a secrets manager, never from the image.
Running as root
Covered above, but worth restating with the consequence: root in the container plus a writable volume mount means an attacker who achieves code execution in your Python process can write to host paths. The USER instruction is one line and removes the whole category. Also add --read-only and --cap-drop ALL at run time where the application allows it.
Unscanned base images
Your image inherits every CVE in its base. Scan on every build and fail the pipeline on high-severity findings:
trivy image --severity HIGH,CRITICAL --exit-code 1 defect-api:2.1.0docker scout cves defect-api:2.1.0Rebuilding weekly against a patched base image resolves most findings for free, which is a good argument for keeping rebuilds fast.
Floating tags
FROM python:3.11 resolves to a different image every few weeks. Pin by digest when reproducibility genuinely matters:
FROM python:3.11-slim-bookworm@sha256:d5b1fbbc00fd7ec7e4a4dcbd5ce0d4b6c6f8f89a3e2d1c0b9a8f7e6d5c4b3a29What containerising well actually gets you
The payoff is not "it runs anywhere". It is that four separate operational problems stop being problems.
Autoscaling becomes viable. A 1.4 GB image pulls in under a minute on a typical node; a 10 GB image takes about six. If your traffic spike lasts five minutes, only one of those images helps you.
Iteration stays fast. A correctly ordered Dockerfile rebuilds in seconds after a code change, so developers actually test in the container rather than only in their virtualenv — which is where environment drift gets caught.
Rollback becomes trivial. If weights, code, and dependencies all live in one immutable digest, rolling back is repointing a tag. If any of the three lives outside the image, rollback is a procedure with steps that can be performed wrong under pressure.
Incidents get shorter. docker inspect on the running container tells you the exact image digest, and the digest tells you the exact model, code, and every dependency version. That turns "which version is running?" from an investigation into a command.
A useful discipline: before merging any Dockerfile change, run docker history on the result and look at the three largest layers. If any of them surprises you, you have found something worth fixing — and it is almost always either a CUDA build you do not need or a directory you forgot to put in .dockerignore.