Model Deployment for AI Engineers

Containerization with Docker


An engineer pushes a working model server to the deployment repo. CI builds it. The image is 10.6 GB. Pulling it onto a new node takes about six minutes, so every autoscale event adds six minutes of unserved traffic. Worse, each code change — a one-line fix to a log message — triggers a full rebuild that reinstalls PyTorch from scratch and takes eleven minutes.

The same application, built properly, is about 1.4 GB and rebuilds in nine seconds after a code change. Nothing about the model or the API changed. What changed was the order of four lines in a Dockerfile and the choice of one base image.

Containers are not complicated, but the difference between a naive Dockerfile and a considered one is roughly an order of magnitude on every metric that matters in production: image size, rebuild time, cold-start time, and attack surface. This lesson is about that difference.

Where the 1.4 GB and the rebuild time livepython:3.11-slim— 130 MBapt runtimelibrariespip install, CPUtorch — 960 MBmodel weights— per releaseCOPY app code— 9 s rebuildtopbottomCopying code before pip install reinstalls PyTorch on every one-line change — eleven minutes, not nine seconds.
Order the Dockerfile by how often each layer changes, and the slowest layer is built once.

What a container actually is

The problem containers solve is specific. A Python model server depends on far more than its Python packages: a particular glibc, libgomp for OpenMP threading, libjpeg and libpng for Pillow, sometimes CUDA runtime libraries. "It works on my machine" usually means "my machine happens to have libjpeg-turbo 2.1 and the CI runner has 1.5".

A container image bundles all of that — the application, its Python packages, and the system libraries underneath — into a single addressable artefact. The container is that image running as an isolated process on a shared kernel.

TermWhat it isAnalogy
ImageRead-only stack of filesystem layers plus metadataA class
ContainerA running image with a thin writable layer on topAn instance
LayerThe filesystem diff produced by one build instructionA git commit
DockerfileThe recipe that produces the layersSource code
RegistryWhere images are stored and pulled fromA package index
VolumeStorage that outlives the containerA mounted disk

The distinction from a virtual machine is worth being precise about, because it explains the performance characteristics. A VM ships a whole guest kernel and boots it — hundreds of milliseconds to seconds, gigabytes of overhead. A container is just a Linux process with namespaces restricting what it can see (its own view of the filesystem, process table, and network stack) and cgroups restricting what it can consume (CPU shares, memory limits). Start-up is process start-up: tens of milliseconds. There is no virtualisation overhead on the inference itself.

A container gives you the reproducibility of shipping a whole machine at roughly the cost of starting a process.

Layers and the cache — the part that determines rebuild time

Each instruction in a Dockerfile produces a layer. Docker caches layers by a hash of the instruction and its inputs. On rebuild, it reuses cached layers until it hits the first changed instruction — and then every layer after that is rebuilt, whether or not it depended on the change.

That cascade is the whole game. Consider the naive ordering:

Text
FROM python:3.11          cachedCOPY . .                  ← your source changed, so this layer is invalidatedRUN pip install -r requirements.txt   ← rebuilt: reinstalls PyTorch, 11 minutes

versus the correct ordering:

Text
FROM python:3.11-slim     cachedCOPY requirements.txt .   cached (the file did not change)RUN pip install -r requirements.txt   cached (960 MB of wheels, untouched)COPY . .                  ← rebuilt: 9 seconds

The rule generalises: order instructions from least likely to change to most likely to change. System packages, then Python dependencies, then model weights, then application code. Application code changes fifty times a day; the dependency list changes twice a month.

A production Dockerfile, line by line

Dockerfile
FROM python:3.11-slim-bookworm# System libraries the Python wheels link against. Not optional for Pillow.RUN apt-get update \ && apt-get install -y --no-install-recommends libgomp1 libjpeg62-turbo curl \ && rm -rf /var/lib/apt/lists/*ENV PYTHONUNBUFFERED=1 \    PYTHONDONTWRITEBYTECODE=1 \    PIP_NO_CACHE_DIR=1 \    OMP_NUM_THREADS=1WORKDIR /app# Dependencies first: this layer survives every code change.COPY requirements.txt .RUN pip install --no-cache-dir -r requirements.txt# Model weights next: change per release, not per commit.COPY models/ /app/models/# Application code last: changes constantly.COPY app/ /app/app/# Never run as root.RUN useradd --create-home --uid 10001 appuser \ && chown -R appuser:appuser /appUSER appuserEXPOSE 8000HEALTHCHECK --interval=30s --timeout=3s --start-period=40s --retries=3 \  CMD curl -fsS http://localhost:8000/health || exit 1CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]

Several of those lines fix a specific failure.

python:3.11-slim-bookworm, not python:3.11. The full image is about 1.0 GB because it carries build toolchains, documentation, and a large set of development headers. Slim is around 130 MB. Both run the same Python. Pinning the Debian release (bookworm) matters too: slim alone will silently change base OS at the next Debian release, which is exactly the kind of surprise containers are supposed to prevent.

apt-get update and install in one RUN. Split across two instructions, Docker caches the update layer, and weeks later your install pulls package versions from a stale index that no longer exist. Deleting /var/lib/apt/lists in the same instruction is also required — do it in a later RUN and the files still exist in the earlier layer, so the image is just as big with an extra layer of confusion.

PYTHONUNBUFFERED=1. Without it, Python buffers stdout when it is not a terminal. Your logs appear in 8 KB chunks, and if the container is killed the last chunk is lost — which is precisely the output describing the crash. Skipping this line is why "the container died and there's nothing in the logs".

OMP_NUM_THREADS=1. PyTorch and NumPy default to one compute thread per visible core. Inside a container with a 2-core CPU limit, they often still see all 32 host cores and spawn 32 threads that thrash against a cgroup quota. Setting this explicitly, and running multiple uvicorn workers instead, gives markedly better and far more predictable throughput.

USER appuser. Containers run as root by default. Combined with any container-escape vulnerability, or a plain volume mount, root inside means root-adjacent damage outside. The line costs nothing.

--start-period=40s on the health check. Loading a PyTorch model takes 15–30 seconds. Without a grace period the health check fails three times during start-up and the orchestrator kills the container before it ever becomes ready — producing a restart loop that looks like a crash but is purely a timing misconfiguration.

CMD in exec form. The JSON-array form runs uvicorn as PID 1 directly. The shell form (CMD uvicorn ...) wraps it in /bin/sh -c, and sh does not forward SIGTERM to its child. Your container then ignores graceful shutdown and gets SIGKILLed after the timeout, dropping every in-flight request on every deploy.

The .dockerignore file

Without one, COPY . . sends your entire working directory to the build daemon — including .git (often hundreds of megabytes), virtual environments, datasets, and notebook checkpoints. It also bakes any .env file into the image, where anyone who can pull it can read your credentials.

Text
.git.gitignore__pycache__/*.py[cod].venv/venv/.pytest_cache/.ipynb_checkpoints/notebooks/data/tests/*.md.env.env.*Dockerfiledocker-compose.yml

Multi-stage builds

Some Python packages need a C compiler to install but not to run. Installing build-essential adds roughly 300 MB that is pure dead weight at runtime. A multi-stage build compiles in one image and copies only the results into a clean final image.

Dockerfile
FROM python:3.11-slim-bookworm AS builderRUN apt-get update \ && apt-get install -y --no-install-recommends build-essential gcc g++ \ && rm -rf /var/lib/apt/lists/*RUN python -m venv /opt/venvENV PATH="/opt/venv/bin:$PATH"COPY requirements.txt .RUN pip install --no-cache-dir -r requirements.txt# ---------- runtime ----------FROM python:3.11-slim-bookwormRUN apt-get update \ && apt-get install -y --no-install-recommends libgomp1 libjpeg62-turbo curl \ && rm -rf /var/lib/apt/lists/*COPY --from=builder /opt/venv /opt/venvENV PATH="/opt/venv/bin:$PATH" PYTHONUNBUFFERED=1 OMP_NUM_THREADS=1WORKDIR /appCOPY models/ /app/models/COPY app/ /app/app/RUN useradd --create-home --uid 10001 appuser && chown -R appuser:appuser /appUSER appuserEXPOSE 8000CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000"]

The virtual environment is the trick. Everything pip installed lands in /opt/venv, so a single COPY --from=builder moves the whole dependency set without dragging the compilers along. The builder stage never appears in the final image or the registry.

Where the gigabytes actually are

For an ML image, the base image is not the problem — the dependencies are. Numbers for a PyTorch image classifier on Python 3.11 and x86-64 Linux, with PyTorch 2.14 (installed sizes, not download sizes):

ChoiceContributionBetter choiceContribution
python:3.111,010 MBpython:3.11-slim-bookworm130 MB
torch (default = CUDA build, with its NVIDIA libraries)~5,200 MBtorch from the CPU index~750 MB
build-essential left in~300 MBMulti-stage build0 MB
pip cache retained (the downloaded wheels, kept a second time)~3,000 MB--no-cache-dir0 MB
.git and datasets copied in~600 MB.dockerignore0 MB
App, model weights, other deps~515 MBunchanged515 MB
Total~10,600 MB~1,400 MB

The single biggest line is one flag. If you are serving on CPU, the default Linux build of PyTorch pulls in around 4.5 GB of CUDA, cuDNN, NCCL and Triton libraries you will never execute:

Text
# requirements.txt--extra-index-url https://download.pytorch.org/whl/cputorch==2.14.0+cputorchvision==0.29.0+cpufastapi==0.141.1uvicorn[standard]==0.53.0pillow==12.3.0

Before optimising anything else, check whether you are shipping GPU kernels to a CPU-only service — it is usually the largest single item in the image.

Pin exact versions, not ranges. torch>=2.0 means today's build and next month's build are different software, which defeats the purpose of containerising in the first place.

Building, running, and inspecting

Bash
# Build with a real tag. "latest" tells you nothing during an incident.docker build -t defect-api:2.1.0 .docker tag defect-api:2.1.0 registry.example.com/ml/defect-api:2.1.0# Where did the size go?docker history defect-api:2.1.0 --human --format "{{.Size}}\t{{.CreatedBy}}"# Run with the resource limits production will actually impose.docker run -d --name defect-api \  -p 8000:8000 \  --memory 2g --cpus 2 \  --restart unless-stopped \  -e LOG_LEVEL=info \  -v /srv/models:/app/models:ro \  defect-api:2.1.0docker logs -f defect-apidocker stats defect-api          # live CPU / memory / networkdocker exec -it defect-api /bin/bashdocker inspect defect-api --format '{{.State.Health.Status}}'

Two of those flags matter more than they look.

--memory 2g --cpus 2 during local testing is how you discover, on your laptop rather than at 03:00, that four uvicorn workers each holding a 180 MB model need 720 MB plus overhead. Testing without limits on a 32 GB machine tells you nothing about a 2 GB pod.

-v /srv/models:/app/models:ro mounts weights from the host read-only rather than baking them into the image. This is a genuine architectural fork:

Weights baked into the imageWeights mounted at runtime
ReproducibilityTotal — image digest pins the exact modelPartial — same image, different model
Image size+ model size, on every pullSmall image
Swapping a modelRebuild and redeployChange the mount, restart
RollbackOne image tag rolls back everythingTwo things to roll back independently
Best forProduction releasesDevelopment, or models over a few GB

For production, prefer baking them in. One immutable digest that fully determines behaviour is worth the extra pull time, and it removes the "which model was actually loaded?" question from every future incident.

Docker Compose for the whole stack

A model service rarely runs alone. Compose describes the set and their relationships in one file.

YAML
services:  api:    build: .    image: defect-api:2.1.0    expose: ["8000"]            # reachable by the proxy, not published to the host    environment:      REDIS_URL: redis://cache:6379/0      LOG_LEVEL: info      OMP_NUM_THREADS: "1"    depends_on:      cache:        condition: service_healthy    deploy:      resources:        limits: {cpus: "2.0", memory: 2G}    healthcheck:      test: ["CMD", "curl", "-fsS", "http://localhost:8000/health"]      interval: 30s      timeout: 3s      start_period: 40s      retries: 3    restart: unless-stopped  cache:    image: redis:7-alpine    command: ["redis-server", "--maxmemory", "256mb", "--maxmemory-policy", "allkeys-lru"]    volumes: ["redis-data:/data"]    healthcheck:      test: ["CMD", "redis-cli", "ping"]      interval: 10s      retries: 5  proxy:    image: nginx:1.27-alpine    ports: ["80:80"]    volumes: ["./nginx.conf:/etc/nginx/nginx.conf:ro"]    depends_on: [api]volumes:  redis-data:

The subtlety worth knowing is depends_on. The plain list form only waits for the container to start, not to be usable — so your API connects to a Redis that is still loading its dump file and crashes on the first command. The condition: service_healthy form waits for the health check to pass. Use it, and give every dependency a health check.

Note also that services address each other by service name (redis://cache:6379), resolved by Compose's internal DNS. Only proxy publishes a port to the host; api and cache are reachable only from inside the network.

Bash
docker compose up -d --builddocker compose logs -f apidocker compose psdocker compose up -d --scale api=3     # three API replicas behind the proxydocker compose down -v                 # -v also deletes named volumes

Security: the four things that go wrong

Secrets baked into layers

This is the most common serious mistake, and it survives deletion:

Dockerfile
# WRONG — the key is now permanently in layer 3COPY .env /app/.envRUN python fetch_model.pyRUN rm /app/.env          # deletes it from the final filesystem, NOT from the image

Layers are immutable and additive. Anyone who pulls the image can extract layer 3 and read the key. The correct approach is a build secret, which is mounted only for the duration of one instruction and never written to a layer:

Dockerfile
# syntax=docker/dockerfile:1.7RUN --mount=type=secret,id=hf_token \    HF_TOKEN=$(cat /run/secrets/hf_token) python fetch_model.py
Bash
docker build --secret id=hf_token,src=./hf_token.txt -t defect-api:2.1.0 .

Runtime secrets come from the environment or a secrets manager, never from the image.

Running as root

Covered above, but worth restating with the consequence: root in the container plus a writable volume mount means an attacker who achieves code execution in your Python process can write to host paths. The USER instruction is one line and removes the whole category. Also add --read-only and --cap-drop ALL at run time where the application allows it.

Unscanned base images

Your image inherits every CVE in its base. Scan on every build and fail the pipeline on high-severity findings:

Bash
trivy image --severity HIGH,CRITICAL --exit-code 1 defect-api:2.1.0docker scout cves defect-api:2.1.0

Rebuilding weekly against a patched base image resolves most findings for free, which is a good argument for keeping rebuilds fast.

Floating tags

FROM python:3.11 resolves to a different image every few weeks. Pin by digest when reproducibility genuinely matters:

Dockerfile
FROM python:3.11-slim-bookworm@sha256:d5b1fbbc00fd7ec7e4a4dcbd5ce0d4b6c6f8f89a3e2d1c0b9a8f7e6d5c4b3a29

What containerising well actually gets you

The payoff is not "it runs anywhere". It is that four separate operational problems stop being problems.

Autoscaling becomes viable. A 1.4 GB image pulls in under a minute on a typical node; a 10 GB image takes about six. If your traffic spike lasts five minutes, only one of those images helps you.

Iteration stays fast. A correctly ordered Dockerfile rebuilds in seconds after a code change, so developers actually test in the container rather than only in their virtualenv — which is where environment drift gets caught.

Rollback becomes trivial. If weights, code, and dependencies all live in one immutable digest, rolling back is repointing a tag. If any of the three lives outside the image, rollback is a procedure with steps that can be performed wrong under pressure.

Incidents get shorter. docker inspect on the running container tells you the exact image digest, and the digest tells you the exact model, code, and every dependency version. That turns "which version is running?" from an investigation into a command.

A useful discipline: before merging any Dockerfile change, run docker history on the result and look at the three largest layers. If any of them surprises you, you have found something worth fixing — and it is almost always either a CUDA build you do not need or a directory you forgot to put in .dockerignore.