Course Content
MLOps for AI
4 sections · 9 lessons
Container Orchestration with Kubernetes
The recommendation service runs as one container on one virtual machine, started months ago with docker run -d. It has worked fine at 40 requests per second.
On the first day of a sale, traffic reaches 400 requests per second. The p99 latency goes from 90 ms to 14 seconds because a single Python process cannot serve ten times the load. Then a batch request arrives asking for recommendations for 5,000 users at once; the process allocates about 9 GB, the kernel's out-of-memory killer terminates it, and nothing restarts it. The service is down for 47 minutes, until a customer complains loudly enough to reach an engineer.
Every one of those failures has a name and a standard answer: no horizontal scaling, no restart policy, no memory bound, no health checking, no alert. Assembling those answers yourself — a supervisor, a load balancer, a scaler, a health checker, a deployer — is roughly what Kubernetes is. It is worth understanding it as a set of answers to specific failures rather than as a platform you adopt wholesale.
The building blocks, and the failure each one prevents
Kubernetes works on one principle: you declare the state you want, and a control loop continuously works to make reality match. You never say "start a container". You say "four of these should exist", and something else notices when three do.
| Object | What it is | Failure it answers |
|---|---|---|
| Pod | One or more containers sharing a network namespace and lifecycle | None directly; it is the unit everything else manages |
| Deployment | "Keep N healthy pods of this spec"; handles rolling updates | Crashed process never restarts; no way to update without downtime |
| Service | Stable virtual IP and DNS name that load-balances across matching pods | Pod IPs change on every restart; callers hardcode a dead address |
| Ingress | HTTP routing from outside the cluster, with TLS termination | No public entry point; certificates managed by hand |
| ConfigMap / Secret | Configuration and credentials injected as env vars or files | Secrets baked into images; config changes need a rebuild |
| HorizontalPodAutoscaler | Adjusts replica count from a metric | The sale-day capacity problem |
| PersistentVolumeClaim | Storage that outlives a pod | Model cache re-downloaded on every restart |
| Namespace | A scope for names, quotas and access control | Staging and production sharing a blast radius |
Nothing in Kubernetes is a command. Everything is a statement about the desired world, plus a controller whose whole job is to notice disagreement and fix it.
A manifest for a model server, line by line
apiVersion: apps/v1kind: Deploymentmetadata: name: recommender namespace: ml-production labels: {app: recommender}spec: replicas: 4 strategy: type: RollingUpdate rollingUpdate: {maxSurge: 1, maxUnavailable: 0} selector: matchLabels: {app: recommender} template: metadata: labels: {app: recommender} annotations: prometheus.io/scrape: "true" prometheus.io/port: "8000" spec: securityContext: runAsNonRoot: true runAsUser: 1001 containers: - name: server image: ghcr.io/acme/recommender@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d ports: - {containerPort: 8000, name: http} env: - name: MODEL_ALIAS value: champion - name: MLFLOW_TRACKING_URI valueFrom: configMapKeyRef: {name: ml-config, key: tracking_uri} - name: AWS_SECRET_ACCESS_KEY valueFrom: secretKeyRef: {name: model-store, key: secret_key} resources: requests: {cpu: "1000m", memory: "4Gi"} limits: {cpu: "2000m", memory: "6Gi"} startupProbe: httpGet: {path: /health, port: 8000} failureThreshold: 30 periodSeconds: 5 readinessProbe: httpGet: {path: /ready, port: 8000} periodSeconds: 10 failureThreshold: 3 livenessProbe: httpGet: {path: /health, port: 8000} periodSeconds: 20 failureThreshold: 3 lifecycle: preStop: exec: {command: ["sleep", "10"]} terminationGracePeriodSeconds: 60Requests and limits are two different things
This is the single most misunderstood pair of numbers in the file.
requests | limits | |
|---|---|---|
| Used by | The scheduler, to decide which node has room | The kernel, to enforce a ceiling at runtime |
| Memory behaviour | Reserves capacity; the pod will not be placed without it | Exceeding it means an immediate OOMKill — no warning, no grace |
| CPU behaviour | Guarantees a share under contention | Throttles: the process is stopped until the next scheduling period |
| Set too low | Pods packed onto a node, then evicted under pressure | Constant OOMKills or crippling throttling |
| Set too high | Nodes half-empty; you pay for reserved-but-unused capacity | Little harm for memory; none for CPU |
Work the packing arithmetic. A node has 16 GiB of RAM, of which roughly 1.5 GiB goes to the kubelet, the container runtime and system daemons — about 14.5 GiB schedulable. With requests.memory: 4Gi, the scheduler fits ⌊14.5 / 4⌋ = 3 pods per node. Four replicas therefore need two nodes. Drop the request to 3Gi and ⌊14.5 / 3⌋ = 4 pods fit on one node — but if the model genuinely resides at 3.2 GiB, you have set a request below actual usage, and under memory pressure the kubelet evicts your pods before anything else because they exceeded their request.
CPU limits deserve their own warning because the failure is invisible. The Linux scheduler enforces a CPU limit as a quota per 100 ms period: limits.cpu: 2000m means 200 ms of CPU time per 100 ms of wall clock, spread across all threads. A model server whose inference library spawns 8 BLAS threads can burn that entire quota in 25 ms and then be frozen for the remaining 75 ms of the period. Latency graphs show mysterious 75 ms plateaus that no profiler explains. The fix is usually to cap the library's thread count to match the limit — OMP_NUM_THREADS=2 — rather than to raise the limit.
A memory limit is a promise the kernel makes to kill you; a memory request is a promise the scheduler makes to you. Confusing them yields either evictions or half-empty nodes.
Three probes, three jobs
| Probe | Question | On failure | ML-specific reason it matters |
|---|---|---|---|
startupProbe | Has it finished booting? | Kills the pod after the threshold | Loading a multi-gigabyte model can take 90 s; without this, the liveness probe kills it mid-load, forever |
readinessProbe | Can it take traffic now? | Removes it from the Service endpoints; does not kill it | Lets a pod refreshing its model temporarily leave the rotation |
livenessProbe | Is it wedged and unrecoverable? | Restarts the container | Catches a deadlocked inference thread that still accepts connections |
The startupProbe here allows 30 × 5 = 150 seconds to load. Until it succeeds, the liveness and readiness probes are suspended entirely. This is the fix for the classic crash loop where a large model never finishes loading because the liveness probe kills it at 30 seconds, every time, indefinitely.
The preStop: sleep 10 paired with terminationGracePeriodSeconds: 60 solves a subtler problem. When a pod is deleted, two things happen concurrently: the container gets SIGTERM, and the endpoint is removed from the Service. Endpoint removal propagates through kube-proxy on every node and is not instant. Without the sleep, the process exits before some nodes have stopped routing to it, and a handful of requests get connection-refused during every deploy. Ten seconds of doing nothing before shutdown eliminates that entirely.
Finally, maxUnavailable: 0 with maxSurge: 1 means a new pod must be ready before an old one is removed. Capacity never dips during a rollout — at the cost of needing room for one extra pod.
Autoscaling with arithmetic you can predict
apiVersion: autoscaling/v2kind: HorizontalPodAutoscalermetadata: {name: recommender, namespace: ml-production}spec: scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: recommender} minReplicas: 4 maxReplicas: 30 metrics: - type: Resource resource: name: cpu target: {type: Utilization, averageUtilization: 60} - type: Pods pods: metric: {name: inference_queue_depth} target: {type: AverageValue, averageValue: "5"} behavior: scaleUp: stabilizationWindowSeconds: 0 policies: [{type: Percent, value: 100, periodSeconds: 30}] scaleDown: stabilizationWindowSeconds: 300 policies: [{type: Percent, value: 25, periodSeconds: 60}]The scaling formula is simply desired = ceil(current × currentMetric / targetMetric). With 12 pods averaging 85% CPU against a 60% target: 12 × 85 / 60 = 17.0, so it scales to 17 pods. If utilisation then settles at 45%, the next evaluation gives 17 × 45 / 60 = 12.75, which rounds up to 13 — note that scaling down also rounds up, which is why the count drifts down slowly rather than overshooting. Had utilisation settled at 55% instead, nothing would happen at all: 55 / 60 is within 10% of the target, and the autoscaler ignores ratios inside that default tolerance so that small wobbles do not cause flapping.
The asymmetric behavior block is what makes autoscaling usable in practice. Scale-up has no stabilisation window and may double the replica count every 30 seconds, so a traffic spike is absorbed quickly. Scale-down waits 300 seconds and removes at most 25% per minute, so a brief lull does not tear down capacity you need again two minutes later. Get this backwards and you get flapping: pods start, load a 3 GB model for 90 seconds, become ready, and are terminated before serving a single request.
The second metric matters for ML specifically. CPU utilisation is a poor proxy for a GPU-bound model server — the GPU can be saturated while CPU sits at 20%. Scaling on a queue-depth metric exported by the application itself tracks the thing users actually feel.
Driving deployments from Python
A promotion pipeline needs to change the running deployment and wait for the result, not fire and forget.
1from kubernetes import client, config, watch23try:4 config.load_incluster_config() # running as a pod5except config.ConfigException:6 config.load_kube_config() # running on a laptop78apps = client.AppsV1Api()910def set_image(name: str, ns: str, image_digest: str) -> None:11 apps.patch_namespaced_deployment(12 name=name, namespace=ns,13 body={"spec": {"template": {"spec": {"containers": [14 {"name": "server", "image": image_digest}]}}}},15 )1617def wait_for_rollout(name: str, ns: str, timeout: int = 600) -> bool:18 """True if all replicas became ready and updated within the timeout."""19 w = watch.Watch()20 for event in w.stream(apps.list_namespaced_deployment, namespace=ns,21 field_selector=f"metadata.name={name}",22 timeout_seconds=timeout):23 s = event["object"].status24 spec_replicas = event["object"].spec.replicas25 if (s.updated_replicas == spec_replicas26 and s.ready_replicas == spec_replicas27 and s.unavailable_replicas in (None, 0)28 and s.observed_generation >= event["object"].metadata.generation):29 w.stop()30 return True31 return False3233def rollback(name: str, ns: str) -> None:34 """Undo by re-applying the image of the most recent scaled-down ReplicaSet."""35 rs_list = apps.list_namespaced_replica_set(36 ns, label_selector=f"app={name}").items37 previous = sorted(38 [r for r in rs_list if r.spec.replicas == 0],39 key=lambda r: int(r.metadata.annotations40 .get("deployment.kubernetes.io/revision", 0)),41 )[-1]42 set_image(name, ns,43 previous.spec.template.spec.containers[0].image)The observed_generation >= metadata.generation check is the detail that stops a subtle race. Immediately after a patch, the deployment's status still describes the previous rollout, which was healthy — so a naive readiness check returns true instantly and your pipeline concludes a broken deploy succeeded. Comparing generations proves the controller has actually seen your change.
Kubernetes-native ML: KServe and the Training Operator
Everything above works, and it is a lot of YAML to write for every model. Two custom resource types collapse most of it.
KServe: serving as a single object
apiVersion: serving.kserve.io/v1beta1kind: InferenceServicemetadata: name: recommender namespace: ml-productionspec: predictor: minReplicas: 1 maxReplicas: 20 scaleTarget: 40 # concurrent requests per replica scaleMetric: concurrency canaryTrafficPercent: 10 # set per component, on the predictor sklearn: storageUri: "s3://ml-models/recommender/v14/" resources: requests: {cpu: "1", memory: "4Gi"} limits: {cpu: "2", memory: "6Gi"} transformer: # preprocessing runs as its own container containers: - name: kserve-container image: ghcr.io/acme/recommender-transformer@sha256:8b2e...That replaces a Deployment, a Service, an Ingress, an HPA and a router. Three things it adds that plain Deployments do not have:
- Concurrency-based scaling. Scaling on in-flight requests rather than CPU is far closer to what a model server actually experiences, particularly on GPUs.
canaryTrafficPercent. ChangingstorageUriand setting this to 10 sends a tenth of traffic to the new model. Setting it to 0 is an instant rollback with no pod churn.- Scale-to-zero. With
minReplicas: 0, an idle model costs nothing until a request arrives. The cold start is the price — expect 20–60 seconds to schedule a pod, pull an image and load weights. Excellent for a dozen rarely-used internal models; wrong behind a user-facing form.
| Plain Deployment + Service + HPA | KServe InferenceService | |
|---|---|---|
| YAML per model | ~120 lines across 4 objects | ~20 lines in 1 object |
| Canary | Build it yourself (two deployments, weighted routing) | One field |
| Scale to zero | Not supported | Supported |
| Prerequisites | Just Kubernetes | KServe, and usually Knative and a service mesh |
| Debuggability | Everything is standard and familiar | More layers between a request and your code |
The Training Operator: distributed training as an object
The manifest below is the Training Operator's v1 PyTorchJob, which still runs on clusters that have it installed and which most existing examples use. Kubeflow has since released Trainer v2, which replaces the per-framework job types with a single TrainJob resource plus reusable training runtimes, and puts v1 in maintenance mode. For a new cluster, start with Trainer v2; the ideas below — one object describing a whole distributed job, with the rendezvous wired up for you — carry over directly.
apiVersion: kubeflow.org/v1kind: PyTorchJobmetadata: {name: rec-train-2024-04-15, namespace: ml-training}spec: runPolicy: cleanPodPolicy: Running ttlSecondsAfterFinished: 86400 pytorchReplicaSpecs: Master: replicas: 1 restartPolicy: OnFailure template: spec: containers: - name: pytorch image: ghcr.io/acme/rec-train@sha256:9e1a... command: ["torchrun", "--nproc_per_node=1", "train.py", "--epochs", "10"] resources: limits: {nvidia.com/gpu: 1, memory: "32Gi"} Worker: replicas: 3 restartPolicy: OnFailure template: spec: containers: - name: pytorch image: ghcr.io/acme/rec-train@sha256:9e1a... command: ["torchrun", "--nproc_per_node=1", "train.py", "--epochs", "10"] resources: limits: {nvidia.com/gpu: 1, memory: "32Gi"}The operator injects MASTER_ADDR, MASTER_PORT, WORLD_SIZE and RANK into every pod, so your script calls torch.distributed.init_process_group() and the rendezvous just works. Doing this by hand means allocating stable network identities and coordinating start-up ordering yourself.
cleanPodPolicy: Running is the field that protects your GPUs: when the job ends, any pod still running — a worker hung in a collective after the master exited, say — is deleted instead of sitting on its GPU indefinitely. ttlSecondsAfterFinished is the one people forget: it garbage-collects the finished job and its pods after a day, so thousands of completed jobs do not pile up in the namespace.
Observability: the four numbers that matter
A model server needs standard service telemetry plus one thing no ordinary service has — the distribution of its own predictions.
1from prometheus_client import Counter, Histogram, Gauge, generate_latest2from fastapi import FastAPI, Response34REQUESTS = Counter("inference_requests_total", "Requests",5 ["model_version", "status"])6LATENCY = Histogram("inference_latency_seconds", "End-to-end latency",7 ["model_version"],8 buckets=(.005, .01, .025, .05, .1, .25, .5, 1, 2.5, 5))9SCORE = Histogram("prediction_score", "Predicted probability",10 buckets=[i / 20 for i in range(21)])11QUEUE = Gauge("inference_queue_depth", "Requests waiting for a worker")12MODEL_VERSION = Gauge("model_version_info", "Live model version", ["alias"])1314app = FastAPI()1516@app.post("/predict")17def predict(req: PredictRequest): # plain def: the model call blocks18 QUEUE.inc()19 try:20 with LATENCY.labels(model.version).time():21 score = model.predict_proba(req.to_frame())[0, 1]22 SCORE.observe(float(score))23 REQUESTS.labels(model.version, "ok").inc()24 return {"score": float(score), "model_version": model.version}25 except Exception:26 REQUESTS.labels(model.version, "error").inc()27 raise28 finally:29 QUEUE.dec()3031@app.get("/metrics")32def metrics():33 return Response(generate_latest(), media_type="text/plain")SCORE is the ML-specific metric and the one that catches silent failures. Every operational signal can look perfect — zero errors, 40 ms p99, healthy pods — while the mean predicted probability slides from 0.11 to 0.29 because an upstream feature changed units. A histogram of scores turns that into a visible, alertable shift within minutes instead of within however long labels take to arrive.
Every operational dashboard can be green while the model is wrong. The distribution of its own predictions is the only cheap signal that notices.
The model_version label on every metric is what lets you attribute a change to a deployment. Without it, a latency regression during a canary is indistinguishable from a general slowdown.
apiVersion: monitoring.coreos.com/v1kind: PrometheusRulemetadata: {name: recommender-alerts, namespace: ml-production}spec: groups: - name: recommender rules: - alert: InferenceErrorRateHigh expr: | sum(rate(inference_requests_total{status="error"}[5m])) / sum(rate(inference_requests_total[5m])) > 0.02 for: 5m labels: {severity: page} - alert: InferenceLatencyP99High expr: | histogram_quantile(0.99, sum by (le) (rate(inference_latency_seconds_bucket[5m]))) > 1.0 for: 10m labels: {severity: page} - alert: PredictionDistributionShift expr: | abs( sum(rate(prediction_score_sum[1h])) / sum(rate(prediction_score_count[1h])) - 0.11 ) > 0.05 for: 30m labels: {severity: ticket}Note the severities. Error rate and latency page a human; a distribution shift opens a ticket. Distribution shift is real and important, and it is almost never something you fix at 03:00 — paging on it trains people to ignore pages.
Closing the loop with GitOps
Manifests in a repository are only useful if something applies them. ArgoCD runs inside the cluster, watches a repository path, and reconciles continuously.
apiVersion: argoproj.io/v1alpha1kind: Applicationmetadata: {name: recommender, namespace: argocd}spec: project: ml source: repoURL: https://github.com/acme/ml-platform targetRevision: main path: deploy/production/recommender destination: server: https://kubernetes.default.svc namespace: ml-production syncPolicy: automated: {prune: true, selfHeal: true} syncOptions: [CreateNamespace=true] retry: limit: 3 backoff: {duration: 30s, factor: 2, maxDuration: 5m}selfHeal: true means a manual kubectl edit is reverted within a minute. That is the intended behaviour, not a bug: it guarantees the repository describes reality, so the answer to "what is running in production" is a file, not an investigation. prune: true means deleting a manifest from Git deletes the object from the cluster — powerful and worth enabling only once you trust your review process, because a bad merge now deletes a Service.
The security consequence is the underrated one. With this in place, your CI system never needs cluster credentials. It builds images and opens pull requests; the agent inside the cluster does the deploying. A compromised CI pipeline can push a container image, but it cannot deploy one without a merged, reviewed commit.
Deciding whether you need any of this
Kubernetes has a real cost. Somebody must upgrade the cluster, manage node pools, understand why a pod is Pending, and be woken by it. If you run three models with predictable traffic, a managed serving product or a couple of containers behind a load balancer is the correct engineering decision, and choosing Kubernetes instead means you now have two problems.
The threshold is usually crossed by one of three things: many models (a dozen services with the same shape, where a template beats a bespoke deployment each time), heterogeneous compute (some steps want a GPU for twenty minutes and others want 32 CPUs for two hours, and bin-packing that onto shared nodes is genuinely valuable), or an existing cluster that another team already operates, in which case most of the cost is already paid.
If you do cross it, the order of adoption matters. Start with plain Deployments, Services and HPAs — they are standard, well documented, and every engineer can debug them. Add Prometheus and the prediction-distribution histogram before anything clever, because you cannot make good scaling or rollout decisions without telemetry. Add GitOps next, since it is what turns the cluster from a mutable pet into a reviewable artefact. Reach for KServe and the Training Operator last, when the YAML repetition is genuinely painful — they are excellent, and every layer they add is another thing standing between a failing request and the log line that explains it.