MLOps for AI

Container Orchestration with Kubernetes


The recommendation service runs as one container on one virtual machine, started months ago with docker run -d. It has worked fine at 40 requests per second.

On the first day of a sale, traffic reaches 400 requests per second. The p99 latency goes from 90 ms to 14 seconds because a single Python process cannot serve ten times the load. Then a batch request arrives asking for recommendations for 5,000 users at once; the process allocates about 9 GB, the kernel's out-of-memory killer terminates it, and nothing restarts it. The service is down for 47 minutes, until a customer complains loudly enough to reach an engineer.

Every one of those failures has a name and a standard answer: no horizontal scaling, no restart policy, no memory bound, no health checking, no alert. Assembling those answers yourself — a supervisor, a load balancer, a scaler, a health checker, a deployer — is roughly what Kubernetes is. It is worth understanding it as a set of answers to specific failures rather than as a platform you adopt wholesale.

Five objects, and the failure each one preventsOne modelserver,run properlyDeployment —replicas and rolloutService — one stable addressHPA — replicasfrom a real metricProbes — startup,readiness, livenessRequests and limits— schedule vs cap
A single docker run has none of these, which is why it survives 40 requests a second and nothing above it.

The building blocks, and the failure each one prevents

Kubernetes works on one principle: you declare the state you want, and a control loop continuously works to make reality match. You never say "start a container". You say "four of these should exist", and something else notices when three do.

ObjectWhat it isFailure it answers
PodOne or more containers sharing a network namespace and lifecycleNone directly; it is the unit everything else manages
Deployment"Keep N healthy pods of this spec"; handles rolling updatesCrashed process never restarts; no way to update without downtime
ServiceStable virtual IP and DNS name that load-balances across matching podsPod IPs change on every restart; callers hardcode a dead address
IngressHTTP routing from outside the cluster, with TLS terminationNo public entry point; certificates managed by hand
ConfigMap / SecretConfiguration and credentials injected as env vars or filesSecrets baked into images; config changes need a rebuild
HorizontalPodAutoscalerAdjusts replica count from a metricThe sale-day capacity problem
PersistentVolumeClaimStorage that outlives a podModel cache re-downloaded on every restart
NamespaceA scope for names, quotas and access controlStaging and production sharing a blast radius

Nothing in Kubernetes is a command. Everything is a statement about the desired world, plus a controller whose whole job is to notice disagreement and fix it.

A manifest for a model server, line by line

Text
apiVersion: apps/v1kind: Deploymentmetadata:  name: recommender  namespace: ml-production  labels: {app: recommender}spec:  replicas: 4  strategy:    type: RollingUpdate    rollingUpdate: {maxSurge: 1, maxUnavailable: 0}  selector:    matchLabels: {app: recommender}  template:    metadata:      labels: {app: recommender}      annotations:        prometheus.io/scrape: "true"        prometheus.io/port: "8000"    spec:      securityContext:        runAsNonRoot: true        runAsUser: 1001      containers:      - name: server        image: ghcr.io/acme/recommender@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d        ports:        - {containerPort: 8000, name: http}        env:        - name: MODEL_ALIAS          value: champion        - name: MLFLOW_TRACKING_URI          valueFrom:            configMapKeyRef: {name: ml-config, key: tracking_uri}        - name: AWS_SECRET_ACCESS_KEY          valueFrom:            secretKeyRef: {name: model-store, key: secret_key}        resources:          requests: {cpu: "1000m", memory: "4Gi"}          limits:   {cpu: "2000m", memory: "6Gi"}        startupProbe:          httpGet: {path: /health, port: 8000}          failureThreshold: 30          periodSeconds: 5        readinessProbe:          httpGet: {path: /ready, port: 8000}          periodSeconds: 10          failureThreshold: 3        livenessProbe:          httpGet: {path: /health, port: 8000}          periodSeconds: 20          failureThreshold: 3        lifecycle:          preStop:            exec: {command: ["sleep", "10"]}      terminationGracePeriodSeconds: 60

Requests and limits are two different things

This is the single most misunderstood pair of numbers in the file.

requestslimits
Used byThe scheduler, to decide which node has roomThe kernel, to enforce a ceiling at runtime
Memory behaviourReserves capacity; the pod will not be placed without itExceeding it means an immediate OOMKill — no warning, no grace
CPU behaviourGuarantees a share under contentionThrottles: the process is stopped until the next scheduling period
Set too lowPods packed onto a node, then evicted under pressureConstant OOMKills or crippling throttling
Set too highNodes half-empty; you pay for reserved-but-unused capacityLittle harm for memory; none for CPU

Work the packing arithmetic. A node has 16 GiB of RAM, of which roughly 1.5 GiB goes to the kubelet, the container runtime and system daemons — about 14.5 GiB schedulable. With requests.memory: 4Gi, the scheduler fits ⌊14.5 / 4⌋ = 3 pods per node. Four replicas therefore need two nodes. Drop the request to 3Gi and ⌊14.5 / 3⌋ = 4 pods fit on one node — but if the model genuinely resides at 3.2 GiB, you have set a request below actual usage, and under memory pressure the kubelet evicts your pods before anything else because they exceeded their request.

CPU limits deserve their own warning because the failure is invisible. The Linux scheduler enforces a CPU limit as a quota per 100 ms period: limits.cpu: 2000m means 200 ms of CPU time per 100 ms of wall clock, spread across all threads. A model server whose inference library spawns 8 BLAS threads can burn that entire quota in 25 ms and then be frozen for the remaining 75 ms of the period. Latency graphs show mysterious 75 ms plateaus that no profiler explains. The fix is usually to cap the library's thread count to match the limit — OMP_NUM_THREADS=2 — rather than to raise the limit.

A memory limit is a promise the kernel makes to kill you; a memory request is a promise the scheduler makes to you. Confusing them yields either evictions or half-empty nodes.

Three probes, three jobs

ProbeQuestionOn failureML-specific reason it matters
startupProbeHas it finished booting?Kills the pod after the thresholdLoading a multi-gigabyte model can take 90 s; without this, the liveness probe kills it mid-load, forever
readinessProbeCan it take traffic now?Removes it from the Service endpoints; does not kill itLets a pod refreshing its model temporarily leave the rotation
livenessProbeIs it wedged and unrecoverable?Restarts the containerCatches a deadlocked inference thread that still accepts connections

The startupProbe here allows 30 × 5 = 150 seconds to load. Until it succeeds, the liveness and readiness probes are suspended entirely. This is the fix for the classic crash loop where a large model never finishes loading because the liveness probe kills it at 30 seconds, every time, indefinitely.

The preStop: sleep 10 paired with terminationGracePeriodSeconds: 60 solves a subtler problem. When a pod is deleted, two things happen concurrently: the container gets SIGTERM, and the endpoint is removed from the Service. Endpoint removal propagates through kube-proxy on every node and is not instant. Without the sleep, the process exits before some nodes have stopped routing to it, and a handful of requests get connection-refused during every deploy. Ten seconds of doing nothing before shutdown eliminates that entirely.

Finally, maxUnavailable: 0 with maxSurge: 1 means a new pod must be ready before an old one is removed. Capacity never dips during a rollout — at the cost of needing room for one extra pod.

Autoscaling with arithmetic you can predict

Text
apiVersion: autoscaling/v2kind: HorizontalPodAutoscalermetadata: {name: recommender, namespace: ml-production}spec:  scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: recommender}  minReplicas: 4  maxReplicas: 30  metrics:  - type: Resource    resource:      name: cpu      target: {type: Utilization, averageUtilization: 60}  - type: Pods    pods:      metric: {name: inference_queue_depth}      target: {type: AverageValue, averageValue: "5"}  behavior:    scaleUp:      stabilizationWindowSeconds: 0      policies: [{type: Percent, value: 100, periodSeconds: 30}]    scaleDown:      stabilizationWindowSeconds: 300      policies: [{type: Percent, value: 25, periodSeconds: 60}]

The scaling formula is simply desired = ceil(current × currentMetric / targetMetric). With 12 pods averaging 85% CPU against a 60% target: 12 × 85 / 60 = 17.0, so it scales to 17 pods. If utilisation then settles at 45%, the next evaluation gives 17 × 45 / 60 = 12.75, which rounds up to 13 — note that scaling down also rounds up, which is why the count drifts down slowly rather than overshooting. Had utilisation settled at 55% instead, nothing would happen at all: 55 / 60 is within 10% of the target, and the autoscaler ignores ratios inside that default tolerance so that small wobbles do not cause flapping.

The asymmetric behavior block is what makes autoscaling usable in practice. Scale-up has no stabilisation window and may double the replica count every 30 seconds, so a traffic spike is absorbed quickly. Scale-down waits 300 seconds and removes at most 25% per minute, so a brief lull does not tear down capacity you need again two minutes later. Get this backwards and you get flapping: pods start, load a 3 GB model for 90 seconds, become ready, and are terminated before serving a single request.

The second metric matters for ML specifically. CPU utilisation is a poor proxy for a GPU-bound model server — the GPU can be saturated while CPU sits at 20%. Scaling on a queue-depth metric exported by the application itself tracks the thing users actually feel.

Driving deployments from Python

A promotion pipeline needs to change the running deployment and wait for the result, not fire and forget.

Python
from kubernetes import client, config, watchtry:    config.load_incluster_config()      # running as a podexcept config.ConfigException:    config.load_kube_config()           # running on a laptopapps = client.AppsV1Api()def set_image(name: str, ns: str, image_digest: str) -> None:    apps.patch_namespaced_deployment(        name=name, namespace=ns,        body={"spec": {"template": {"spec": {"containers": [            {"name": "server", "image": image_digest}]}}}},    )def wait_for_rollout(name: str, ns: str, timeout: int = 600) -> bool:    """True if all replicas became ready and updated within the timeout."""    w = watch.Watch()    for event in w.stream(apps.list_namespaced_deployment, namespace=ns,                          field_selector=f"metadata.name={name}",                          timeout_seconds=timeout):        s = event["object"].status        spec_replicas = event["object"].spec.replicas        if (s.updated_replicas == spec_replicas                and s.ready_replicas == spec_replicas                and s.unavailable_replicas in (None, 0)                and s.observed_generation >= event["object"].metadata.generation):            w.stop()            return True    return Falsedef rollback(name: str, ns: str) -> None:    """Undo by re-applying the image of the most recent scaled-down ReplicaSet."""    rs_list = apps.list_namespaced_replica_set(        ns, label_selector=f"app={name}").items    previous = sorted(        [r for r in rs_list if r.spec.replicas == 0],        key=lambda r: int(r.metadata.annotations                          .get("deployment.kubernetes.io/revision", 0)),    )[-1]    set_image(name, ns,              previous.spec.template.spec.containers[0].image)

The observed_generation >= metadata.generation check is the detail that stops a subtle race. Immediately after a patch, the deployment's status still describes the previous rollout, which was healthy — so a naive readiness check returns true instantly and your pipeline concludes a broken deploy succeeded. Comparing generations proves the controller has actually seen your change.

Kubernetes-native ML: KServe and the Training Operator

Everything above works, and it is a lot of YAML to write for every model. Two custom resource types collapse most of it.

KServe: serving as a single object

Text
apiVersion: serving.kserve.io/v1beta1kind: InferenceServicemetadata:  name: recommender  namespace: ml-productionspec:  predictor:    minReplicas: 1    maxReplicas: 20    scaleTarget: 40                 # concurrent requests per replica    scaleMetric: concurrency    canaryTrafficPercent: 10        # set per component, on the predictor    sklearn:      storageUri: "s3://ml-models/recommender/v14/"      resources:        requests: {cpu: "1", memory: "4Gi"}        limits:   {cpu: "2", memory: "6Gi"}  transformer:                      # preprocessing runs as its own container    containers:    - name: kserve-container      image: ghcr.io/acme/recommender-transformer@sha256:8b2e...

That replaces a Deployment, a Service, an Ingress, an HPA and a router. Three things it adds that plain Deployments do not have:

  • Concurrency-based scaling. Scaling on in-flight requests rather than CPU is far closer to what a model server actually experiences, particularly on GPUs.
  • canaryTrafficPercent. Changing storageUri and setting this to 10 sends a tenth of traffic to the new model. Setting it to 0 is an instant rollback with no pod churn.
  • Scale-to-zero. With minReplicas: 0, an idle model costs nothing until a request arrives. The cold start is the price — expect 20–60 seconds to schedule a pod, pull an image and load weights. Excellent for a dozen rarely-used internal models; wrong behind a user-facing form.
Plain Deployment + Service + HPAKServe InferenceService
YAML per model~120 lines across 4 objects~20 lines in 1 object
CanaryBuild it yourself (two deployments, weighted routing)One field
Scale to zeroNot supportedSupported
PrerequisitesJust KubernetesKServe, and usually Knative and a service mesh
DebuggabilityEverything is standard and familiarMore layers between a request and your code

The Training Operator: distributed training as an object

The manifest below is the Training Operator's v1 PyTorchJob, which still runs on clusters that have it installed and which most existing examples use. Kubeflow has since released Trainer v2, which replaces the per-framework job types with a single TrainJob resource plus reusable training runtimes, and puts v1 in maintenance mode. For a new cluster, start with Trainer v2; the ideas below — one object describing a whole distributed job, with the rendezvous wired up for you — carry over directly.

Text
apiVersion: kubeflow.org/v1kind: PyTorchJobmetadata: {name: rec-train-2024-04-15, namespace: ml-training}spec:  runPolicy:    cleanPodPolicy: Running    ttlSecondsAfterFinished: 86400  pytorchReplicaSpecs:    Master:      replicas: 1      restartPolicy: OnFailure      template:        spec:          containers:          - name: pytorch            image: ghcr.io/acme/rec-train@sha256:9e1a...            command: ["torchrun", "--nproc_per_node=1", "train.py",                      "--epochs", "10"]            resources:              limits: {nvidia.com/gpu: 1, memory: "32Gi"}    Worker:      replicas: 3      restartPolicy: OnFailure      template:        spec:          containers:          - name: pytorch            image: ghcr.io/acme/rec-train@sha256:9e1a...            command: ["torchrun", "--nproc_per_node=1", "train.py",                      "--epochs", "10"]            resources:              limits: {nvidia.com/gpu: 1, memory: "32Gi"}

The operator injects MASTER_ADDR, MASTER_PORT, WORLD_SIZE and RANK into every pod, so your script calls torch.distributed.init_process_group() and the rendezvous just works. Doing this by hand means allocating stable network identities and coordinating start-up ordering yourself.

cleanPodPolicy: Running is the field that protects your GPUs: when the job ends, any pod still running — a worker hung in a collective after the master exited, say — is deleted instead of sitting on its GPU indefinitely. ttlSecondsAfterFinished is the one people forget: it garbage-collects the finished job and its pods after a day, so thousands of completed jobs do not pile up in the namespace.

Observability: the four numbers that matter

A model server needs standard service telemetry plus one thing no ordinary service has — the distribution of its own predictions.

Python
from prometheus_client import Counter, Histogram, Gauge, generate_latestfrom fastapi import FastAPI, ResponseREQUESTS = Counter("inference_requests_total", "Requests",                   ["model_version", "status"])LATENCY = Histogram("inference_latency_seconds", "End-to-end latency",                    ["model_version"],                    buckets=(.005, .01, .025, .05, .1, .25, .5, 1, 2.5, 5))SCORE = Histogram("prediction_score", "Predicted probability",                  buckets=[i / 20 for i in range(21)])QUEUE = Gauge("inference_queue_depth", "Requests waiting for a worker")MODEL_VERSION = Gauge("model_version_info", "Live model version", ["alias"])app = FastAPI()@app.post("/predict")def predict(req: PredictRequest):          # plain def: the model call blocks    QUEUE.inc()    try:        with LATENCY.labels(model.version).time():            score = model.predict_proba(req.to_frame())[0, 1]        SCORE.observe(float(score))        REQUESTS.labels(model.version, "ok").inc()        return {"score": float(score), "model_version": model.version}    except Exception:        REQUESTS.labels(model.version, "error").inc()        raise    finally:        QUEUE.dec()@app.get("/metrics")def metrics():    return Response(generate_latest(), media_type="text/plain")

SCORE is the ML-specific metric and the one that catches silent failures. Every operational signal can look perfect — zero errors, 40 ms p99, healthy pods — while the mean predicted probability slides from 0.11 to 0.29 because an upstream feature changed units. A histogram of scores turns that into a visible, alertable shift within minutes instead of within however long labels take to arrive.

Every operational dashboard can be green while the model is wrong. The distribution of its own predictions is the only cheap signal that notices.

The model_version label on every metric is what lets you attribute a change to a deployment. Without it, a latency regression during a canary is indistinguishable from a general slowdown.

Text
apiVersion: monitoring.coreos.com/v1kind: PrometheusRulemetadata: {name: recommender-alerts, namespace: ml-production}spec:  groups:  - name: recommender    rules:    - alert: InferenceErrorRateHigh      expr: |        sum(rate(inference_requests_total{status="error"}[5m]))        / sum(rate(inference_requests_total[5m])) > 0.02      for: 5m      labels: {severity: page}    - alert: InferenceLatencyP99High      expr: |        histogram_quantile(0.99,          sum by (le) (rate(inference_latency_seconds_bucket[5m]))) > 1.0      for: 10m      labels: {severity: page}    - alert: PredictionDistributionShift      expr: |        abs(          sum(rate(prediction_score_sum[1h]))          / sum(rate(prediction_score_count[1h]))          - 0.11        ) > 0.05      for: 30m      labels: {severity: ticket}

Note the severities. Error rate and latency page a human; a distribution shift opens a ticket. Distribution shift is real and important, and it is almost never something you fix at 03:00 — paging on it trains people to ignore pages.

Closing the loop with GitOps

Manifests in a repository are only useful if something applies them. ArgoCD runs inside the cluster, watches a repository path, and reconciles continuously.

Text
apiVersion: argoproj.io/v1alpha1kind: Applicationmetadata: {name: recommender, namespace: argocd}spec:  project: ml  source:    repoURL: https://github.com/acme/ml-platform    targetRevision: main    path: deploy/production/recommender  destination:    server: https://kubernetes.default.svc    namespace: ml-production  syncPolicy:    automated: {prune: true, selfHeal: true}    syncOptions: [CreateNamespace=true]    retry:      limit: 3      backoff: {duration: 30s, factor: 2, maxDuration: 5m}

selfHeal: true means a manual kubectl edit is reverted within a minute. That is the intended behaviour, not a bug: it guarantees the repository describes reality, so the answer to "what is running in production" is a file, not an investigation. prune: true means deleting a manifest from Git deletes the object from the cluster — powerful and worth enabling only once you trust your review process, because a bad merge now deletes a Service.

The security consequence is the underrated one. With this in place, your CI system never needs cluster credentials. It builds images and opens pull requests; the agent inside the cluster does the deploying. A compromised CI pipeline can push a container image, but it cannot deploy one without a merged, reviewed commit.

Deciding whether you need any of this

Kubernetes has a real cost. Somebody must upgrade the cluster, manage node pools, understand why a pod is Pending, and be woken by it. If you run three models with predictable traffic, a managed serving product or a couple of containers behind a load balancer is the correct engineering decision, and choosing Kubernetes instead means you now have two problems.

The threshold is usually crossed by one of three things: many models (a dozen services with the same shape, where a template beats a bespoke deployment each time), heterogeneous compute (some steps want a GPU for twenty minutes and others want 32 CPUs for two hours, and bin-packing that onto shared nodes is genuinely valuable), or an existing cluster that another team already operates, in which case most of the cost is already paid.

If you do cross it, the order of adoption matters. Start with plain Deployments, Services and HPAs — they are standard, well documented, and every engineer can debug them. Add Prometheus and the prediction-distribution histogram before anything clever, because you cannot make good scaling or rollout decisions without telemetry. Add GitOps next, since it is what turns the cluster from a mutable pet into a reviewable artefact. Reach for KServe and the Training Operator last, when the YAML repetition is genuinely painful — they are excellent, and every layer they add is another thing standing between a failing request and the log line that explains it.