LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What challenges arise when managing GPUs for LLM workloads?


Where 80 GB goes: an 8B model on one H100Weights,bf16: 16 GBRuntime andCUDA graphs: ~4 GBKV cache: ~52 GBLeft free at0.90: 8 GBtopbottom52 GB at 128 KB per token is about 425,000 tokens — roughly 100 requests of 4,000 tokens.
The weights are the small part; the KV cache sets how many users one GPU can serve.

What you need to know

1. Memory sizing

Text
GPU memory = weights + KV cache + activations/runtime + headroomKV bytes per token = 2 (K and V) × layers × kv_heads × head_dim × bytes_per_valueLlama 3.1 8B,  bf16: 2 × 32 × 8 × 128 × 2 = 131,072 B ≈ 128 KBLlama 3.1 70B, bf16: 2 × 80 × 8 × 128 × 2 = 327,680 B ≈ 320 KB

Worked example: an 8B model on one 80 GB H100 with --gpu-memory-utilization 0.90:

ItemGB
Left free (10%)8
Weights (bf16)16
Activations, CUDA graphs, runtime~4
KV cache~52

52 GB ÷ 128 KB ≈ 425,000 tokens of context in flight — about 100 requests of 4,000 tokens, or 52 of 8,000. Concurrency is limited by memory, not compute. FP8 weights or an FP8 KV cache raise it further.

2. Utilisation and cost

A GPU bills 24 hours a day. One request at a time uses a small fraction of it, because decode is limited by memory bandwidth. Continuous batching, prefix caching and routing enough traffic to each GPU are what make the economics work. Track tokens per GPU-hour, not only GPU utilisation percentage (which reads high even when batches are small).

3. Cold starts

A new replica needs a node, a multi-gigabyte container image, and 16–140 GB of weights loaded into GPU memory. That is minutes. Pre-pull images, keep weights on local NVMe or a fast cache, keep warm spare capacity, and pre-scale for known peaks.

4. Sharing and packing

  • A model larger than one GPU (70B in bf16 is ~140 GB) needs tensor parallelism across a matched pair or group of GPUs on the same fast interconnect, so the scheduling unit becomes "two H100s on one node".
  • Small models waste big GPUs. MIG (Multi-Instance GPU) on A100/H100-class GPUs splits one GPU into up to seven isolated slices; time-slicing shares a GPU without isolation.

5. Fleet operations

Driver and CUDA version compatibility with the inference engine; the NVIDIA GPU Operator on Kubernetes; DCGM metrics for temperature, clocks, memory errors (ECC) and XID errors; hardware failures that take a node out; and quotas and reservations, because capacity for the newest GPUs is limited.

A real-life example

A hospital runs its AI workloads on-prem on four L40S GPUs (48 GB each): the discharge summariser, a document embedding model for search, and an OCR model for scanned forms.

The first design gave each model a whole GPU and kept one spare. Monitoring showed the embedding model used 3 GB and about 5% of its GPU, while the summariser queued at peak. The redesign:

  • The summariser (70B in 4-bit, tensor-parallel) gets two GPUs as one unit, with ~40 GB of KV cache — about 130,000 tokens in flight.
  • Embedding and OCR share the third GPU (they fit easily in memory, and both are batch-friendly).
  • The fourth GPU runs an 8B copy of the summariser as a warm standby and handles overflow.

When a GPU reports repeated ECC memory errors in DCGM, the node is drained, the gateway sends all summaries to the 8B standby for two days, and the vendor replaces the card.

Follow-up questions to expect

  • "What limits concurrency on a GPU?" — KV-cache memory, then the latency target (bigger batches make each token slower).
  • "How do you reduce KV-cache size?" — Models with grouped-query attention (fewer KV heads), FP8 KV cache, shorter prompts and capped outputs, and prefix sharing.
  • "How would you share one GPU between small models?" — MIG for isolation and predictable performance; time-slicing or running several models in one server when isolation matters less.