Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you scale an LLM system to handle high traffic loads?
What you need to know
Layer 1: reduce the work
Every cached answer, trimmed prompt or request routed to a small model is capacity you do not have to buy. Offline jobs (nightly summaries, backfills) go to batch APIs or a separate low-priority pool.
Layer 2: use each GPU fully
- Continuous batching — vLLM, SGLang and TensorRT-LLM add new requests to the running batch at every decode step, so the GPU never waits for the slowest sequence.
- Prefix caching — shared system prompts are computed once.
- Tuning —
--max-num-seqs(batch size cap),--max-num-batched-tokens(work per step) and--gpu-memory-utilizationtrade latency for throughput.
Layer 3: replicas and smart routing
Keep servers stateless, one replica per GPU (or per tensor-parallel group). Route by least outstanding requests, not round-robin — LLM requests vary 100× in cost. Better still, prefix-aware routing sends requests with the same system prompt or conversation to the replica that already holds that KV cache (supported by projects such as llm-d and the Kubernetes Gateway API Inference Extension).
Layer 4: autoscale on the right signal
GPU utilisation sits near 100% under any load, so it cannot tell you when to add capacity. Scale on in-flight requests per replica or queue depth. With KEDA on Kubernetes:
1apiVersion: keda.sh/v1alpha12kind: ScaledObject3metadata:4 name: support-bot-vllm5spec:6 scaleTargetRef:7 name: vllm-support-bot # the vLLM Deployment8 minReplicaCount: 29 maxReplicaCount: 1610 advanced:11 horizontalPodAutoscalerConfig:12 behavior:13 scaleDown:14 stabilizationWindowSeconds: 600 # wait 10 min before removing GPUs15 triggers:16 - type: prometheus17 metadata:18 serverAddress: http://prometheus.monitoring:909019 query: sum(vllm:num_requests_running{model_name="support-8b"}) + sum(vllm:num_requests_waiting{model_name="support-8b"})20 threshold: "40" # target in-flight requests per replicaKEDA sets the replica count to roughly total in-flight requests divided by 40. The long scale-down window stops it removing GPUs during a short dip and then paying a cold start again minutes later.
Sizing with Little's law
For example, 10 requests per second that each take 6 seconds means 60 requests in flight. If load tests show one replica holds 30 in flight while meeting the TPOT target, you need 2 replicas, plus headroom. Get the per-replica number from a load test with your real prompt and output lengths, not from a benchmark.
Cold starts
A new replica needs a GPU node, an image pull and a model load — often several minutes for a multi-gigabyte model. So scale early, keep headroom, pre-pull images, keep weights on fast local disk, and pre-scale before known events.
A real-life example
A fintech support bot normally peaks at about 5.3 requests per second. Each request takes about 7.8 s (0.6 s TTFT plus 400 tokens at 18 ms). Little's law: 5.3 × 7.8 ≈ 42 requests in flight. At the chosen 40 per replica (the level that keeps TPOT under 20 ms), that is about 1 replica; they run 2 for redundancy.
For the Diwali sale, marketing expects 5× traffic: 26.7 requests per second × 7.8 s ≈ 208 in flight, which is 5.2 replicas. With 30% headroom, 7 replicas. Because new replicas take about 5 minutes to become ready, the team raises minReplicaCount to 7 at 7:30 pm, before the 8 pm sale email goes out, and lets KEDA handle anything above that up to 16. They also turn on an exact-match cache for the ten most common sale questions, which serves 12% of traffic without touching a GPU.
Follow-up questions to expect
- "Why not autoscale on GPU utilisation?" — It reads near 100% whenever there is any batch running, so it does not show how much demand is waiting.
- "How do you handle scale-to-zero?" — Fine for internal or dev models where a minutes-long first request is acceptable; not for user-facing traffic.
- "What about managed APIs — do you still scale?" — You scale quotas instead: tokens-per-minute limits, provisioned throughput, and multiple deployments behind the gateway.