LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you scale an LLM system to handle high traffic loads?


Sizing the Diwali sale with Little's law26.7requestsper secondtimes 7.8 s each= 208 in flightdivided by40 perreplica = 5.2plus 30%headroom =7 replicaspre-scaleat 7:30 pm,KEDA aboveNew replicas take about five minutes, so the sale email cannot be the trigger.
Autoscale on in-flight requests, not GPU utilisation, and pre-scale for any spike you can see coming.

What you need to know

Layer 1: reduce the work

Every cached answer, trimmed prompt or request routed to a small model is capacity you do not have to buy. Offline jobs (nightly summaries, backfills) go to batch APIs or a separate low-priority pool.

Layer 2: use each GPU fully

  • Continuous batching — vLLM, SGLang and TensorRT-LLM add new requests to the running batch at every decode step, so the GPU never waits for the slowest sequence.
  • Prefix caching — shared system prompts are computed once.
  • Tuning — --max-num-seqs (batch size cap), --max-num-batched-tokens (work per step) and --gpu-memory-utilization trade latency for throughput.

Layer 3: replicas and smart routing

Keep servers stateless, one replica per GPU (or per tensor-parallel group). Route by least outstanding requests, not round-robin — LLM requests vary 100× in cost. Better still, prefix-aware routing sends requests with the same system prompt or conversation to the replica that already holds that KV cache (supported by projects such as llm-d and the Kubernetes Gateway API Inference Extension).

Layer 4: autoscale on the right signal

GPU utilisation sits near 100% under any load, so it cannot tell you when to add capacity. Scale on in-flight requests per replica or queue depth. With KEDA on Kubernetes:

YAML
apiVersion: keda.sh/v1alpha1kind: ScaledObjectmetadata:  name: support-bot-vllmspec:  scaleTargetRef:    name: vllm-support-bot            # the vLLM Deployment  minReplicaCount: 2  maxReplicaCount: 16  advanced:    horizontalPodAutoscalerConfig:      behavior:        scaleDown:          stabilizationWindowSeconds: 600   # wait 10 min before removing GPUs  triggers:    - type: prometheus      metadata:        serverAddress: http://prometheus.monitoring:9090        query: sum(vllm:num_requests_running{model_name="support-8b"}) + sum(vllm:num_requests_waiting{model_name="support-8b"})        threshold: "40"                 # target in-flight requests per replica

KEDA sets the replica count to roughly total in-flight requests divided by 40. The long scale-down window stops it removing GPUs during a short dip and then paying a cold start again minutes later.

Sizing with Little's law

For example, 10 requests per second that each take 6 seconds means 60 requests in flight. If load tests show one replica holds 30 in flight while meeting the TPOT target, you need 2 replicas, plus headroom. Get the per-replica number from a load test with your real prompt and output lengths, not from a benchmark.

Cold starts

A new replica needs a GPU node, an image pull and a model load — often several minutes for a multi-gigabyte model. So scale early, keep headroom, pre-pull images, keep weights on fast local disk, and pre-scale before known events.

A real-life example

A fintech support bot normally peaks at about 5.3 requests per second. Each request takes about 7.8 s (0.6 s TTFT plus 400 tokens at 18 ms). Little's law: 5.3 × 7.8 ≈ 42 requests in flight. At the chosen 40 per replica (the level that keeps TPOT under 20 ms), that is about 1 replica; they run 2 for redundancy.

For the Diwali sale, marketing expects 5× traffic: 26.7 requests per second × 7.8 s ≈ 208 in flight, which is 5.2 replicas. With 30% headroom, 7 replicas. Because new replicas take about 5 minutes to become ready, the team raises minReplicaCount to 7 at 7:30 pm, before the 8 pm sale email goes out, and lets KEDA handle anything above that up to 16. They also turn on an exact-match cache for the ten most common sale questions, which serves 12% of traffic without touching a GPU.

Follow-up questions to expect

  • "Why not autoscale on GPU utilisation?" — It reads near 100% whenever there is any batch running, so it does not show how much demand is waiting.
  • "How do you handle scale-to-zero?" — Fine for internal or dev models where a minutes-long first request is acceptable; not for user-facing traffic.
  • "What about managed APIs — do you still scale?" — You scale quotas instead: tokens-per-minute limits, provisioned throughput, and multiple deployments behind the gateway.