LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What causes latency spikes, and how can you stabilize performance?


What you need to know

The causes, from most to least common

  • Queueing. When arrival rate passes capacity, the queue grows without limit and wait time dominates TTFT.
  • Long prompts. Prefill is heavy and grows with input length. One pasted 80,000-token log file can delay other requests in the same batch.
  • KV-cache exhaustion. When the cache is full, vLLM preempts requests — pauses them, frees their memory, and later recomputes them. Throughput collapses right at the edge of capacity.
  • Cold starts after scale-out or a crash.
  • Retry storms. Clients retry timeouts immediately; load doubles on an already saturated system.
  • Slow dependencies. Vector search, reranker, or a tool call with no timeout.
  • Provider variance on managed APIs, especially at peak hours.

Why the curve is a cliff, not a slope

Queueing theory says wait time grows sharply as utilisation approaches 100%. At 70% load, the queue is small; at 95%, small bursts cause long waits; above 100%, it grows forever. So a system that is "fine at 90%" can spike from a tiny extra load.

Stabilising

  1. Cap inputs and outputs — per-route max_tokens and maximum prompt length (--max-model-len); reject or summarise oversized input at the edge.
  2. Keep prefill from blocking decode — chunked prefill (on by default in current vLLM) mixes long prefills with other users' decode steps.
  3. Admission control — bounded queue; when the expected wait passes the SLO, return 429 with Retry-After at once.
  4. Provision for p95 demand, with 20–30% headroom, and scale on queue signals.
  5. Separate pools for long-context and batch work.
  6. Timeouts and retry budgets everywhere, with jitter.

On managed APIs you cannot change the server, but you can cap tokens, set client timeouts, use hedged requests (send a second request if the first is slow past p95, cancel the loser) for short critical calls, and spread load across deployments.

A real-life example

A code assistant for 2,000 engineers has a p95 TTFT of 0.9 s. Every day around 11 am, it spikes to 12 s for a few minutes.

The dashboard shows vllm:kv_cache_usage_perc hitting 100% and the preemption counter jumping at the same times. Traces show the cause: engineers pasting whole CI logs — 60,000 to 90,000 tokens — into the chat after the morning build. At 128 KB of KV cache per token on the 8B model, one 80,000-token request takes about 10 GB, a fifth of the replica's cache. Three of them at once push out dozens of normal requests.

Fixes: the IDE plugin now sends only the last 300 lines of a log plus lines matching "error" (usually under 6,000 tokens); requests over 32,000 tokens go to a separate "long" pool with its own replica; and an alert fires when KV-cache usage stays above 90% for two minutes. The 11 am spike disappears, and p99 falls from 12 s to 2.1 s.

Follow-up questions to expect

  • "What is preemption in vLLM?" — When KV-cache blocks run out, the scheduler pauses some running requests and frees their blocks; they are recomputed later, which wastes work and adds latency.
  • "How would you find the cause of a spike?" — Split TTFT from TPOT: TTFT spikes mean queueing or prefill; TPOT spikes mean too large a batch or memory pressure. Then check queue depth, KV usage and prompt-length distribution at that time.
  • "What are hedged requests?" — Send a duplicate request after a delay and use whichever returns first; it cuts tail latency but costs extra, so use it only for short, important calls.