Course Content
LLMOps & Deployment
6 sections · 40 lessons
What causes latency spikes, and how can you stabilize performance?
What you need to know
The causes, from most to least common
- Queueing. When arrival rate passes capacity, the queue grows without limit and wait time dominates TTFT.
- Long prompts. Prefill is heavy and grows with input length. One pasted 80,000-token log file can delay other requests in the same batch.
- KV-cache exhaustion. When the cache is full, vLLM preempts requests — pauses them, frees their memory, and later recomputes them. Throughput collapses right at the edge of capacity.
- Cold starts after scale-out or a crash.
- Retry storms. Clients retry timeouts immediately; load doubles on an already saturated system.
- Slow dependencies. Vector search, reranker, or a tool call with no timeout.
- Provider variance on managed APIs, especially at peak hours.
Why the curve is a cliff, not a slope
Queueing theory says wait time grows sharply as utilisation approaches 100%. At 70% load, the queue is small; at 95%, small bursts cause long waits; above 100%, it grows forever. So a system that is "fine at 90%" can spike from a tiny extra load.
Stabilising
- Cap inputs and outputs — per-route
max_tokensand maximum prompt length (--max-model-len); reject or summarise oversized input at the edge. - Keep prefill from blocking decode — chunked prefill (on by default in current vLLM) mixes long prefills with other users' decode steps.
- Admission control — bounded queue; when the expected wait passes the SLO, return 429 with
Retry-Afterat once. - Provision for p95 demand, with 20–30% headroom, and scale on queue signals.
- Separate pools for long-context and batch work.
- Timeouts and retry budgets everywhere, with jitter.
On managed APIs you cannot change the server, but you can cap tokens, set client timeouts, use hedged requests (send a second request if the first is slow past p95, cancel the loser) for short critical calls, and spread load across deployments.
A real-life example
A code assistant for 2,000 engineers has a p95 TTFT of 0.9 s. Every day around 11 am, it spikes to 12 s for a few minutes.
The dashboard shows vllm:kv_cache_usage_perc hitting 100% and the preemption counter jumping at the same times. Traces show the cause: engineers pasting whole CI logs — 60,000 to 90,000 tokens — into the chat after the morning build. At 128 KB of KV cache per token on the 8B model, one 80,000-token request takes about 10 GB, a fifth of the replica's cache. Three of them at once push out dozens of normal requests.
Fixes: the IDE plugin now sends only the last 300 lines of a log plus lines matching "error" (usually under 6,000 tokens); requests over 32,000 tokens go to a separate "long" pool with its own replica; and an alert fires when KV-cache usage stays above 90% for two minutes. The 11 am spike disappears, and p99 falls from 12 s to 2.1 s.
Follow-up questions to expect
- "What is preemption in vLLM?" — When KV-cache blocks run out, the scheduler pauses some running requests and frees their blocks; they are recomputed later, which wastes work and adds latency.
- "How would you find the cause of a spike?" — Split TTFT from TPOT: TTFT spikes mean queueing or prefill; TPOT spikes mean too large a batch or memory pressure. Then check queue depth, KV usage and prompt-length distribution at that time.
- "What are hedged requests?" — Send a duplicate request after a delay and use whichever returns first; it cuts tail latency but costs extra, so use it only for short, important calls.