Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Differentiate prefill vs decode phase in LLM inference and what dominates cost in each?
What you need to know
Arithmetic intensity decides the bottleneck
Arithmetic intensity is FLOPs done per byte read from memory. A GPU has a "ridge point": above it, compute is the limit; below it, memory bandwidth is.
H100 SXM: about 989 TFLOP/s (dense bf16) / 3.35 TB/s ≈ 295 FLOPs per byteeach weight (2 bytes in bf16) does 2 FLOPs per token it is applied tointensity ≈ tokens processed per weight read ≈ batch of tokens in this pass- Prefill of a 2,000-token prompt uses each weight 2,000 times: intensity about 2,000, far above 295. Compute-bound.
- Decode at batch size 1 uses each weight once: intensity about 1. The tensor cores sit idle more than 99% of the time, waiting for memory.
Put numbers on it (8B model, H100)
Prefill, 2,000 tokens: 2 × 8e9 × 2,000 = 3.2e13 FLOPs at ~50% of 989 TFLOP/s → about 65 ms of TTFTDecode, one token: read 16 GB of weights at 3.35 TB/s → about 4.8 ms per step, at least → at most ~200 tokens/s for one userWith a batch of 64 users, one decode step still reads the 16 GB once (plus each user's cache), so the step takes only a little longer while producing 64 tokens. That is why batching is almost free throughput in decode.
| Prefill | Decode | |
|---|---|---|
| Tokens per pass | all prompt tokens | 1 per sequence |
| Matrix shape | large matrix × matrix | thin matrix × vector |
| Bottleneck | compute (FLOPs) | memory bandwidth |
| Latency metric | TTFT | TPOT |
| Grows with | prompt length (attention grows as T²) | model size + cache size |
Techniques that follow from this
- Continuous batching (Orca, Yu et al., 2022) — add and remove requests from the running batch at every step, instead of waiting for a whole batch to finish.
- Chunked prefill — split a long prompt into pieces and mix them into decode batches, so one 50,000-token prompt does not freeze everyone else's stream.
- Disaggregated serving (DistServe, Splitwise, 2024) — run prefill and decode on separate GPU pools and ship the KV cache between them.
- Speculative decoding (Leviathan et al., 2023; Chen et al., 2023) — a small draft model proposes several tokens and the big model checks them all in one pass, turning several bandwidth-bound steps into one.
- Weight quantisation — 4-bit weights mean a quarter of the bytes per decode step, so decode speeds up; prefill gains only if the hardware also computes in low precision.
A real-life example
A code-completion assistant in an IDE sends about 3,000 tokens of surrounding code and asks for about 30 tokens back. Almost all the latency is prefill: users expect a suggestion within roughly 300 ms of pausing. The team keeps the model small, and uses prefix caching so that between two keystrokes only the few changed tokens at the end are prefilled again.
A support chatbot is the opposite: a 300-token question and a 600-token answer. Here decode dominates, the user watches tokens stream, and the team's levers are batching, quantisation and speculative decoding. Same model family, opposite tuning.
Follow-up questions to expect
- "Why does TTFT grow with prompt length?" — Prefill work grows linearly for the weight matmuls and quadratically for attention, so a 100k-token prompt can take seconds before the first token.
- "How does speculative decoding keep the output identical?" — The big model verifies each draft token with a rejection-sampling rule, so the output distribution is exactly the big model's; only speed changes.
- "Does int4 quantisation speed up prefill?" — Mostly not; prefill is compute-bound and int4 weights are usually de-quantised to 16-bit for the math. It mainly speeds up decode.