Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Differentiate prefill vs decode phase in LLM inference and what dominates cost in each?


Same model, opposite bottlenecksPrefill — the whole prompt• All prompt tokens in one pass• Each weight used about T times• Compute-bound: sets time to first token• 2,000 tokens on 8B: about 65 msDecode — one token per pass• One new token per sequence• Each weight used once per user• Bandwidth-bound: sets time per token• 16 GB read per step: about 5 ms
Decode leaves the tensor cores idle waiting on memory, so batching many users into one step is almost free throughput.

What you need to know

Arithmetic intensity decides the bottleneck

Arithmetic intensity is FLOPs done per byte read from memory. A GPU has a "ridge point": above it, compute is the limit; below it, memory bandwidth is.

Text
H100 SXM: about 989 TFLOP/s (dense bf16) / 3.35 TB/s ≈ 295 FLOPs per byteeach weight (2 bytes in bf16) does 2 FLOPs per token it is applied tointensity ≈ tokens processed per weight read ≈ batch of tokens in this pass
  • Prefill of a 2,000-token prompt uses each weight 2,000 times: intensity about 2,000, far above 295. Compute-bound.
  • Decode at batch size 1 uses each weight once: intensity about 1. The tensor cores sit idle more than 99% of the time, waiting for memory.

Put numbers on it (8B model, H100)

Text
Prefill, 2,000 tokens: 2 × 8e9 × 2,000 = 3.2e13 FLOPs  at ~50% of 989 TFLOP/s  → about 65 ms of TTFTDecode, one token: read 16 GB of weights at 3.35 TB/s  → about 4.8 ms per step, at least → at most ~200 tokens/s for one user

With a batch of 64 users, one decode step still reads the 16 GB once (plus each user's cache), so the step takes only a little longer while producing 64 tokens. That is why batching is almost free throughput in decode.

PrefillDecode
Tokens per passall prompt tokens1 per sequence
Matrix shapelarge matrix × matrixthin matrix × vector
Bottleneckcompute (FLOPs)memory bandwidth
Latency metricTTFTTPOT
Grows withprompt length (attention grows as T²)model size + cache size

Techniques that follow from this

  • Continuous batching (Orca, Yu et al., 2022) — add and remove requests from the running batch at every step, instead of waiting for a whole batch to finish.
  • Chunked prefill — split a long prompt into pieces and mix them into decode batches, so one 50,000-token prompt does not freeze everyone else's stream.
  • Disaggregated serving (DistServe, Splitwise, 2024) — run prefill and decode on separate GPU pools and ship the KV cache between them.
  • Speculative decoding (Leviathan et al., 2023; Chen et al., 2023) — a small draft model proposes several tokens and the big model checks them all in one pass, turning several bandwidth-bound steps into one.
  • Weight quantisation — 4-bit weights mean a quarter of the bytes per decode step, so decode speeds up; prefill gains only if the hardware also computes in low precision.

A real-life example

A code-completion assistant in an IDE sends about 3,000 tokens of surrounding code and asks for about 30 tokens back. Almost all the latency is prefill: users expect a suggestion within roughly 300 ms of pausing. The team keeps the model small, and uses prefix caching so that between two keystrokes only the few changed tokens at the end are prefilled again.

A support chatbot is the opposite: a 300-token question and a 600-token answer. Here decode dominates, the user watches tokens stream, and the team's levers are batching, quantisation and speculative decoding. Same model family, opposite tuning.

Follow-up questions to expect

  • "Why does TTFT grow with prompt length?" — Prefill work grows linearly for the weight matmuls and quadratically for attention, so a 100k-token prompt can take seconds before the first token.
  • "How does speculative decoding keep the output identical?" — The big model verifies each draft token with a rejection-sampling rule, so the output distribution is exactly the big model's; only speed changes.
  • "Does int4 quantisation speed up prefill?" — Mostly not; prefill is compute-bound and int4 weights are usually de-quantised to 16-bit for the math. It mainly speeds up decode.