Course Content
LLMOps & Deployment
6 sections · 40 lessons
What are the trade-offs between static and dynamic quantization?
What you need to know
Two separate questions
- What is quantized? Weights only, or weights and activations (written W8A8, W4A16 and so on: bits for weights, bits for activations).
- When are activation scales decided? Offline (static) or at runtime (dynamic).
Weights never change after training, so their scales are always computed offline. The static-versus-dynamic question is about activations, which change with every input.
Static versus dynamic activation scales
| Static | Dynamic | |
|---|---|---|
| Scales computed | Offline, from calibration data | At runtime, per tensor or per token |
| Needs calibration set | Yes, and it must match production | No |
| Runtime overhead | None | Small (find the max, then scale) |
| Unexpected outliers | Clipped — accuracy loss | Handled |
| Typical LLM use | Static FP8/INT8 for maximum speed | FP8 with per-token dynamic scales |
A tiny example shows the trade-off:
1def quantize(xs, scale):2 """Symmetric int8: map [-scale, scale] onto [-127, 127], then back."""3 q = [max(-127, min(127, round(x / scale * 127))) for x in xs]4 return [v * scale / 127 for v in q]56calibration_max = 4.0 # static: largest value seen offline7token = [0.02, -1.3, 2.7, 9.0] # a production token with an outlier89static = quantize(token, calibration_max)10dynamic = quantize(token, max(abs(x) for x in token)) # scale from this token1112print("static :", [round(v, 2) for v in static]) # [0.03, -1.29, 2.71, 4.0]13print("dynamic:", [round(v, 2) for v in dynamic]) # [0.0, -1.28, 2.69, 9.0]The static scale, from calibration data that never saw a value above 4.0, clips the outlier 9.0 to 4.0 — a big error. The dynamic scale keeps 9.0, but the wider range costs precision on small values (0.02 becomes 0.0). LLM activations have exactly these outliers, which is why methods like SmoothQuant move them into the weights, and why per-token dynamic scaling is popular.
Weight-only versus W8A8
Weight-only (e.g. AWQ, GPTQ 4-bit)
- Weights 4-bit, math still in bf16
- Biggest memory saving
- Fastest at low batch (bandwidth-bound)
- Small calibration set for AWQ/GPTQ
W8A8 (e.g. FP8)
- Weights and activations 8-bit
- Math runs on fast 8-bit tensor cores
- Best at high batch (compute-bound)
- Needs hardware support (Hopper and newer for FP8)
Rule of thumb: few users, small GPU, memory-tight → 4-bit weight-only. Many users on H100-class GPUs → FP8. Tools such as vLLM's llm-compressor produce both kinds of checkpoints, and vLLM can apply dynamic FP8 on the fly at load time.
A real-life example
A code assistant for 2,000 engineers self-hosts a 32B coding model on H100 GPUs. At peak, each replica serves 60 concurrent requests, so the GPUs are compute-bound, not memory-bound.
The team tries two options. A 4-bit AWQ checkpoint halves memory again but barely raises throughput at batch 60, because the math still runs in bf16. An FP8 checkpoint with dynamic per-token activation scales raises throughput per GPU by roughly 1.6× in their load test and needs no calibration data. A static FP8 version with scales calibrated on public code samples was slightly faster, but lost 3 points on their internal-code eval — their code uses patterns the calibration set did not cover. They ship dynamic FP8, and cut the GPU count from 10 to 7.
Follow-up questions to expect
- "Why do LLMs have activation outliers?" — A few channels in some layers carry very large values; a single per-tensor scale then wastes most of the range on them.
- "How big should a calibration set be?" — Usually a few hundred samples, but they must look like production traffic: same languages, domains and lengths.
- "Is dynamic always better?" — No; it costs a little runtime and still loses precision on small values when outliers widen the range. Measure both.