LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What are the trade-offs between static and dynamic quantization?


One token with an outlier, quantized to int80.02-1.32.79.00.03-1.292.714.00.0-1.282.699.0abcdoriginalstatic, max 4.0dynamic, max 9.0
Static scales clip the outlier the calibration set never saw; dynamic scales keep it but lose the smallest values.

What you need to know

Two separate questions

  1. What is quantized? Weights only, or weights and activations (written W8A8, W4A16 and so on: bits for weights, bits for activations).
  2. When are activation scales decided? Offline (static) or at runtime (dynamic).

Weights never change after training, so their scales are always computed offline. The static-versus-dynamic question is about activations, which change with every input.

Static versus dynamic activation scales

StaticDynamic
Scales computedOffline, from calibration dataAt runtime, per tensor or per token
Needs calibration setYes, and it must match productionNo
Runtime overheadNoneSmall (find the max, then scale)
Unexpected outliersClipped — accuracy lossHandled
Typical LLM useStatic FP8/INT8 for maximum speedFP8 with per-token dynamic scales

A tiny example shows the trade-off:

Python
def quantize(xs, scale):    """Symmetric int8: map [-scale, scale] onto [-127, 127], then back."""    q = [max(-127, min(127, round(x / scale * 127))) for x in xs]    return [v * scale / 127 for v in q]calibration_max = 4.0                 # static: largest value seen offlinetoken = [0.02, -1.3, 2.7, 9.0]        # a production token with an outlierstatic = quantize(token, calibration_max)dynamic = quantize(token, max(abs(x) for x in token))   # scale from this tokenprint("static :", [round(v, 2) for v in static])    # [0.03, -1.29, 2.71, 4.0]print("dynamic:", [round(v, 2) for v in dynamic])   # [0.0, -1.28, 2.69, 9.0]

The static scale, from calibration data that never saw a value above 4.0, clips the outlier 9.0 to 4.0 — a big error. The dynamic scale keeps 9.0, but the wider range costs precision on small values (0.02 becomes 0.0). LLM activations have exactly these outliers, which is why methods like SmoothQuant move them into the weights, and why per-token dynamic scaling is popular.

Weight-only versus W8A8

Weight-only (e.g. AWQ, GPTQ 4-bit)

  • Weights 4-bit, math still in bf16
  • Biggest memory saving
  • Fastest at low batch (bandwidth-bound)
  • Small calibration set for AWQ/GPTQ

W8A8 (e.g. FP8)

  • Weights and activations 8-bit
  • Math runs on fast 8-bit tensor cores
  • Best at high batch (compute-bound)
  • Needs hardware support (Hopper and newer for FP8)

Rule of thumb: few users, small GPU, memory-tight → 4-bit weight-only. Many users on H100-class GPUs → FP8. Tools such as vLLM's llm-compressor produce both kinds of checkpoints, and vLLM can apply dynamic FP8 on the fly at load time.

A real-life example

A code assistant for 2,000 engineers self-hosts a 32B coding model on H100 GPUs. At peak, each replica serves 60 concurrent requests, so the GPUs are compute-bound, not memory-bound.

The team tries two options. A 4-bit AWQ checkpoint halves memory again but barely raises throughput at batch 60, because the math still runs in bf16. An FP8 checkpoint with dynamic per-token activation scales raises throughput per GPU by roughly 1.6× in their load test and needs no calibration data. A static FP8 version with scales calibrated on public code samples was slightly faster, but lost 3 points on their internal-code eval — their code uses patterns the calibration set did not cover. They ship dynamic FP8, and cut the GPU count from 10 to 7.

Follow-up questions to expect

  • "Why do LLMs have activation outliers?" — A few channels in some layers carry very large values; a single per-tensor scale then wastes most of the range on them.
  • "How big should a calibration set be?" — Usually a few hundred samples, but they must look like production traffic: same languages, domains and lengths.
  • "Is dynamic always better?" — No; it costs a little runtime and still loses precision on small values when outliers widen the range. Measure both.