LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What is model quantization, and how does it improve efficiency?


Weight memory by precision16 GB8 GB~5 GB140 GB70 GB~40 GBbf16FP8 or int84-bit8B70BTwo L40S GPUs hold 96 GB: the 70B fits only at 4-bit, with about 40 GB left for KV cache.
Halving the bytes halves the memory and, at small batch, roughly doubles decode speed, because every token reads every weight.

What you need to know

Memory, with numbers

Weights take bytes per parameter × number of parameters:

Modelbf16 (2 bytes)FP8 / int8 (1 byte)4-bit (~0.5 byte + scales)
8B16 GB8 GB~5 GB
70B140 GB70 GB~38–40 GB

So a 70B model needs two 80 GB GPUs in bf16. In FP8 its 70 GB of weights technically fit one 80 GB GPU but leave almost no room for KV cache, so it is usually served on two GPUs or on one 141 GB H200. In 4-bit it fits one 80 GB GPU with room for KV cache.

Why it also makes generation faster

In the decode phase, each step reads every weight once. A rough ceiling for one request's speed is memory bandwidth ÷ weight size. On an L40S (about 864 GB/s):

Text
8B bf16 : 864 / 16  ≈  54 tokens/s ceiling8B 4-bit: 864 / 4.5 ≈ 190 tokens/s ceiling

Real numbers are lower, but the ratio holds at small batch sizes. At large batch sizes the GPU becomes compute-bound, and then only formats that also run the math in low precision (FP8 or int8 activations, FP4 on Blackwell GPUs) keep helping.

The main formats (2026)

  • FP8 (W8A8) — weights and activations in 8-bit float; native on NVIDIA Hopper (H100/H200) and newer. Near-lossless; the usual choice for high-throughput serving on these GPUs.
  • INT8 — weight-only or W8A8 (for example with SmoothQuant); good on older GPUs such as A100.
  • 4-bit weight-only — AWQ and GPTQ for GPU serving; GGUF quantization types for llama.cpp and Ollama. Biggest memory saving; some quality loss.
  • FP4 (NVFP4, MXFP4) — 4-bit floats with small-block scales, supported by Blackwell GPUs.
  • KV-cache quantization — storing the KV cache in FP8 halves its size, so twice as many tokens fit.

Quality

8-bit is usually within noise of bf16 on most tasks. 4-bit loses a little on average, but the loss is uneven: maths, code, long context and less common languages often suffer more. Published numbers are for someone else's tasks — run your own eval set.

A real-life example

A hospital wants a stronger summariser than its 8B model, using a 70B model, on a server with two 48 GB L40S GPUs (96 GB total). In bf16, the 70B model needs 140 GB — it does not fit.

With 4-bit AWQ weights (~40 GB) split across both GPUs with tensor parallelism, about 40 GB remains for KV cache after overheads. At 320 KB per token for this model, that is about 130,000 tokens — roughly 16 discharge notes of 8,000 tokens in flight at once, plenty for their volume. The speed ceiling is about 43 tokens a second per request, so a 500-token summary takes 12–15 seconds, which doctors accept for a background task.

On the 400-case eval set, the 4-bit 70B scores 91% against 88% for the bf16 8B model, but misses slightly more dosage details in long notes. The team keeps the medication-extraction step on a separate structured-output call and ships the 70B.

Follow-up questions to expect

  • "Why is decoding memory-bound?" — Each step does little math per weight (one token's worth) but must read every weight, so memory bandwidth, not compute, is the limit.
  • "Would you quantize the KV cache?" — Often yes (FP8), when memory limits concurrency; check long-context quality.
  • "AWQ or GPTQ?" — Both are 4-bit weight-only methods using a small calibration set; quality is similar, so choose by what your engine supports best and by your own evals.