LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What is pruning, and how does it reduce model size?


What you need to know

Three kinds of pruning

KindWhat is removedReal speed-up?Notes
UnstructuredIndividual weights, anywhereRarely on GPUsSparseGPT, Wanda: ~50% with small loss, no retraining
Semi-structured 2:42 of every 4 consecutive weightsSome, on Ampere and newer sparse tensor coresUp to 2× on the matrix math in theory; less end to end; more quality loss
StructuredWhole heads, channels, layersYes — smaller matricesNeeds retraining or distillation to "heal"

How pruning decides what to remove

  • Magnitude — remove the smallest weights.
  • Wanda — weight magnitude × size of the input activation, so a small weight on a busy input is kept.
  • SparseGPT — removes weights and adjusts the remaining ones to compensate, layer by layer.
  • Structured — score each head, channel or layer by its importance on sample data, remove the least important, then train the smaller model to copy the original (distillation).

Where it fits in practice

Structured pruning plus distillation is a model-production technique. For example, NVIDIA's Minitron work pruned larger models (such as Llama 3.1 8B) in width or depth and then distilled them into smaller models with much less training than starting from scratch. For an application team, the practical version is: pick a smaller model someone has already pruned and distilled, then quantize it.

A real-life example

A hospital wants its summariser to run on CPU-only servers in small district clinics, with 32 GB of RAM and no GPU. An engineer proposes pruning their 8B model to 50% sparsity with Wanda.

The test shows the problem: 50% unstructured sparsity keeps quality within 2 points, but llama.cpp on CPU runs no faster, because the matrices are the same shape with zeros inside. File size falls only if a sparse format is used, and the runtime does not support it.

What works instead: a 3–4B model that was already made smaller by structured pruning and distillation, quantized to 4-bit GGUF (about 2–2.5 GB), running with llama.cpp at a usable speed on the clinic CPUs. On the eval set it scores 81% against 88% for the 8B model on the GPU server, so clinics use it for first drafts that a doctor reviews, and complex cases are sent to the central server.

Follow-up questions to expect

  • "Why doesn't unstructured sparsity speed up GPUs?" — GPU matrix units work on dense blocks; skipping random zeros needs special kernels, and the overhead usually cancels the saving at 50% sparsity.
  • "Pruning or quantization first?" — Quantization: it is simpler, well-supported by serving engines, and gives reliable speed and memory gains.
  • "Can you combine them?" — Yes; pruned-and-distilled models are often quantized for serving.