Course Content
Local LLM Deployment and Quantization
3 sections · 7 lessons
GGUF, GGML, and Quantized Model Formats
Two files claim to be the same model. One is Meta-Llama-3.1-8B-Instruct-F16.gguf at 16.07 GB. The other is Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf at 4.92 GB. Same 8.03 billion parameters, same architecture, same training. The second is 3.3 times smaller and, on the standard perplexity benchmark, about 3% worse.
That trade — under a third of the size for a few percent of quality — is the entire reason anyone can run a serious language model on a laptop. It sounds like it should be impossible. Throwing away 12 of every 16 bits should ruin a system with billions of interacting parameters.
It does not, and the reason is specific and worth deriving rather than believing. By the end of this you should be able to take a block of eight weights, compute its scale and zero-point by hand, quantise it to 4 bits, dequantise it back, and state the error you introduced — and then explain why 4.92 GB is the right file size for a "4-bit" 8-billion-parameter model rather than the 4.0 GB you would naively predict.
Quantisation, derived from scratch
A trained weight is a 32-bit or 16-bit float — a number like -0.4713821. Quantisation asks: how much of that precision is load-bearing? If we only kept 16 distinct values instead of 65,536, how wrong would the model be?
The mechanism is a linear map from a real interval onto a small set of integers. Take a group of weights with minimum α and maximum β, and a target bit width b, giving 2b available integer levels. Two constants define the map:
s is the scale: the width of one integer step in real units. z is the zero-point: the integer that represents real zero. Quantising and dequantising are then:
The zero-point exists because weight distributions are not symmetric. If a group runs from −0.6 to +0.9, a scheme that assumed symmetry would have to cover [−0.9,+0.9] and waste a third of its levels on a region containing no weights.
A worked example, by hand
Take a block of eight weights, quantising to b=4 bits, so 16 levels numbered 0 to 15.
Weights: [−0.60,−0.21,0.04,0.33,0.90,−0.47,0.12,0.68]
Step one, find the range: α=−0.60, β=0.90, so β−α=1.50.
Step two, the scale: s=1.50/(24−1)=1.50/15=0.10.
Step three, the zero-point: z=round(0.60/0.10)=6. Integer 6 now means real zero, and you can verify the endpoints — −0.60 maps to round(−6)+6=0, and 0.90 maps to round(9)+6=15. The map uses the full range.
| Original x | x/s | q (4-bit) | Dequantised x^=0.1(q−6) | Error |
|---|---|---|---|---|
| −0.60 | −6.0 | 0 | −0.60 | 0.00 |
| −0.21 | −2.1 | 4 | −0.20 | +0.01 |
| 0.04 | 0.4 | 6 | 0.00 | −0.04 |
| 0.33 | 3.3 | 9 | 0.30 | −0.03 |
| 0.90 | 9.0 | 15 | 0.90 | 0.00 |
| −0.47 | −4.7 | 1 | −0.50 | −0.03 |
| 0.12 | 1.2 | 7 | 0.10 | −0.02 |
| 0.68 | 6.8 | 13 | 0.70 | +0.02 |
Every error is at most s/2=0.05, which is guaranteed by construction — rounding to the nearest grid point cannot be off by more than half a grid step. The root-mean-square error here is
which sits just under the theoretical value for uniform rounding error, s/12=0.1/3.464=0.029. The maths is behaving.
Symmetric quantisation, and why it is often used anyway
Drop the zero-point and centre the map on zero. With signed 4-bit integers running −8 to 7:
Now −0.60 quantises to round(−4.67)=−5, dequantising to −0.643 — an error of 0.043, where the affine map reproduced that weight exactly. Symmetric quantisation wastes levels whenever the distribution is lopsided.
It survives because it is cheaper at inference time. Dequantisation becomes a single multiply instead of a subtract-then-multiply, and in a matrix multiplication the scale can be factored out of the whole dot product. When the underlying distribution is roughly symmetric — which trained weights usually are — the accuracy cost is small and the speed gain is real.
| Affine (asymmetric) | Symmetric | |
|---|---|---|
| Parameters stored per block | scale and zero-point | scale only |
| Dequantisation | s(q−z) | s⋅q |
| Handles lopsided ranges | Yes | Poorly |
| Storage overhead per 32 weights | 4 bytes (two fp16) | 2 bytes (one fp16) |
| GGUF examples | Q4_1, Q5_1, Q2_K, Q4_K, Q5_K | Q4_0, Q5_0, Q8_0, Q3_K, Q6_K |
The scale is the width of one integer step; the zero-point is which integer means zero. Every quantisation scheme in this field is a variation on where those two numbers are computed and how often they are recomputed.
Why blocks, not whole tensors
The single most important design choice is how many weights share one scale. Consider a row of 4,096 weights where 4,095 lie within ±0.5 and one outlier sits at 12.0 — a completely normal situation in transformer layers, where a handful of channels carry outsized activations.
Share one scale across the whole row: s=(12.0−(−0.5))/15=0.833. Every weight smaller than 0.42 in magnitude now rounds into the same bucket. The row is destroyed to accommodate one number.
Split the row into 128 blocks of 32 weights each: the outlier is confined to one block, which gets a coarse scale of its own. The other 127 blocks get scales near 1.0/15=0.067 and keep their resolution. You have paid 127 extra scale factors — 254 bytes — and saved the layer.
This is why real quantised files carry per-block metadata, and it is why their true bit rate exceeds the nominal one.
What the file sizes actually are
Take GGUF's Q4_0: 32 weights per block, 4 bits each, plus one fp16 scale.
32 × 4 bits = 16 bytes, plus 2 bytes of scale = 18 bytes per 32 weights = 4.5 bits per weight.
Q4_1 adds an fp16 minimum for the affine map: 20 bytes per 32 weights = 5.0 bits per weight.
The K-quants are cleverer. Q4_K uses a superblock of 256 weights, split into eight sub-blocks of 32. The sub-block scales and minima are themselves quantised to 6 bits and stored in 12 bytes, with two fp16 values for the superblock. That gives 128 + 12 + 4 = 144 bytes per 256 weights = 4.5 bits per weight with affine-quality accuracy — a strictly better deal than Q4_1.
Then the _M and _S suffixes describe a mixture. Q4_K_M keeps about half of the attention value projections and feed-forward down projections at Q6_K (the first and last few layers, and every third layer in between), because those tensors are measurably more sensitive, and stores the output tensor at Q6_K too. Averaged over the whole model that lifts it to about 4.9 bits per weight:
8.03×109 × 4.9 / 8 = 4.92×109 bytes = 4.92 GB
which is exactly the published file size. The naive "4 bits means half a byte" calculation gives 4.02 GB and is wrong by 22%.
GGML and GGUF
GGML is the tensor library: C code that defines tensor types (including all the quantised block formats above), builds a computation graph, and executes it on CPU with AVX2/AVX-512/NEON, or on GPU via CUDA, Metal, Vulkan, or ROCm. Its distinguishing property is that matrix-multiply kernels consume quantised blocks directly, dequantising a block at a time inside registers. Nothing ever materialises the full FP16 tensor in memory.
GGUF is the file format that feeds it. It replaced the earlier GGML file format in August 2023, after several rounds of breaking changes taught everyone why the old design was inadequate.
| Problem with the old format | What GGUF does |
|---|---|
| Architecture hyperparameters hard-coded in the loader | Arbitrary key-value metadata in the file; new architectures need no format change |
| Tokeniser shipped separately, easily mismatched | Vocabulary, merges, and special token IDs embedded |
| Chat template lived in the runtime or the user's head | Jinja chat template stored as metadata |
| Every change broke old files | Versioned, with backwards compatibility as a design goal |
| Whole file read into RAM | Aligned tensor data, designed for mmap: pages load on demand and are shared between processes |
A GGUF file is, in order: a magic number and version; counts of tensors and metadata entries; the metadata key-value block; a tensor directory giving each tensor's name, shape, type, and byte offset; then the tensor data itself, aligned. You can read all of that without loading the weights:
1from gguf import GGUFReader23r = GGUFReader("Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf")45for field in r.fields.values():6 if "attention" in field.name or "block_count" in field.name:7 print(field.name, field.parts[field.data[0]])89total = 010for t in r.tensors:11 total += t.n_bytes12 if t.name.startswith("blk.0."):13 print(f"{t.name:32} {str(t.tensor_type):12} {t.shape} {t.n_bytes/1e6:8.2f} MB")1415print(f"total tensor bytes: {total/1e9:.2f} GB")Run that on a _M file and you will see the mixture directly: most tensors reporting Q4_K, about half of the down projections and value projections reporting Q6_K. The naming convention is not decoration; it describes a real per-tensor policy.
Reading a quantisation name
Q4_K_M decomposes as: Q for quantised, 4 for the nominal bits per weight, K for the K-quant family (superblocks with quantised sub-scales), M for the medium mixture. _0 means legacy symmetric, _1 means legacy affine, _S/_M/_L are small, medium, and large mixtures within a K-quant family.
| Format | Bits/weight | 8B file | Perplexity increase vs FP16 | Use it when |
|---|---|---|---|---|
| F16 | 16 | 16.07 GB | baseline | You are converting, or measuring a baseline |
| Q8_0 | 8.5 | 8.54 GB | < 0.1% | Memory is free and you want a reference-quality local model |
| Q6_K | 6.6 | 6.60 GB | ~0.3% | You have the VRAM. Effectively indistinguishable from FP16 |
| Q5_K_M | 5.7 | 5.73 GB | ~0.9% | Best quality that still fits many 8–12 GB budgets |
| Q4_K_M | 4.9 | 4.92 GB | ~2.8% | The default. Best size-to-quality point for almost everyone |
| Q4_K_S | 4.6 | 4.69 GB | ~4.3% | You need the extra 200 MB back |
| Q3_K_M | 4.0 | 4.02 GB | ~10% | Only if Q4 genuinely will not fit. Degradation is now visible |
| Q2_K | 3.2 | 3.18 GB | ~56% | Rarely. Instruction-following and formatting break |
The perplexity figures are the ones llama.cpp publishes for Llama 3 8B on Wikitext-2, without an importance matrix. Older models such as Llama 2 7B lose noticeably less at each step, so tables copied from 2023 look rosier than this. The figures also hide something important: the curve is flat then it is a cliff. Between 8 and 5 bits almost nothing happens. Between 5 and 4 bits a little happens. Below 4 bits, quality falls away quickly, and it falls away first in exactly the behaviours you care about — following formatting instructions, staying consistent over long outputs, arithmetic.
A larger model at a lower precision usually beats a smaller model at high precision: a 13B at Q4_K_M outperforms a 7B at Q8_0 while occupying similar memory. Spend your bytes on parameters before you spend them on precision — but not below about 4 bits.
Converting a model yourself
You need this when a model has no community GGUF, or when it is your own fine-tune.
1# 1. Get the source weights (safetensors, FP16/BF16)2# Gated model: accept the licence on its Hugging Face page, then run `hf auth login`3pip install -U huggingface_hub4hf download meta-llama/Llama-3.1-8B-Instruct \5 --local-dir ./llama-3.1-8b --exclude "original/*"67# 2. Convert to a full-precision GGUF (no quality loss, just a repack)8cd llama.cpp9python convert_hf_to_gguf.py ../llama-3.1-8b \10 --outfile llama-3.1-8b-f16.gguf --outtype f161112# 3. Quantise. This is where the arithmetic above happens, tensor by tensor.13./build/bin/llama-quantize llama-3.1-8b-f16.gguf \14 llama-3.1-8b-Q4_K_M.gguf Q4_K_M1516# 4. Check it before trusting it17./build/bin/llama-cli -m llama-3.1-8b-Q4_K_M.gguf \18 -p "List three prime numbers above 100." -st -n 64Step 2 is lossless and step 3 is where every byte is decided. You need disk space for both: 16 GB of source weights plus 16 GB of F16 GGUF plus 5 GB of output, so budget 40 GB even though the artefact you keep is 5 GB.
An importance matrix improves low-bit results substantially. You run a sample of representative text through the F16 model, record which weights matter most to the activations, and let the quantiser allocate its error budget accordingly:
1./build/bin/llama-imatrix -m llama-3.1-8b-f16.gguf \2 -f calibration.txt -o imatrix.gguf --chunks 20034./build/bin/llama-quantize --imatrix imatrix.gguf \5 llama-3.1-8b-f16.gguf llama-3.1-8b-IQ3_M.gguf IQ3_MFor Q4 and above the gain is small. At 3 bits and below it is the difference between usable and broken, which is why the modern IQ formats assume an importance matrix.
Using GGUF from Python
1from llama_cpp import Llama23llm = Llama(4 model_path="llama-3.1-8b-Q4_K_M.gguf",5 n_ctx=8192, # KV cache is sized from this: 8192 tokens x 128 KiB = 1 GiB6 n_gpu_layers=-1, # -1 offloads every layer; use a number to split with the CPU7 n_threads=8, # physical performance cores8 verbose=False,9)1011out = llm.create_chat_completion(12 messages=[{"role": "user", "content": "Why do 4-bit files store 4.9 bits per weight?"}],13 temperature=0.3,14 max_tokens=300,15)16print(out["choices"][0]["message"]["content"])17print(out["usage"]) # prompt_tokens, completion_tokensNote what n_ctx costs. It is not a limit you set generously "just in case" — it allocates the KV cache up front. Setting n_ctx=32768 on this model reserves 4 GiB of memory whether or not you ever send a long prompt.
Other quantisation families
GGUF is not the only scheme, and the alternatives are worth recognising because they optimise for different hardware.
| Scheme | Approach | Runs on | Best for |
|---|---|---|---|
| GGUF / K-quants | Block-wise, round-to-nearest, optional importance matrix | CPU, and every GPU backend | CPU inference, Apple Silicon, mixed CPU/GPU offload |
| GPTQ | Layer-wise error compensation: quantise one column, adjust the rest to absorb the error | NVIDIA GPU | Fully GPU-resident serving; strong 4-bit accuracy |
| AWQ | Finds the ~1% of salient weight channels from activation statistics and scales them to protect them | NVIDIA GPU | Fast 4-bit serving with high accuracy; popular with vLLM |
| bitsandbytes NF4 | 4-bit datatype whose levels match a normal distribution; quantises on load | NVIDIA GPU | Training-time quantisation, where the base model must stay frozen and small |
| EXL2 / EXL3 | Variable bit rate per layer to hit a target average | NVIDIA GPU | Squeezing the largest possible model onto a fixed VRAM budget |
The dividing line is simple: GGUF is the format that does not assume a GPU. Everything else assumes one, and buys accuracy or speed with that assumption.
Where this bites when you build
The failure mode people hit is picking a quantisation from a table and then discovering the model is subtly worse at the one thing their application needs. Aggregate perplexity is a poor proxy for that. If your product depends on the model emitting valid JSON, measure valid-JSON rate at Q4 and at Q6, on your prompts. It is a twenty-minute experiment and it occasionally reveals a two-percentage-point gap that no perplexity table would have shown you.
The second thing to internalise is that quantisation degrades non-uniformly across tasks. Free-form prose survives aggressive quantisation almost untouched, because there are many acceptable next words and small logit perturbations rarely change which one is chosen. Tasks with one correct answer — code that must compile, arithmetic, strict schema adherence — are where a 3% perplexity increase can show up as a noticeably higher error rate. Quantise conversational features harder than you quantise structured ones.
Third, remember what the file size does not include. A 4.92 GB Q4_K_M download plus an 8K context is 4.58 + 1.0 GiB, and by the time compute buffers are counted you are near 6.1 GiB — still fine on an 8 GB card, and impossible at 32K context where the cache alone wants 4 GiB. When you are one gigabyte short, dropping the KV cache to 8-bit is usually a better move than dropping the weights from Q4_K_M to Q3_K_M: you halve the cache for a barely-measurable quality cost, where the weight step costs you four times as much perplexity.
Finally, keep the F16 GGUF if you have the disk space. Every requantisation starts from full precision — you cannot go from Q4 to Q5, only from F16 to either. Re-downloading 16 GB because you deleted the intermediate is the most avoidable half-hour in this workflow.