Local LLM Deployment and Quantization

GGUF, GGML, and Quantized Model Formats


Two files claim to be the same model. One is Meta-Llama-3.1-8B-Instruct-F16.gguf at 16.07 GB. The other is Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf at 4.92 GB. Same 8.03 billion parameters, same architecture, same training. The second is 3.3 times smaller and, on the standard perplexity benchmark, about 3% worse.

That trade — under a third of the size for a few percent of quality — is the entire reason anyone can run a serious language model on a laptop. It sounds like it should be impossible. Throwing away 12 of every 16 bits should ruin a system with billions of interacting parameters.

It does not, and the reason is specific and worth deriving rather than believing. By the end of this you should be able to take a block of eight weights, compute its scale and zero-point by hand, quantise it to 4 bits, dequantise it back, and state the error you introduced — and then explain why 4.92 GB is the right file size for a "4-bit" 8-billion-parameter model rather than the 4.0 GB you would naively predict.

Llama 3.1 8B at every quantisation level16.07 GB16.0baseline8.54 GB8.5under 0.1 pct6.60 GB6.6about 0.3 pct5.73 GB5.7about 0.9 pct4.92 GB4.9about 2.8 pct4.02 GB4.0about 10 pctFile sizeBits per weightPerplexity changeF16Q8_0Q6_KQ5_K_MQ4_K_MQ3_K_MQuantisation is per block of 32 weights, so one outlier widens one scale rather than the whole tensor.
The curve is nearly flat down to Q4_K_M and falls off a cliff below it — 3.3 times smaller for about 3 percent more perplexity.

Quantisation, derived from scratch

A trained weight is a 32-bit or 16-bit float — a number like -0.4713821. Quantisation asks: how much of that precision is load-bearing? If we only kept 16 distinct values instead of 65,536, how wrong would the model be?

The mechanism is a linear map from a real interval onto a small set of integers. Take a group of weights with minimum α\alpha and maximum β\beta, and a target bit width bb, giving 2b2^b available integer levels. Two constants define the map:

s=β−α2b−1z=round⁡ ⁣(−αs)s = \frac{\beta - \alpha}{2^b - 1} \qquad\qquad z = \operatorname{round}\!\left(-\frac{\alpha}{s}\right)

ss is the scale: the width of one integer step in real units. zz is the zero-point: the integer that represents real zero. Quantising and dequantising are then:

q=clip⁡ ⁣(round⁡ ⁣(xs)+z,  0,  2b−1)x^=s (q−z)q = \operatorname{clip}\!\left(\operatorname{round}\!\left(\frac{x}{s}\right) + z,\; 0,\; 2^b - 1\right) \qquad\qquad \hat{x} = s\,(q - z)

The zero-point exists because weight distributions are not symmetric. If a group runs from −0.6-0.6 to +0.9+0.9, a scheme that assumed symmetry would have to cover [−0.9,+0.9][-0.9, +0.9] and waste a third of its levels on a region containing no weights.

A worked example, by hand

Take a block of eight weights, quantising to b=4b = 4 bits, so 16 levels numbered 0 to 15.

Weights: [−0.60,  −0.21,  0.04,  0.33,  0.90,  −0.47,  0.12,  0.68][-0.60,\; -0.21,\; 0.04,\; 0.33,\; 0.90,\; -0.47,\; 0.12,\; 0.68]

Step one, find the range: α=−0.60\alpha = -0.60, β=0.90\beta = 0.90, so β−α=1.50\beta - \alpha = 1.50.

Step two, the scale: s=1.50/(24−1)=1.50/15=0.10s = 1.50 / (2^4 - 1) = 1.50 / 15 = 0.10.

Step three, the zero-point: z=round⁡(0.60/0.10)=6z = \operatorname{round}(0.60 / 0.10) = 6. Integer 6 now means real zero, and you can verify the endpoints — −0.60-0.60 maps to round⁡(−6)+6=0\operatorname{round}(-6) + 6 = 0, and 0.900.90 maps to round⁡(9)+6=15\operatorname{round}(9) + 6 = 15. The map uses the full range.

Original xxx/sx/sqq (4-bit)Dequantised x^=0.1(q−6)\hat{x} = 0.1(q-6)Error
−0.60−6.00−0.600.00
−0.21−2.14−0.20+0.01
0.040.460.00−0.04
0.333.390.30−0.03
0.909.0150.900.00
−0.47−4.71−0.50−0.03
0.121.270.10−0.02
0.686.8130.70+0.02

Every error is at most s/2=0.05s/2 = 0.05, which is guaranteed by construction — rounding to the nearest grid point cannot be off by more than half a grid step. The root-mean-square error here is

18(02+0.012+0.042+0.032+02+0.032+0.022+0.022)=0.000538=0.023\sqrt{\tfrac{1}{8}\left(0^2 + 0.01^2 + 0.04^2 + 0.03^2 + 0^2 + 0.03^2 + 0.02^2 + 0.02^2\right)} = \sqrt{0.000538} = 0.023

which sits just under the theoretical value for uniform rounding error, s/12=0.1/3.464=0.029s/\sqrt{12} = 0.1/3.464 = 0.029. The maths is behaving.

Symmetric quantisation, and why it is often used anyway

Drop the zero-point and centre the map on zero. With signed 4-bit integers running −8-8 to 77:

s=max⁡∣x∣2b−1−1=0.907=0.1286,z=0s = \frac{\max|x|}{2^{b-1} - 1} = \frac{0.90}{7} = 0.1286, \qquad z = 0

Now −0.60-0.60 quantises to round⁡(−4.67)=−5\operatorname{round}(-4.67) = -5, dequantising to −0.643-0.643 — an error of 0.043, where the affine map reproduced that weight exactly. Symmetric quantisation wastes levels whenever the distribution is lopsided.

It survives because it is cheaper at inference time. Dequantisation becomes a single multiply instead of a subtract-then-multiply, and in a matrix multiplication the scale can be factored out of the whole dot product. When the underlying distribution is roughly symmetric — which trained weights usually are — the accuracy cost is small and the speed gain is real.

Affine (asymmetric)Symmetric
Parameters stored per blockscale and zero-pointscale only
Dequantisations(q−z)s(q - z)s⋅qs \cdot q
Handles lopsided rangesYesPoorly
Storage overhead per 32 weights4 bytes (two fp16)2 bytes (one fp16)
GGUF examplesQ4_1, Q5_1, Q2_K, Q4_K, Q5_KQ4_0, Q5_0, Q8_0, Q3_K, Q6_K

The scale is the width of one integer step; the zero-point is which integer means zero. Every quantisation scheme in this field is a variation on where those two numbers are computed and how often they are recomputed.

Why blocks, not whole tensors

The single most important design choice is how many weights share one scale. Consider a row of 4,096 weights where 4,095 lie within ±0.5\pm 0.5 and one outlier sits at 12.012.0 — a completely normal situation in transformer layers, where a handful of channels carry outsized activations.

Share one scale across the whole row: s=(12.0−(−0.5))/15=0.833s = (12.0 - (-0.5))/15 = 0.833. Every weight smaller than 0.42 in magnitude now rounds into the same bucket. The row is destroyed to accommodate one number.

Split the row into 128 blocks of 32 weights each: the outlier is confined to one block, which gets a coarse scale of its own. The other 127 blocks get scales near 1.0/15=0.0671.0/15 = 0.067 and keep their resolution. You have paid 127 extra scale factors — 254 bytes — and saved the layer.

This is why real quantised files carry per-block metadata, and it is why their true bit rate exceeds the nominal one.

What the file sizes actually are

Take GGUF's Q4_0: 32 weights per block, 4 bits each, plus one fp16 scale.

32 × 4 bits = 16 bytes, plus 2 bytes of scale = 18 bytes per 32 weights = 4.5 bits per weight.

Q4_1 adds an fp16 minimum for the affine map: 20 bytes per 32 weights = 5.0 bits per weight.

The K-quants are cleverer. Q4_K uses a superblock of 256 weights, split into eight sub-blocks of 32. The sub-block scales and minima are themselves quantised to 6 bits and stored in 12 bytes, with two fp16 values for the superblock. That gives 128 + 12 + 4 = 144 bytes per 256 weights = 4.5 bits per weight with affine-quality accuracy — a strictly better deal than Q4_1.

Then the _M and _S suffixes describe a mixture. Q4_K_M keeps about half of the attention value projections and feed-forward down projections at Q6_K (the first and last few layers, and every third layer in between), because those tensors are measurably more sensitive, and stores the output tensor at Q6_K too. Averaged over the whole model that lifts it to about 4.9 bits per weight:

8.03×1098.03 \times 10^{9} × 4.9 / 8 = 4.92×1094.92 \times 10^{9} bytes = 4.92 GB

which is exactly the published file size. The naive "4 bits means half a byte" calculation gives 4.02 GB and is wrong by 22%.

GGML and GGUF

GGML is the tensor library: C code that defines tensor types (including all the quantised block formats above), builds a computation graph, and executes it on CPU with AVX2/AVX-512/NEON, or on GPU via CUDA, Metal, Vulkan, or ROCm. Its distinguishing property is that matrix-multiply kernels consume quantised blocks directly, dequantising a block at a time inside registers. Nothing ever materialises the full FP16 tensor in memory.

GGUF is the file format that feeds it. It replaced the earlier GGML file format in August 2023, after several rounds of breaking changes taught everyone why the old design was inadequate.

Problem with the old formatWhat GGUF does
Architecture hyperparameters hard-coded in the loaderArbitrary key-value metadata in the file; new architectures need no format change
Tokeniser shipped separately, easily mismatchedVocabulary, merges, and special token IDs embedded
Chat template lived in the runtime or the user's headJinja chat template stored as metadata
Every change broke old filesVersioned, with backwards compatibility as a design goal
Whole file read into RAMAligned tensor data, designed for mmap: pages load on demand and are shared between processes

A GGUF file is, in order: a magic number and version; counts of tensors and metadata entries; the metadata key-value block; a tensor directory giving each tensor's name, shape, type, and byte offset; then the tensor data itself, aligned. You can read all of that without loading the weights:

Python
from gguf import GGUFReaderr = GGUFReader("Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf")for field in r.fields.values():    if "attention" in field.name or "block_count" in field.name:        print(field.name, field.parts[field.data[0]])total = 0for t in r.tensors:    total += t.n_bytes    if t.name.startswith("blk.0."):        print(f"{t.name:32} {str(t.tensor_type):12} {t.shape} {t.n_bytes/1e6:8.2f} MB")print(f"total tensor bytes: {total/1e9:.2f} GB")

Run that on a _M file and you will see the mixture directly: most tensors reporting Q4_K, about half of the down projections and value projections reporting Q6_K. The naming convention is not decoration; it describes a real per-tensor policy.

Reading a quantisation name

Q4_K_M decomposes as: Q for quantised, 4 for the nominal bits per weight, K for the K-quant family (superblocks with quantised sub-scales), M for the medium mixture. _0 means legacy symmetric, _1 means legacy affine, _S/_M/_L are small, medium, and large mixtures within a K-quant family.

FormatBits/weight8B filePerplexity increase vs FP16Use it when
F161616.07 GBbaselineYou are converting, or measuring a baseline
Q8_08.58.54 GB< 0.1%Memory is free and you want a reference-quality local model
Q6_K6.66.60 GB~0.3%You have the VRAM. Effectively indistinguishable from FP16
Q5_K_M5.75.73 GB~0.9%Best quality that still fits many 8–12 GB budgets
Q4_K_M4.94.92 GB~2.8%The default. Best size-to-quality point for almost everyone
Q4_K_S4.64.69 GB~4.3%You need the extra 200 MB back
Q3_K_M4.04.02 GB~10%Only if Q4 genuinely will not fit. Degradation is now visible
Q2_K3.23.18 GB~56%Rarely. Instruction-following and formatting break

The perplexity figures are the ones llama.cpp publishes for Llama 3 8B on Wikitext-2, without an importance matrix. Older models such as Llama 2 7B lose noticeably less at each step, so tables copied from 2023 look rosier than this. The figures also hide something important: the curve is flat then it is a cliff. Between 8 and 5 bits almost nothing happens. Between 5 and 4 bits a little happens. Below 4 bits, quality falls away quickly, and it falls away first in exactly the behaviours you care about — following formatting instructions, staying consistent over long outputs, arithmetic.

A larger model at a lower precision usually beats a smaller model at high precision: a 13B at Q4_K_M outperforms a 7B at Q8_0 while occupying similar memory. Spend your bytes on parameters before you spend them on precision — but not below about 4 bits.

Converting a model yourself

You need this when a model has no community GGUF, or when it is your own fine-tune.

Bash
# 1. Get the source weights (safetensors, FP16/BF16)# Gated model: accept the licence on its Hugging Face page, then run `hf auth login`pip install -U huggingface_hubhf download meta-llama/Llama-3.1-8B-Instruct \  --local-dir ./llama-3.1-8b --exclude "original/*"# 2. Convert to a full-precision GGUF (no quality loss, just a repack)cd llama.cpppython convert_hf_to_gguf.py ../llama-3.1-8b \  --outfile llama-3.1-8b-f16.gguf --outtype f16# 3. Quantise. This is where the arithmetic above happens, tensor by tensor../build/bin/llama-quantize llama-3.1-8b-f16.gguf \  llama-3.1-8b-Q4_K_M.gguf Q4_K_M# 4. Check it before trusting it./build/bin/llama-cli -m llama-3.1-8b-Q4_K_M.gguf \  -p "List three prime numbers above 100." -st -n 64

Step 2 is lossless and step 3 is where every byte is decided. You need disk space for both: 16 GB of source weights plus 16 GB of F16 GGUF plus 5 GB of output, so budget 40 GB even though the artefact you keep is 5 GB.

An importance matrix improves low-bit results substantially. You run a sample of representative text through the F16 model, record which weights matter most to the activations, and let the quantiser allocate its error budget accordingly:

Bash
./build/bin/llama-imatrix -m llama-3.1-8b-f16.gguf \  -f calibration.txt -o imatrix.gguf --chunks 200./build/bin/llama-quantize --imatrix imatrix.gguf \  llama-3.1-8b-f16.gguf llama-3.1-8b-IQ3_M.gguf IQ3_M

For Q4 and above the gain is small. At 3 bits and below it is the difference between usable and broken, which is why the modern IQ formats assume an importance matrix.

Using GGUF from Python

Python
from llama_cpp import Llamallm = Llama(    model_path="llama-3.1-8b-Q4_K_M.gguf",    n_ctx=8192,        # KV cache is sized from this: 8192 tokens x 128 KiB = 1 GiB    n_gpu_layers=-1,   # -1 offloads every layer; use a number to split with the CPU    n_threads=8,       # physical performance cores    verbose=False,)out = llm.create_chat_completion(    messages=[{"role": "user", "content": "Why do 4-bit files store 4.9 bits per weight?"}],    temperature=0.3,    max_tokens=300,)print(out["choices"][0]["message"]["content"])print(out["usage"])   # prompt_tokens, completion_tokens

Note what n_ctx costs. It is not a limit you set generously "just in case" — it allocates the KV cache up front. Setting n_ctx=32768 on this model reserves 4 GiB of memory whether or not you ever send a long prompt.

Other quantisation families

GGUF is not the only scheme, and the alternatives are worth recognising because they optimise for different hardware.

SchemeApproachRuns onBest for
GGUF / K-quantsBlock-wise, round-to-nearest, optional importance matrixCPU, and every GPU backendCPU inference, Apple Silicon, mixed CPU/GPU offload
GPTQLayer-wise error compensation: quantise one column, adjust the rest to absorb the errorNVIDIA GPUFully GPU-resident serving; strong 4-bit accuracy
AWQFinds the ~1% of salient weight channels from activation statistics and scales them to protect themNVIDIA GPUFast 4-bit serving with high accuracy; popular with vLLM
bitsandbytes NF44-bit datatype whose levels match a normal distribution; quantises on loadNVIDIA GPUTraining-time quantisation, where the base model must stay frozen and small
EXL2 / EXL3Variable bit rate per layer to hit a target averageNVIDIA GPUSqueezing the largest possible model onto a fixed VRAM budget

The dividing line is simple: GGUF is the format that does not assume a GPU. Everything else assumes one, and buys accuracy or speed with that assumption.

Where this bites when you build

The failure mode people hit is picking a quantisation from a table and then discovering the model is subtly worse at the one thing their application needs. Aggregate perplexity is a poor proxy for that. If your product depends on the model emitting valid JSON, measure valid-JSON rate at Q4 and at Q6, on your prompts. It is a twenty-minute experiment and it occasionally reveals a two-percentage-point gap that no perplexity table would have shown you.

The second thing to internalise is that quantisation degrades non-uniformly across tasks. Free-form prose survives aggressive quantisation almost untouched, because there are many acceptable next words and small logit perturbations rarely change which one is chosen. Tasks with one correct answer — code that must compile, arithmetic, strict schema adherence — are where a 3% perplexity increase can show up as a noticeably higher error rate. Quantise conversational features harder than you quantise structured ones.

Third, remember what the file size does not include. A 4.92 GB Q4_K_M download plus an 8K context is 4.58 + 1.0 GiB, and by the time compute buffers are counted you are near 6.1 GiB — still fine on an 8 GB card, and impossible at 32K context where the cache alone wants 4 GiB. When you are one gigabyte short, dropping the KV cache to 8-bit is usually a better move than dropping the weights from Q4_K_M to Q3_K_M: you halve the cache for a barely-measurable quality cost, where the weight step costs you four times as much perplexity.

Finally, keep the F16 GGUF if you have the disk space. Every requantisation starts from full precision — you cannot go from Q4 to Q5, only from F16 to either. Re-downloading 16 GB because you deleted the intermediate is the most avoidable half-hour in this workflow.