Course Content
Local LLM Deployment and Quantization
3 sections · 7 lessons
Hardware Optimization (CPU / GPU / Apple Silicon / TPU)
Someone posts a benchmark: Llama 3 8B at Q4_K_M, 4 tokens per second. They have a 16-core Ryzen 9, 64 GB of DDR5, and an RTX 4070. The obvious advice arrives — more threads, bigger batch, a faster CPU. All of it is wrong. The problem is not the CPU at all, and the fix is one setting.
Their model file is 4.92 GB and their GPU has 12 GB of VRAM. The model fits with room to spare. But they were loading it through llama-cpp-python, where n_gpu_layers defaults to zero, so all 32 transformer layers were running on the CPU, and the CPU's memory bandwidth is 89.6 GB/s against the GPU's 504 GB/s. (The llama.cpp command-line tools now default to auto and offload what fits; many bindings and older wrappers do not.) Setting n_gpu_layers=-1 — -ngl 99 on the command line — took them from 4 tokens per second to about 95.
That factor of twenty-four was never about cores or threads. It was about how many bytes per second the processor can pull from memory, because generating one token requires reading every weight in the model exactly once. Get that one idea straight and most hardware tuning becomes arithmetic instead of folklore.
The equation that predicts your speed
Token generation in a transformer is memory-bandwidth-bound. To produce token number 501, the model must multiply a single vector against every weight matrix in sequence. Each weight is touched once and then discarded. There is almost no arithmetic per byte loaded — roughly two floating-point operations per byte — so the processor spends nearly all its time waiting for memory.
That gives a ceiling you can compute before buying anything:
Apply it to a 4.92 GB Q4_K_M file:
| Hardware | Bandwidth | Theoretical ceiling | Observed, roughly | Efficiency |
|---|---|---|---|---|
| DDR4-3200 dual channel | 51.2 GB/s | 10.4 tok/s | 6–7 | ~62% |
| DDR5-5600 dual channel | 89.6 GB/s | 18.2 tok/s | 10–12 | ~60% |
| Apple M1 Pro | 200 GB/s | 40.7 tok/s | 25–30 | ~68% |
| Apple M3 Max | 400 GB/s | 81 tok/s | 55–65 | ~74% |
| RTX 3060 12 GB | 360 GB/s | 73 tok/s | 50–60 | ~75% |
| RTX 4070 | 504 GB/s | 102 tok/s | 75–95 | ~83% |
| Apple M2 Ultra | 800 GB/s | 163 tok/s | 95–115 | ~64% |
| RTX 4090 | 1008 GB/s | 205 tok/s | 130–160 | ~71% |
Two consequences follow immediately, and they overturn most intuitions:
A smaller quantisation is faster, not just smaller. Going from Q8_0 (8.54 GB) to Q4_K_M (4.92 GB) on the same hardware nearly halves the bytes read per token, so it nearly doubles throughput. The speed-up is a side effect of the size reduction, not of any cleverness in the arithmetic.
Core count barely matters for generation. A 32-core CPU on the same memory bus as a 16-core CPU generates at nearly the same speed, because both are waiting on the same 89.6 GB/s. Cores matter for prompt processing, which is a different regime entirely.
Generation speed is bandwidth divided by model size. Everything else is a second-order correction to that number.
The other regime: prefill
Reading a 900-token prompt is not the same operation as writing a token. Prefill processes all 900 tokens at once, so each weight loaded is used 900 times. The ratio of arithmetic to memory traffic is now enormous, and prefill is compute-bound.
Its cost is roughly 2×Nparams×ntokens floating-point operations. For 8.03B parameters and 900 tokens:
2 × 8.03×109 × 900 = 1.45×1013 = 14.5 TFLOPs
An RTX 4090 delivering around 66 TFLOP/s in practice finishes that in 0.22 s. A good desktop CPU managing 0.5 TFLOP/s takes 29 seconds. That is why CPU-only inference feels acceptable in a chat toy with three-word prompts and unusable the moment you paste in a document: the token rate is tolerable, the time-to-first-token is not.
| Prefill (prompt processing) | Decode (generation) | |
|---|---|---|
| Bound by | Compute (FLOP/s) | Memory bandwidth |
| Helped by | More cores, GPU, larger batch | Faster memory, smaller model |
| Scales with | Prompt length | Output length |
| User perceives as | Time to first token | Words per second |
| CPU vs GPU gap | 50–100× | 5–15× |
CPU inference
CPU is not a fallback to be ashamed of. For an 8B model at Q4_K_M with short prompts, a modern desktop produces 10–12 tokens per second, which is faster than most people read. The tuning that matters is narrow.
Threads
Set -t to the number of physical performance cores. Not logical threads, not efficiency cores. Two hyperthreads on one physical core share one set of load/store units and one L1 cache, so pairing them on a bandwidth-bound workload adds contention and no throughput. On an Apple M-series or Intel P/E-core chip, scheduling work onto efficiency cores makes every step wait for the slowest thread.
1# Find physical cores2lscpu | grep -E "^CPU\(s\)|Core\(s\) per socket|Thread\(s\) per core" # Linux3sysctl -n hw.perflevel0.physicalcpu # macOS45# Then sweep, do not guess6for t in 4 6 8 12 16; do7 echo "threads=$t"8 ./build/bin/llama-bench -m model-Q4_K_M.gguf -p 512 -n 128 -t $t9doneThe sweep almost always shows a rise, a plateau, and then a decline. The peak is usually at physical core count or slightly below.
Memory configuration
Because bandwidth is the ceiling, the DIMM layout matters more than the DIMM speed rating. A single 32 GB stick in a dual-channel board runs at half bandwidth: two 16 GB sticks in the correct slots is a free doubling. On desktop platforms, enabling the XMP/EXPO profile in firmware often moves DDR5 from a JEDEC 4800 MT/s default to its rated 6000 MT/s, which is a direct 25% throughput gain on generation.
Two loading modes round out CPU tuning. Current builds set them with --load-mode (older builds had separate --mlock and --no-mmap flags). --load-mode mlock pins the model in RAM so the OS cannot page it out — essential if you are near your memory limit, dangerous if you are over it. --load-mode none reads the whole file into process memory instead of memory-mapping it; it costs a slower first load but avoids page-fault stalls on some filesystems and on network storage.
GPU inference
A GPU wins for two independent reasons: VRAM bandwidth is 4–20 times higher than system RAM, and its arithmetic throughput is 50–100 times higher, which collapses prefill.
Offload and the partial-offload cliff
The critical flag is -ngl (or n_gpu_layers). It sets how many of the model's transformer layers live in VRAM. If all of them fit, set it to 99 and stop tuning.
If they do not, the arithmetic tells you how many will. For Llama 3 8B at Q4_K_M: total weights 4.58 GiB, of which the embedding and output tensors are about 0.55 GiB, leaving 4.03 GiB across 32 layers, or 0.126 GiB per layer.
On a 6 GB card, usable VRAM is around 5.3 GiB. Reserve 0.5 GiB for a 4K KV cache and 0.3 GiB for compute buffers, and you have 4.5 GiB for weights. Put the 0.55 GiB of embeddings there too and 3.95 GiB remains: 3.95 / 0.126 = 31 layers.
But be careful what that buys. Partial offload is not proportional. With 28 of 32 layers on the GPU, every token still waits for four CPU layers, and those four dominate the step time:
| Layers on GPU (of 32) | Typical tokens/s | Comment |
|---|---|---|
| 0 | 10 | Pure CPU baseline |
| 16 | 16 | Half the work moved, 60% faster. Not 2× |
| 28 | 28 | 87% offloaded, still under a third of full speed |
| 32 | 95 | The last four layers are worth more than the first 28 |
This shape is the single most useful thing to know about GPU offload. If a smaller quantisation lets you reach 32/32, take it — a Q3_K_M running entirely on the GPU beats a Q4_K_M running 28/32 by a wide margin, and the quality difference between them is a fraction of the speed difference.
NVIDIA and AMD
NVIDIA is the smoothest path. Build llama.cpp with -DGGML_CUDA=ON, or install a prebuilt CUDA wheel of llama-cpp-python. Make sure flash attention is on (-fa on; the default, auto, enables it where supported): the fused attention kernel reduces both memory traffic and the size of intermediate buffers, which often matters more than its raw speed gain when you are near the VRAM limit.
AMD works through ROCm on Linux with -DGGML_HIP=ON, and through Vulkan everywhere else. ROCm is faster where it is supported; Vulkan is the universal fallback and now performs respectably. Check rocminfo for your card's gfx target before committing to ROCm, because official support is narrower than the marketing suggests, and HSA_OVERRIDE_GFX_VERSION is a common workaround for near-miss cards.
1nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv # NVIDIA2rocm-smi --showmeminfo vram # AMD34# Watch VRAM while a long generation runs — this is how you catch KV cache growth5watch -n 1 nvidia-smi --query-gpu=memory.used,utilization.gpu --format=csvVRAM budgeting
Combine the weight table with the KV cache. For Llama 3 8B, one token of context costs 2 × 32 layers × 8 KV heads × 128 dims × 2 bytes = 131,072 bytes, so 128 KiB per token, or 1 GiB per 8,192 tokens.
| Quantisation | Weights | + 4K ctx | + 16K ctx | + 32K ctx | Smallest card |
|---|---|---|---|---|---|
| Q4_K_M | 4.58 GiB | 5.6 | 7.1 | 9.1 | 8 GB (4K only) |
| Q5_K_M | 5.34 GiB | 6.3 | 7.8 | 9.8 | 8 GB (4K), 12 GB (32K) |
| Q6_K | 6.15 GiB | 7.2 | 8.7 | 10.7 | 12 GB |
| Q8_0 | 7.95 GiB | 9.0 | 10.5 | 12.5 | 12 GB (4K), 16 GB (32K) |
| F16 | 14.96 GiB | 16.0 | 17.5 | 19.5 | 24 GB |
Figures include roughly 0.5 GiB of compute buffers. When you are a gigabyte short, quantise the cache before you quantise the weights: --cache-type-k q8_0 --cache-type-v q8_0 halves the context column for a quality cost far below a step down in weight precision. (The V-cache setting requires flash attention.)
Apple Silicon
Apple's architecture is unusual in a way that suits this workload precisely: CPU and GPU share one pool of memory with one high-bandwidth bus. There is no PCIe copy, and no separate VRAM ceiling — a 64 GB Mac can hand roughly 48 GB to the GPU, which is more than any consumer graphics card offers.
llama.cpp builds with Metal enabled by default on macOS, and -ngl 99 works as it does on any GPU. The one system setting worth knowing is the GPU memory cap, which defaults to about 75% of physical RAM:
sudo sysctl iogpu.wired_limit_mb=57344 # allow 56 GiB on a 64 GB machineApple's MLX framework is the native alternative. It targets Metal directly, uses a lazy NumPy-like API, and is typically 10–25% faster than llama.cpp on the same Mac for the same model, with its own quantised format.
pip install mlx-lmmlx_lm.generate --model mlx-community/Meta-Llama-3.1-8B-Instruct-4bit \ --prompt "Explain memory bandwidth in two sentences." --max-tokens 200| llama.cpp with Metal | MLX | |
|---|---|---|
| Model availability | Everything, in GGUF | Smaller catalogue, growing |
| Speed on Apple Silicon | Baseline | 10–25% faster |
| Portability | Runs everywhere | macOS only |
| Fine-tuning | Limited | LoRA training built in |
| Use when | You want one setup across machines | You are Mac-only and want the last 20% |
The real constraint on a Mac is bandwidth by tier, not capacity. A base M3 has 100 GB/s, an M3 Pro 150, an M3 Max 400, an M2 Ultra 800. A 70B model at Q4_K_M is about 42 GB: give it enough memory to load but only 100 GB/s to read it through, and it generates at about two tokens per second (100 / 42 ≈ 2.4 at best). Capacity determines what runs; bandwidth determines whether you want to wait for it.
TPUs and other accelerators
Google's TPUs are systolic-array chips optimised for large dense matrix multiplication with high-bandwidth memory attached, reached through JAX or PyTorch/XLA on Cloud TPU or Colab. They are excellent at training and at high-batch serving, and a poor fit for what local deployment usually means.
The mismatch is structural. TPUs want big batches to fill the array; single-user interactive decoding runs at batch size one. They are cloud-hosted, so "local" no longer applies. And GGUF does not run on them at all — you need the model in a JAX or XLA-compatible form. If you are serving thousands of concurrent requests and can batch aggressively, they become interesting. For one person on one machine, they are the wrong shape.
Edge accelerators such as Coral Edge TPU or Hexagon DSPs have the opposite problem: a few megabytes to a few hundred megabytes of on-chip memory, against the several gigabytes an LLM needs. They serve small vision and audio models well. For language models below about 4 GB, a Raspberry Pi 5 running a 1–3B model on its CPU is usually the more practical edge answer.
Choosing a configuration for your machine
| Profile | Model and format | Flags | Expect |
|---|---|---|---|
| 8 GB RAM laptop, no GPU | 3B at Q4_K_M | -t 4 -c 2048 --load-mode mlock | 12–18 tok/s, slow prefill |
| 16 GB RAM laptop, no GPU | 8B at Q4_K_M | -t 6 -c 4096 | 8–12 tok/s; keep prompts short |
| 16 GB RAM, 8 GB VRAM | 8B at Q4_K_M | -ngl 99 -c 4096 -fa on | 50–70 tok/s; 4K context is the limit |
| 32 GB RAM, 12 GB VRAM | 8B at Q6_K or 14B at Q4_K_M | -ngl 99 -c 8192 -fa on | 45–70 tok/s with room for context |
| 64 GB RAM, 24 GB VRAM | 32B at Q4_K_M | -ngl 99 -c 16384 -fa on --cache-type-k q8_0 --cache-type-v q8_0 | 30–40 tok/s at genuinely strong quality |
| Apple Silicon, 36 GB | 32B at Q5_K_M | -ngl 99 -c 8192 | 15–25 tok/s on an M3 Max |
| Raspberry Pi 5, 8 GB | 1–3B at Q4_K_M | -t 4 -c 2048 | 3–6 tok/s |
Fit the whole model on the fastest memory you own, then spend what is left on context. A model that half-fits on a GPU is slower than a smaller one that fits completely.
Measuring instead of guessing
llama-bench separates the two regimes for you, reporting pp512 (prompt processing, tokens per second during prefill) and tg128 (text generation). Those are the two numbers that describe your machine:
./build/bin/llama-bench -m model-Q4_K_M.gguf -p 512 -n 128 -ngl 0,16,32 -r 3To measure inside your own application, time the phases separately — an average tokens-per-second figure that mixes prefill and decode hides exactly the information you need:
1import time2from llama_cpp import Llama34llm = Llama(model_path="model-Q4_K_M.gguf", n_ctx=4096, n_gpu_layers=-1, verbose=False)5prompt = "Summarise the causes of memory-bandwidth bottlenecks. " * 2067t0 = time.perf_counter()8first = None9n = 010for chunk in llm(prompt, max_tokens=200, stream=True):11 if first is None:12 first = time.perf_counter()13 n += 114end = time.perf_counter()1516print(f"time to first token : {first - t0:.2f} s")17print(f"decode rate : {(n - 1) / (end - first):.1f} tok/s")18print(f"predicted ceiling : {504e9 / 4.92e9:.0f} tok/s")Compare the measured decode rate against the bandwidth ceiling. Above roughly 60% of it, you are done — no flag will help, and only different hardware or a smaller model will. Far below it, something is wrong, and there are only a handful of candidates: layers still on the CPU, single-channel memory, thermal throttling, or a context window so large that the attention over the KV cache has itself become the bottleneck.
What to do with all this on Monday morning
Start every investigation by computing the ceiling. Look up your memory bandwidth, divide by your model file size, and write the number down. That single figure tells you whether the system is healthy before you touch a configuration file, and it converts a vague "it feels slow" into a testable claim.
When the measurement falls far short, check offload first, memory channels second, and everything else third. Roughly nine out of ten "my local model is unusably slow" reports are one of those two, and both are visible in seconds — nvidia-smi showing near-zero GPU memory in use, or a single DIMM in dmidecode.
When you are choosing hardware rather than tuning it, buy bandwidth and capacity, in that order, and treat compute as nearly irrelevant for single-user inference. A used card with 24 GB of moderately fast VRAM beats a new card with 12 GB of very fast VRAM for this workload, because the first runs a 32B model at full offload and the second does not run it at all. On Apple hardware the same logic points at the Max and Ultra tiers rather than at more memory on a base chip: memory you cannot feed quickly enough only lets you run models slowly.
And build the habit of recording pp512 and tg128 for every configuration you try, in a file, with the date. Backends change monthly, and a flag that cost you 15% last quarter may be a win now. Without a baseline you will be re-deriving the same conclusions every time you update.