Course Content
Local LLM Deployment and Quantization
3 sections · 7 lessons
Ollama, llama.cpp, and LM Studio
You have a MacBook with 16 GB of unified memory. You install the obvious thing, pip install transformers torch, and write four lines to load Llama 3 8B. The download crawls through 16.1 GB of safetensors files. Then Python allocates, the fans spin, the machine swaps, and about ninety seconds later the process is killed. No answer, no token, nothing.
The arithmetic explains it before you ever open a profiler. Llama 3 8B has 8.03 billion parameters. Loaded in the format it was published in, 16-bit floating point, each parameter costs 2 bytes:
8.03×109 parameters × 2 bytes = 1.606×1010 bytes = 16.06 GB = 14.96 GiB
That is the weights alone, on a machine where macOS and your browser have already taken 4–5 GiB, and before the model has stored a single token of conversation history. There was never room.
Now run ollama run llama3.1:8b on the same laptop. It downloads 4.9 GB, loads in a few seconds, and streams a reply at fifteen to twenty-five tokens per second, depending on the chip, while you keep your other apps open. Same model, same weights, same architecture. What changed is the runtime — the program that owns the memory — and the compressed weight format it knows how to read.
What a runtime actually does
A "local LLM runtime" sounds like a convenience wrapper. It is not. It is a memory manager with a matrix multiplier attached. Everything it does well, it does by controlling exactly how many bytes sit in RAM or VRAM at any moment.
Three things consume memory during inference, and you should be able to size all three from memory:
| Consumer | What it holds | How it scales |
|---|---|---|
| Weights | Every parameter of the model | Fixed. Parameters × bytes-per-parameter |
| KV cache | Keys and values for every token seen so far, so the model does not recompute them | Grows linearly with context length and with concurrent users |
| Compute buffers | Scratch space for the current forward pass — logits, intermediate activations | Roughly constant for one request; grows with batch size |
The whole discipline of local deployment is making those three sum to less than what you have. Llama 3 8B is the reference model for this entire topic.
Llama 3.1 8B dates from 2024. It is the reference here because every runtime supports it and its shapes are fully documented, not because it is the model to pick for a new project in 2026 — for that, look at the current Qwen, Gemma, Mistral and gpt-oss releases on the Ollama library or Hugging Face. The arithmetic is the same for all of them: read the parameter count, layer count, key/value heads and head dimension from the model's config.json. One twist: many current models are mixture-of-experts, where memory is set by the total parameter count but decode speed by the much smaller number of active parameters per token.
Here is the first term, the weights, for Llama 3 8B:
| Format | Bits per weight (effective) | File size | In GiB |
|---|---|---|---|
| FP32 | 32 | 32.1 GB | 29.9 |
| FP16 / BF16 | 16 | 16.1 GB | 14.96 |
| Q8_0 | 8.5 | 8.54 GB | 7.95 |
| Q6_K | 6.6 | 6.60 GB | 6.15 |
| Q5_K_M | 5.7 | 5.73 GB | 5.34 |
| Q4_K_M | 4.9 | 4.92 GB | 4.58 |
| Q3_K_M | 4.0 | 4.02 GB | 3.74 |
| Q2_K | 3.2 | 3.18 GB | 2.96 |
Those "bits per weight" figures are not the nominal bit width. A file labelled 4-bit stores 4.9 bits per weight on average, because each block of weights also carries a scale factor, and because a few sensitive tensors (the output layer, and about half of the attention-value and feed-forward-down projections) are deliberately kept at higher precision. Assume 0.5 bytes per parameter for a 4-bit model and you will under-budget by roughly a fifth.
Every local deployment question reduces to one inequality: weights + KV cache + compute buffers must fit in the memory you actually have, not the memory printed on the box.
The term people forget
The KV cache is where budgets die. For Llama 3 8B — 32 layers, 8 key/value heads after grouped-query attention, head dimension 128 — one token of context costs:
2 (keys and values) × 32 layers × 8 heads × 128 dims × 2 bytes = 131,072 bytes = 128 KiB per token
So an 8,192-token context is 8192 × 128 KiB = exactly 1 GiB. A 32,768-token context is 4 GiB. A 128K context is 16 GiB — larger than the FP16 weights of the model itself.
This is why an 8 GB graphics card that comfortably runs Llama 3 8B at Q4_K_M with a 4K context falls over the instant you paste in a long document. The weights did not change. The cache did.
llama.cpp: the engine underneath
Nearly every friendly local-LLM tool is llama.cpp, or its ggml tensor library, wearing a coat. Ollama grew out of it. LM Studio ships it. Dozens of desktop chat apps link against it. Understanding it means understanding all of them.
What it is and why it exists
llama.cpp is a C/C++ inference engine built on ggml, a small tensor library with no Python, no PyTorch, and no mandatory CUDA. Georgi Gerganov started it in March 2023 to run LLaMA on a MacBook. The design decisions that made that possible are the reasons it still dominates:
- Quantised weights as first-class citizens. Matrix multiplication kernels read 4-bit blocks directly and dequantise inside the register file. There is no "unpack to FP16 first" step eating memory.
- Memory-mapped model files. The weight file is
mmap-ed, so pages load lazily and the OS can evict them under pressure. Two processes running the same model share one copy in physical RAM. - Layer-granularity offload. You can put 20 of 32 layers on a GPU and leave 12 on the CPU. This is the single most useful feature for people whose VRAM is slightly too small.
- Every backend. CUDA, Metal, ROCm, Vulkan, SYCL, and plain AVX2/AVX-512 CPU, from one codebase.
Building and running it
1git clone https://github.com/ggml-org/llama.cpp2cd llama.cpp34# CPU only (Metal is enabled automatically on macOS)5cmake -B build6cmake --build build --config Release -j78# NVIDIA9cmake -B build -DGGML_CUDA=ON10cmake --build build --config Release -j1112# AMD (ROCm); see docs/build.md for setting your GPU target13cmake -B build -DGGML_HIP=ON14cmake --build build --config Release -jThe three binaries that matter are llama-cli for one-shot generation, llama-server for an HTTP endpoint, and llama-bench for measurement.
1# One-shot generation (-st: answer once and exit instead of staying in chat mode)2./build/bin/llama-cli \3 -m models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \4 -p "Explain grouped-query attention in three sentences." \5 -st -n 256 -c 4096 -ngl 99 -t 867# HTTP server on port 80808./build/bin/llama-server \9 -m models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \10 -c 8192 -ngl 99 --host 0.0.0.0 --port 80801112# Measure before you tune anything13./build/bin/llama-bench -m models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf -p 512 -n 128Those flags are worth learning properly, because they are the same knobs every wrapper exposes under different names.
| Flag | Meaning | How to set it |
|---|---|---|
-ngl N | Number of transformer layers offloaded to the GPU | 99 (or all) means every layer. Current builds default to auto, which offloads what fits; set it explicitly when you want to know |
-c N | Context length in tokens | The KV cache is sized from this. The default (0) means the model's full trained context — 128K for Llama 3.1 — which current builds shrink automatically if it will not fit. Set it yourself so you know what you got. Do not set 32768 if your prompts are 900 tokens |
-t N | CPU threads for generation | Physical performance cores, not logical threads. More is often slower |
-b N | Batch size for prompt processing | Larger speeds up prefill, costs compute-buffer memory |
--cache-type-k q8_0 | Quantise the KV cache itself | Add --cache-type-v q8_0 too; together they halve cache memory for a small quality cost. The rescue flags for long contexts. A quantised V cache needs flash attention |
-fa on | Fused (flash) attention kernel | Lowers memory and raises speed. The default, auto, turns it on where the backend supports it |
Where the performance comes from
Once all layers sit on the GPU, generation speed is governed almost entirely by memory bandwidth, because producing each token requires reading every weight once. On a card with 1,008 GB/s of bandwidth running a 4.92 GB model, the ceiling is 1008 / 4.92 ≈ 205 tokens per second, and real throughput lands around 70% of that. Adding threads, changing batch size, or buying a faster CPU will not move a number set by bandwidth.
Ollama: the runtime as a service
Ollama began as llama.cpp plus the four things that make it feel like a product: a background daemon, a model registry, automatic hardware configuration, and a stable HTTP API. It now runs many models on its own engine, built on the same ggml library, and on Apple Silicon recent versions can also run some models through Apple's MLX (tags ending in -mlx). The memory arithmetic in this lesson applies either way.
1curl -fsSL https://ollama.com/install.sh | sh # Linux2# macOS and Windows: download the installer from ollama.com34ollama serve # start the daemon (the app does this for you)5ollama pull llama3.1:8b # download6ollama run llama3.1:8b # interactive chat7ollama list # what is on disk8ollama ps # what is loaded in memory right now9ollama rm llama3.1:8b # reclaim the spaceRead the tags carefully, because they encode exactly the arithmetic above. In llama3.1:8b-instruct-q4_K_M: llama3.1 is the family, 8b the parameter count, instruct the chat-tuned variant rather than the raw base model, and q4_K_M the quantisation. A bare tag such as llama3.1:8b silently resolves to the q4_K_M instruct build, which is why the download is 4.9 GB and not 16.
The daemon exposes a REST API on port 11434.
1curl http://localhost:11434/api/generate -d '{2 "model": "llama3.1:8b",3 "prompt": "Name three causes of high tail latency.",4 "stream": false,5 "options": { "temperature": 0.3, "num_ctx": 4096, "num_predict": 200 }6}'1import ollama23resp = ollama.chat(4 model="llama3.1:8b",5 messages=[6 {"role": "system", "content": "You answer in one paragraph."},7 {"role": "user", "content": "Why does the KV cache grow with context?"},8 ],9 options={"temperature": 0.3, "num_ctx": 8192},10)11print(resp["message"]["content"])1213# Streaming: yields chunks as they are produced14for chunk in ollama.chat(model="llama3.1:8b",15 messages=[{"role": "user", "content": "Count to five."}],16 stream=True):17 print(chunk["message"]["content"], end="", flush=True)The environment variables are where you regain the control that the friendly interface hides:
| Variable | Default | Why you would change it |
|---|---|---|
OLLAMA_KEEP_ALIVE | 5m | Set to -1 to pin a model in memory and never pay reload latency; set to 0 to free memory immediately after each request |
OLLAMA_NUM_PARALLEL | 1 | Concurrent requests per model. Each slot needs its own KV cache slice |
OLLAMA_MAX_LOADED_MODELS | 3 per GPU (3 on CPU) | Stop a second model evicting the first, or allow it deliberately |
OLLAMA_KV_CACHE_TYPE | f16 | q8_0 halves KV cache memory — the fix for long contexts on small cards. Needs flash attention, which Ollama turns on automatically where supported |
OLLAMA_HOST | 127.0.0.1:11434 | Bind to 0.0.0.0 to serve other machines. There is no authentication, so put it behind something |
Ollama stores weights as content-addressed blobs under ~/.ollama/models, so two tags sharing a base layer share the bytes on disk. A Modelfile lets you bake a system prompt, sampling defaults, or your own GGUF file into a named model:
1cat > Modelfile <<'EOF'2FROM llama3.1:8b3PARAMETER temperature 0.24PARAMETER num_ctx 81925SYSTEM "You are a terse code reviewer. Reply with findings only."6EOF78ollama create reviewer -f Modelfile9ollama run reviewerLM Studio: the graphical route
LM Studio is a desktop application for macOS, Windows, and Linux. It bundles llama.cpp (and Apple's MLX engine on Apple Silicon) behind a chat window, a model browser that reads Hugging Face directly, and a one-click OpenAI-compatible server.
Its genuinely distinctive feature is the model browser's fit estimate. Next to each quantisation it tells you whether that file will fit fully in your GPU, fit partially, or not fit — which is the memory inequality above, computed for you. For someone learning what fits on their own hardware, that feedback loop is worth more than any amount of documentation.
Once the local server is started, anything written against the OpenAI SDK works unchanged:
1from openai import OpenAI23client = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio")45resp = client.chat.completions.create(6 model="llama-3.1-8b-instruct", # the model identifier LM Studio shows7 messages=[{"role": "user", "content": "Summarise TCP slow start."}],8 temperature=0.3,9 max_tokens=300,10)11print(resp.choices[0].message.content)The cost of the graphical route is that it is a desktop app. It wants a logged-in session and a window manager. It is not what you run headless on a server.
Choosing between them
| llama.cpp | Ollama | LM Studio | |
|---|---|---|---|
| Interface | CLI and HTTP | CLI, HTTP, daemon | GUI plus HTTP |
| Setup effort | Compile it yourself | One installer | One installer |
| Control over flags | Total | Most, via options and env vars | Sliders for the common ones |
| Model acquisition | Find the GGUF yourself | Curated registry, ollama pull | Hugging Face browser with fit estimates |
| API shape | Own API plus OpenAI-compatible | Own API plus OpenAI-compatible | OpenAI-compatible |
| Headless server | Yes, ideal | Yes, ideal | Awkward |
| Custom builds and new backends | Yes, day one | When Ollama updates its bundled engine | When the app updates |
| Best at | Squeezing the machine, embedding in your own binary | Serving models to applications | Trying models, learning what fits |
A decision rule that holds up in practice:
- Building an application that calls a model? Ollama. The daemon, the model management, and the stable API are exactly the boring infrastructure you do not want to write.
- Fighting for the last 20% of speed, or shipping a product that embeds inference? llama.cpp directly. You need
-ngl,--cache-type-k, and the ability to rebuild with a new backend the week it lands. - Evaluating five models on a laptop to see which is good enough? LM Studio. Clicking through quantisations while watching the fit estimate teaches the memory budget faster than any table.
These are not competitors so much as three depths of the same stack: LM Studio to explore, Ollama to serve, llama.cpp when you need to see the machine.
What fits on what
Applying the inequality with a 4K context (0.5 GiB of KV cache) and roughly 0.5 GiB of compute buffers:
| Hardware | Usable memory | Comfortable choice | Reality |
|---|---|---|---|
| 8 GB RAM laptop, CPU only | ~4 GiB | 3B at Q4_K_M (~1.9 GiB) | Works. An 8B at Q4 will thrash the page cache |
| 16 GB RAM laptop, CPU only | ~11 GiB | 8B at Q4_K_M (4.58 + 1.0 = 5.6 GiB) | Comfortable, but 8–12 tokens/s at best |
| 8 GB VRAM GPU | ~7.2 GiB | 8B at Q4_K_M, 4K context | Fits at 5.6 GiB. At 32K context the cache alone is 4 GiB and it fails |
| 12 GB VRAM GPU | ~11 GiB | 8B at Q6_K, or 14B at Q4_K_M | 8B Q6_K: 6.15 + 0.5 + 0.5 = 7.2 GiB, with room for 16K context (8.7 GiB) |
| 24 GB VRAM GPU | ~23 GiB | 8B at Q8_0, or 32B at Q4_K_M | A 70B at Q4_K_M is 40 GiB and does not fit |
| Apple Silicon, 36 GB unified | ~27 GiB (GPU cap) | 32B at Q5_K_M | Unified memory is the advantage; bandwidth is the limit |
Where people go wrong
"Ollama is a different, smaller model." It is the same weights, compressed. ollama run llama3.1:8b and a Hugging Face FP16 checkpoint are the same 8.03 billion parameters; one stores each at 4.9 bits and the other at 16.
"It fits, because the file is smaller than my VRAM." This is the most expensive mistake in the topic. A 4.58 GiB model on an 8 GB card looks like 3.4 GiB of headroom, until you set a 32K context and the KV cache asks for 4 GiB. Budget the cache first; the file size is only one of three terms.
"More threads means faster." Generation is memory-bandwidth-bound, not compute-bound. Setting -t 32 on a 16-core machine usually makes things slower, because hyperthreads contend for the same cache lines and memory controller. Start at the number of physical performance cores and measure.
"Adding RAM will speed up my GPU." Once every layer is offloaded, system RAM is irrelevant to generation speed. The only thing that helps is more VRAM bandwidth or a smaller model.
"Partial offload is nearly as good." It is not, and the drop is sharp. With 28 of 32 layers on the GPU, the 4 CPU layers dominate the per-token time because the whole token waits for the slowest stage. Going from 32/32 to 28/32 offload can halve throughput. If you cannot fit every layer, a smaller quantisation that does fit almost always wins.
Setting up so you can change your mind later
The practical lesson is that the runtime is a swappable component, and you should keep it that way. Every one of these tools speaks an OpenAI-compatible HTTP interface — Ollama on 127.0.0.1:11434/v1, LM Studio on 127.0.0.1:1234/v1, llama-server on whatever port you gave it. If your application talks to a base URL from configuration rather than importing a vendor SDK, then moving from LM Studio on your laptop to llama-server on a rented GPU box is an environment-variable change.
Then do the sizing before you download anything. Write down your usable memory, subtract the KV cache for the longest context you actually need, subtract half a gigabyte of buffers, and pick the largest quantisation that fits what remains. On a 12 GB card wanting 16K of context, that is 11 − 2 − 0.5 = 8.5 GiB of budget for weights, which comfortably takes an 8B at Q6_K and rules out a 14B at Q6_K. Doing that sum takes a minute and saves an evening of watching downloads finish and then fail to load.
Finally, measure once with llama-bench or a simple timed loop before tuning anything. Most local-inference tuning advice is written by people optimising a different bottleneck than yours, and a single number — tokens per second at your context length, on your machine — tells you whether you are bandwidth-bound, compute-bound, or accidentally running half the model on the CPU.