Local LLM Deployment and Quantization

Ollama, llama.cpp, and LM Studio


You have a MacBook with 16 GB of unified memory. You install the obvious thing, pip install transformers torch, and write four lines to load Llama 3 8B. The download crawls through 16.1 GB of safetensors files. Then Python allocates, the fans spin, the machine swaps, and about ninety seconds later the process is killed. No answer, no token, nothing.

The arithmetic explains it before you ever open a profiler. Llama 3 8B has 8.03 billion parameters. Loaded in the format it was published in, 16-bit floating point, each parameter costs 2 bytes:

8.03×1098.03 \times 10^{9} parameters × 2 bytes = 1.606×10101.606 \times 10^{10} bytes = 16.06 GB = 14.96 GiB

That is the weights alone, on a machine where macOS and your browser have already taken 4–5 GiB, and before the model has stored a single token of conversation history. There was never room.

Now run ollama run llama3.1:8b on the same laptop. It downloads 4.9 GB, loads in a few seconds, and streams a reply at fifteen to twenty-five tokens per second, depending on the chip, while you keep your other apps open. Same model, same weights, same architecture. What changed is the runtime — the program that owns the memory — and the compressed weight format it knows how to read.

Three tools, one engine underneathLM Studio or Ollama — the interface you installllama.cpp, or Ollama's own engine — runs the graphGGML — quantised tensors and their kernelsMetal, CUDA or AVX2 on your actual chip
Whichever of the three you pick, the speed mostly comes from the same ggml kernels reading the same memory — the choice is about workflow.

What a runtime actually does

A "local LLM runtime" sounds like a convenience wrapper. It is not. It is a memory manager with a matrix multiplier attached. Everything it does well, it does by controlling exactly how many bytes sit in RAM or VRAM at any moment.

Three things consume memory during inference, and you should be able to size all three from memory:

ConsumerWhat it holdsHow it scales
WeightsEvery parameter of the modelFixed. Parameters × bytes-per-parameter
KV cacheKeys and values for every token seen so far, so the model does not recompute themGrows linearly with context length and with concurrent users
Compute buffersScratch space for the current forward pass — logits, intermediate activationsRoughly constant for one request; grows with batch size

The whole discipline of local deployment is making those three sum to less than what you have. Llama 3 8B is the reference model for this entire topic.

Llama 3.1 8B dates from 2024. It is the reference here because every runtime supports it and its shapes are fully documented, not because it is the model to pick for a new project in 2026 — for that, look at the current Qwen, Gemma, Mistral and gpt-oss releases on the Ollama library or Hugging Face. The arithmetic is the same for all of them: read the parameter count, layer count, key/value heads and head dimension from the model's config.json. One twist: many current models are mixture-of-experts, where memory is set by the total parameter count but decode speed by the much smaller number of active parameters per token.

Here is the first term, the weights, for Llama 3 8B:

FormatBits per weight (effective)File sizeIn GiB
FP323232.1 GB29.9
FP16 / BF161616.1 GB14.96
Q8_08.58.54 GB7.95
Q6_K6.66.60 GB6.15
Q5_K_M5.75.73 GB5.34
Q4_K_M4.94.92 GB4.58
Q3_K_M4.04.02 GB3.74
Q2_K3.23.18 GB2.96

Those "bits per weight" figures are not the nominal bit width. A file labelled 4-bit stores 4.9 bits per weight on average, because each block of weights also carries a scale factor, and because a few sensitive tensors (the output layer, and about half of the attention-value and feed-forward-down projections) are deliberately kept at higher precision. Assume 0.5 bytes per parameter for a 4-bit model and you will under-budget by roughly a fifth.

Every local deployment question reduces to one inequality: weights + KV cache + compute buffers must fit in the memory you actually have, not the memory printed on the box.

The term people forget

The KV cache is where budgets die. For Llama 3 8B — 32 layers, 8 key/value heads after grouped-query attention, head dimension 128 — one token of context costs:

2 (keys and values) × 32 layers × 8 heads × 128 dims × 2 bytes = 131,072 bytes = 128 KiB per token

So an 8,192-token context is 8192 × 128 KiB = exactly 1 GiB. A 32,768-token context is 4 GiB. A 128K context is 16 GiB — larger than the FP16 weights of the model itself.

This is why an 8 GB graphics card that comfortably runs Llama 3 8B at Q4_K_M with a 4K context falls over the instant you paste in a long document. The weights did not change. The cache did.

llama.cpp: the engine underneath

Nearly every friendly local-LLM tool is llama.cpp, or its ggml tensor library, wearing a coat. Ollama grew out of it. LM Studio ships it. Dozens of desktop chat apps link against it. Understanding it means understanding all of them.

What it is and why it exists

llama.cpp is a C/C++ inference engine built on ggml, a small tensor library with no Python, no PyTorch, and no mandatory CUDA. Georgi Gerganov started it in March 2023 to run LLaMA on a MacBook. The design decisions that made that possible are the reasons it still dominates:

  • Quantised weights as first-class citizens. Matrix multiplication kernels read 4-bit blocks directly and dequantise inside the register file. There is no "unpack to FP16 first" step eating memory.
  • Memory-mapped model files. The weight file is mmap-ed, so pages load lazily and the OS can evict them under pressure. Two processes running the same model share one copy in physical RAM.
  • Layer-granularity offload. You can put 20 of 32 layers on a GPU and leave 12 on the CPU. This is the single most useful feature for people whose VRAM is slightly too small.
  • Every backend. CUDA, Metal, ROCm, Vulkan, SYCL, and plain AVX2/AVX-512 CPU, from one codebase.

Building and running it

Bash
git clone https://github.com/ggml-org/llama.cppcd llama.cpp# CPU only (Metal is enabled automatically on macOS)cmake -B buildcmake --build build --config Release -j# NVIDIAcmake -B build -DGGML_CUDA=ONcmake --build build --config Release -j# AMD (ROCm); see docs/build.md for setting your GPU targetcmake -B build -DGGML_HIP=ONcmake --build build --config Release -j

The three binaries that matter are llama-cli for one-shot generation, llama-server for an HTTP endpoint, and llama-bench for measurement.

Bash
# One-shot generation (-st: answer once and exit instead of staying in chat mode)./build/bin/llama-cli \  -m models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \  -p "Explain grouped-query attention in three sentences." \  -st -n 256 -c 4096 -ngl 99 -t 8# HTTP server on port 8080./build/bin/llama-server \  -m models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \  -c 8192 -ngl 99 --host 0.0.0.0 --port 8080# Measure before you tune anything./build/bin/llama-bench -m models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf -p 512 -n 128

Those flags are worth learning properly, because they are the same knobs every wrapper exposes under different names.

FlagMeaningHow to set it
-ngl NNumber of transformer layers offloaded to the GPU99 (or all) means every layer. Current builds default to auto, which offloads what fits; set it explicitly when you want to know
-c NContext length in tokensThe KV cache is sized from this. The default (0) means the model's full trained context — 128K for Llama 3.1 — which current builds shrink automatically if it will not fit. Set it yourself so you know what you got. Do not set 32768 if your prompts are 900 tokens
-t NCPU threads for generationPhysical performance cores, not logical threads. More is often slower
-b NBatch size for prompt processingLarger speeds up prefill, costs compute-buffer memory
--cache-type-k q8_0Quantise the KV cache itselfAdd --cache-type-v q8_0 too; together they halve cache memory for a small quality cost. The rescue flags for long contexts. A quantised V cache needs flash attention
-fa onFused (flash) attention kernelLowers memory and raises speed. The default, auto, turns it on where the backend supports it

Where the performance comes from

Once all layers sit on the GPU, generation speed is governed almost entirely by memory bandwidth, because producing each token requires reading every weight once. On a card with 1,008 GB/s of bandwidth running a 4.92 GB model, the ceiling is 1008 / 4.92 ≈ 205 tokens per second, and real throughput lands around 70% of that. Adding threads, changing batch size, or buying a faster CPU will not move a number set by bandwidth.

Ollama: the runtime as a service

Ollama began as llama.cpp plus the four things that make it feel like a product: a background daemon, a model registry, automatic hardware configuration, and a stable HTTP API. It now runs many models on its own engine, built on the same ggml library, and on Apple Silicon recent versions can also run some models through Apple's MLX (tags ending in -mlx). The memory arithmetic in this lesson applies either way.

Bash
curl -fsSL https://ollama.com/install.sh | sh   # Linux# macOS and Windows: download the installer from ollama.comollama serve                 # start the daemon (the app does this for you)ollama pull llama3.1:8b      # downloadollama run llama3.1:8b       # interactive chatollama list                  # what is on diskollama ps                    # what is loaded in memory right nowollama rm llama3.1:8b        # reclaim the space

Read the tags carefully, because they encode exactly the arithmetic above. In llama3.1:8b-instruct-q4_K_M: llama3.1 is the family, 8b the parameter count, instruct the chat-tuned variant rather than the raw base model, and q4_K_M the quantisation. A bare tag such as llama3.1:8b silently resolves to the q4_K_M instruct build, which is why the download is 4.9 GB and not 16.

The daemon exposes a REST API on port 11434.

Bash
curl http://localhost:11434/api/generate -d '{  "model": "llama3.1:8b",  "prompt": "Name three causes of high tail latency.",  "stream": false,  "options": { "temperature": 0.3, "num_ctx": 4096, "num_predict": 200 }}'
Python
import ollamaresp = ollama.chat(    model="llama3.1:8b",    messages=[        {"role": "system", "content": "You answer in one paragraph."},        {"role": "user", "content": "Why does the KV cache grow with context?"},    ],    options={"temperature": 0.3, "num_ctx": 8192},)print(resp["message"]["content"])# Streaming: yields chunks as they are producedfor chunk in ollama.chat(model="llama3.1:8b",                         messages=[{"role": "user", "content": "Count to five."}],                         stream=True):    print(chunk["message"]["content"], end="", flush=True)

The environment variables are where you regain the control that the friendly interface hides:

VariableDefaultWhy you would change it
OLLAMA_KEEP_ALIVE5mSet to -1 to pin a model in memory and never pay reload latency; set to 0 to free memory immediately after each request
OLLAMA_NUM_PARALLEL1Concurrent requests per model. Each slot needs its own KV cache slice
OLLAMA_MAX_LOADED_MODELS3 per GPU (3 on CPU)Stop a second model evicting the first, or allow it deliberately
OLLAMA_KV_CACHE_TYPEf16q8_0 halves KV cache memory — the fix for long contexts on small cards. Needs flash attention, which Ollama turns on automatically where supported
OLLAMA_HOST127.0.0.1:11434Bind to 0.0.0.0 to serve other machines. There is no authentication, so put it behind something

Ollama stores weights as content-addressed blobs under ~/.ollama/models, so two tags sharing a base layer share the bytes on disk. A Modelfile lets you bake a system prompt, sampling defaults, or your own GGUF file into a named model:

Bash
cat > Modelfile <<'EOF'FROM llama3.1:8bPARAMETER temperature 0.2PARAMETER num_ctx 8192SYSTEM "You are a terse code reviewer. Reply with findings only."EOFollama create reviewer -f Modelfileollama run reviewer

LM Studio: the graphical route

LM Studio is a desktop application for macOS, Windows, and Linux. It bundles llama.cpp (and Apple's MLX engine on Apple Silicon) behind a chat window, a model browser that reads Hugging Face directly, and a one-click OpenAI-compatible server.

Its genuinely distinctive feature is the model browser's fit estimate. Next to each quantisation it tells you whether that file will fit fully in your GPU, fit partially, or not fit — which is the memory inequality above, computed for you. For someone learning what fits on their own hardware, that feedback loop is worth more than any amount of documentation.

Once the local server is started, anything written against the OpenAI SDK works unchanged:

Python
from openai import OpenAIclient = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio")resp = client.chat.completions.create(    model="llama-3.1-8b-instruct",             # the model identifier LM Studio shows    messages=[{"role": "user", "content": "Summarise TCP slow start."}],    temperature=0.3,    max_tokens=300,)print(resp.choices[0].message.content)

The cost of the graphical route is that it is a desktop app. It wants a logged-in session and a window manager. It is not what you run headless on a server.

Choosing between them

llama.cppOllamaLM Studio
InterfaceCLI and HTTPCLI, HTTP, daemonGUI plus HTTP
Setup effortCompile it yourselfOne installerOne installer
Control over flagsTotalMost, via options and env varsSliders for the common ones
Model acquisitionFind the GGUF yourselfCurated registry, ollama pullHugging Face browser with fit estimates
API shapeOwn API plus OpenAI-compatibleOwn API plus OpenAI-compatibleOpenAI-compatible
Headless serverYes, idealYes, idealAwkward
Custom builds and new backendsYes, day oneWhen Ollama updates its bundled engineWhen the app updates
Best atSqueezing the machine, embedding in your own binaryServing models to applicationsTrying models, learning what fits

A decision rule that holds up in practice:

  • Building an application that calls a model? Ollama. The daemon, the model management, and the stable API are exactly the boring infrastructure you do not want to write.
  • Fighting for the last 20% of speed, or shipping a product that embeds inference? llama.cpp directly. You need -ngl, --cache-type-k, and the ability to rebuild with a new backend the week it lands.
  • Evaluating five models on a laptop to see which is good enough? LM Studio. Clicking through quantisations while watching the fit estimate teaches the memory budget faster than any table.

These are not competitors so much as three depths of the same stack: LM Studio to explore, Ollama to serve, llama.cpp when you need to see the machine.

What fits on what

Applying the inequality with a 4K context (0.5 GiB of KV cache) and roughly 0.5 GiB of compute buffers:

HardwareUsable memoryComfortable choiceReality
8 GB RAM laptop, CPU only~4 GiB3B at Q4_K_M (~1.9 GiB)Works. An 8B at Q4 will thrash the page cache
16 GB RAM laptop, CPU only~11 GiB8B at Q4_K_M (4.58 + 1.0 = 5.6 GiB)Comfortable, but 8–12 tokens/s at best
8 GB VRAM GPU~7.2 GiB8B at Q4_K_M, 4K contextFits at 5.6 GiB. At 32K context the cache alone is 4 GiB and it fails
12 GB VRAM GPU~11 GiB8B at Q6_K, or 14B at Q4_K_M8B Q6_K: 6.15 + 0.5 + 0.5 = 7.2 GiB, with room for 16K context (8.7 GiB)
24 GB VRAM GPU~23 GiB8B at Q8_0, or 32B at Q4_K_MA 70B at Q4_K_M is 40 GiB and does not fit
Apple Silicon, 36 GB unified~27 GiB (GPU cap)32B at Q5_K_MUnified memory is the advantage; bandwidth is the limit

Where people go wrong

"Ollama is a different, smaller model." It is the same weights, compressed. ollama run llama3.1:8b and a Hugging Face FP16 checkpoint are the same 8.03 billion parameters; one stores each at 4.9 bits and the other at 16.

"It fits, because the file is smaller than my VRAM." This is the most expensive mistake in the topic. A 4.58 GiB model on an 8 GB card looks like 3.4 GiB of headroom, until you set a 32K context and the KV cache asks for 4 GiB. Budget the cache first; the file size is only one of three terms.

"More threads means faster." Generation is memory-bandwidth-bound, not compute-bound. Setting -t 32 on a 16-core machine usually makes things slower, because hyperthreads contend for the same cache lines and memory controller. Start at the number of physical performance cores and measure.

"Adding RAM will speed up my GPU." Once every layer is offloaded, system RAM is irrelevant to generation speed. The only thing that helps is more VRAM bandwidth or a smaller model.

"Partial offload is nearly as good." It is not, and the drop is sharp. With 28 of 32 layers on the GPU, the 4 CPU layers dominate the per-token time because the whole token waits for the slowest stage. Going from 32/32 to 28/32 offload can halve throughput. If you cannot fit every layer, a smaller quantisation that does fit almost always wins.

Setting up so you can change your mind later

The practical lesson is that the runtime is a swappable component, and you should keep it that way. Every one of these tools speaks an OpenAI-compatible HTTP interface — Ollama on 127.0.0.1:11434/v1, LM Studio on 127.0.0.1:1234/v1, llama-server on whatever port you gave it. If your application talks to a base URL from configuration rather than importing a vendor SDK, then moving from LM Studio on your laptop to llama-server on a rented GPU box is an environment-variable change.

Then do the sizing before you download anything. Write down your usable memory, subtract the KV cache for the longest context you actually need, subtract half a gigabyte of buffers, and pick the largest quantisation that fits what remains. On a 12 GB card wanting 16K of context, that is 11 − 2 − 0.5 = 8.5 GiB of budget for weights, which comfortably takes an 8B at Q6_K and rules out a 14B at Q6_K. Doing that sum takes a minute and saves an evening of watching downloads finish and then fail to load.

Finally, measure once with llama-bench or a simple timed loop before tuning anything. Most local-inference tuning advice is written by people optimising a different bottleneck than yours, and a single number — tokens per second at your context length, on your machine — tells you whether you are bandwidth-bound, compute-bound, or accidentally running half the model on the CPU.