Course Content
Local LLM Deployment and Quantization
3 sections · 7 lessons
Mini Project: Deploy an 8B Model Locally with Ollama
Four hours from now you should have an 8-billion-parameter model answering questions on your own machine, a Python client that streams tokens and manages conversation history, a benchmark file containing your machine's real time-to-first-token and decode rate at three prompt lengths, and enough evidence to tell someone whether this hardware is good enough for a real workload.
The single most common way this build fails is skipping the arithmetic at the start. Someone downloads the largest quantisation their disk will hold, sets a 32,768-token context because larger sounds better, and gets CUDA out of memory — or worse, no error at all, just an inexplicable four tokens per second because the runtime quietly fell back to the CPU. That failure is entirely predictable from three numbers you can compute in two minutes.
So the build starts with sizing, not downloading.
What you are building
| Artefact | Done when |
|---|---|
| A running model | ollama run gives a sensible answer in under 10 seconds cold, under 3 warm |
client.py | Streams tokens, keeps multi-turn context, trims history without dropping the system prompt |
benchmark.py and results.json | TTFT and decode rate at three prompt lengths, median of at least five timed runs after warm-up |
| A quantisation comparison | Two quantisations measured on both speed and one quality check of your choosing |
| A short written report | Your hardware, your numbers, and a recommendation with a reason |
The success criterion that matters is the last one. Anyone can make a model produce text. The point of the exercise is being able to say "on this machine, an 8B at Q4_K_M gives 0.3-second TTFT and 62 tokens per second at 4K context, which is fine for chat and unusable for 20-page documents" — and to have the measurements behind it.
Phase one: size the machine before downloading anything
Three things consume memory during inference, and you can compute all three in advance.
Weights. An 8B model at Q4_K_M is a 4.92 GB file, which is 4.58 GiB in memory. Note that "4-bit" does not mean half a byte per parameter — the real figure is about 4.9 bits, because every block of weights carries a scale factor and a few sensitive tensors, such as the output layer, are kept at higher precision.
KV cache. This is the term people forget, and it is the one that kills builds. For Llama 3 8B — 32 layers, 8 key/value heads, head dimension 128, 16-bit cache — one token of context costs:
2 (K and V) × 32 × 8 × 128 × 2 bytes = 131,072 bytes = 128 KiB per token
So 4,096 tokens is 0.5 GiB, 8,192 tokens is exactly 1 GiB, and 32,768 tokens is 4 GiB — nearly as much as the model itself. It is allocated when the server starts, not when you send a long prompt.
Compute buffers. Budget 0.4–0.6 GiB.
Now find your usable memory and check the sum:
1nvidia-smi --query-gpu=name,memory.total --format=csv # NVIDIA VRAM2rocm-smi --showmeminfo vram # AMD VRAM3sysctl hw.memsize # macOS unified memory4free -g # Linux system RAMSubtract roughly 10% for driver and display overhead on a GPU, or 4–5 GiB for the operating system on a CPU-only machine.
| Your hardware | Usable | Pull this | Context | The sum |
|---|---|---|---|---|
| 8 GB VRAM | ~7.2 GiB | llama3.1:8b (Q4_K_M) | 4096 | 4.58 + 0.5 + 0.4 = 5.5 ✓ |
| 12 GB VRAM | ~11 GiB | llama3.1:8b-instruct-q6_K | 8192 | 6.15 + 1.0 + 0.5 = 7.7 ✓ |
| 24 GB VRAM | ~23 GiB | llama3.1:8b-instruct-q8_0 | 32768 | 7.95 + 4.0 + 0.6 = 12.6 ✓ |
| 6 GB VRAM | ~5.3 GiB | llama3.1:8b-instruct-q3_K_M | 4096 | 3.74 + 0.5 + 0.4 = 4.6 ✓ |
| 16 GB RAM, no GPU | ~11 GiB | llama3.1:8b (Q4_K_M) | 4096 | 5.5 ✓, but expect 10 tok/s |
| 8 GB RAM, no GPU | ~4 GiB | llama3.2:3b | 2048 | 1.9 + 0.2 + 0.3 = 2.4 ✓ |
| Apple Silicon 16 GB | ~11 GiB | llama3.1:8b | 8192 | 4.58 + 1.0 + 0.5 = 6.1 ✓ |
Write your own sum down before continuing. If it does not fit, the fix in order of preference is: reduce the context, quantise the KV cache to 8-bit, then drop a quantisation level. Dropping context is nearly free; dropping weight precision costs real quality.
Weights plus KV cache plus buffers must fit in memory you actually have. Every "why is my local model broken" question I have ever seen starts with someone not doing this sum.
Phase two: install and verify
1curl -fsSL https://ollama.com/install.sh | sh # Linux2# macOS / Windows: download the installer from ollama.com34ollama --version5curl -s http://localhost:11434/api/version # is the daemon up?6sudo systemctl status ollama # Linux service checkIf the daemon is not running, start it in a terminal you can watch with ollama serve. The log output during the first model load is genuinely informative — it tells you which backend was selected and how many layers were offloaded, and if it says offloaded 0/33 layers to GPU, you have found your problem before it becomes mysterious.
Set the environment before starting the daemon, not after:
1export OLLAMA_KEEP_ALIVE=-1 # never unload; removes the cold-start stall2export OLLAMA_MAX_LOADED_MODELS=1 # stop a second model evicting the first3export OLLAMA_FLASH_ATTENTION=1 # force flash attention on (recent versions enable it automatically where supported)4ollama servePhase three: pull the model and read its tag
ollama pull llama3.1:8bollama listollama ps # confirms it is resident, and its actual memory footprintRead the tag properly, because it encodes the arithmetic from phase one. In llama3.1:8b-instruct-q4_K_M: the family, the parameter count, the instruction-tuned variant rather than the raw base model, and the quantisation. A bare llama3.1:8b resolves to exactly that, which is why the download is 4.9 GB rather than 16.
The project uses llama3.1:8b because every number in this lesson is computed for it. Once the pipeline works, swap in a current model from the Ollama library and redo the phase-one sum with its layer count, key/value heads and file size.
Pull a second quantisation now, while you are here — you will compare them later:
ollama pull llama3.1:8b-instruct-q8_0 # only if your sum from phase one allows itollama run llama3.1:8b "Give three reasons a local model might run slowly."Then verify the offload immediately, because this is where silent failure lives:
1ollama ps2# Look at the PROCESSOR column. "100% GPU" is what you want.3# "48%/52% CPU/GPU" means partial offload and roughly a third of full speed.4nvidia-smi --query-gpu=memory.used --format=csvPartial offload is far worse than it sounds. With 28 of 32 layers on the GPU, every token still waits for four CPU layers, and throughput lands nearer a third of full-offload speed than seven-eighths. If you see it, reduce context or drop one quantisation level until you reach 100% GPU.
A model that half-fits on the GPU is slower than a smaller one that fits entirely. The last four layers are worth more than the first twenty-eight.
Phase four: a client worth keeping
1import json, requests23BASE = "http://localhost:11434"45def stream_chat(messages, model="llama3.1:8b", temperature=0.3, num_ctx=4096):6 """Yield content pieces as they are generated."""7 with requests.post(f"{BASE}/api/chat", stream=True, timeout=600,8 json={"model": model, "messages": messages, "stream": True,9 "keep_alive": -1,10 "options": {"temperature": temperature, "num_ctx": num_ctx}}) as r:11 r.raise_for_status()12 for line in r.iter_lines():13 if not line:14 continue15 obj = json.loads(line)16 if obj.get("done"):17 return18 yield obj["message"]["content"]192021class Conversation:22 """The server is stateless. History management is your job."""23 def __init__(self, system, model="llama3.1:8b", max_pairs=10):24 self.system = {"role": "system", "content": system}25 self.model, self.max_pairs = model, max_pairs26 self.turns = []2728 def ask(self, text):29 self.turns.append({"role": "user", "content": text})30 parts = []31 for piece in stream_chat([self.system] + self.turns, model=self.model):32 parts.append(piece)33 print(piece, end="", flush=True)34 print()35 self.turns.append({"role": "assistant", "content": "".join(parts)})36 if len(self.turns) > self.max_pairs * 2: # never drop the system message37 self.turns = self.turns[-self.max_pairs * 2:]38 return self.turns[-1]["content"]394041if __name__ == "__main__":42 convo = Conversation("You are a concise technical assistant. Answer in under 120 words.")43 while True:44 try:45 q = input("\n> ").strip()46 except (EOFError, KeyboardInterrupt):47 break48 if q in {"exit", "quit"}:49 break50 if q == "reset":51 convo.turns.clear(); print("(history cleared)"); continue52 convo.ask(q)Two details are deliberate. The timeout is 600 seconds because local generation legitimately takes minutes and a default 30-second client timeout produces truncated answers that look like model failures. And history is trimmed from the front while the system message is held separately — the failure mode otherwise is that a long conversation silently pushes the system prompt out of the context window and the model's behaviour changes for no visible reason.
Test three things: a single question, a follow-up that requires memory of the first ("what did I just ask?"), and twelve turns of chatter to confirm trimming works without the model forgetting its instructions.
Phase five: measure your machine
This is the phase that produces something worth keeping. The key idea is that a request has two phases with different bottlenecks: prefill, reading the prompt, which is compute-bound; and decode, writing tokens, which is memory-bandwidth-bound. A single average tokens-per-second number averages the two and describes neither.
1import json, time, statistics, requests23BASE = "http://localhost:11434"45def one_run(prompt, model, n_out, num_ctx):6 t0 = time.perf_counter(); first = None; n = 07 with requests.post(f"{BASE}/api/generate", stream=True, timeout=900,8 json={"model": model, "prompt": prompt, "stream": True,9 "keep_alive": -1,10 "options": {"num_predict": n_out, "temperature": 0,11 "seed": 42, "num_ctx": num_ctx}}) as r:12 for line in r.iter_lines():13 if not line:14 continue15 if first is None:16 first = time.perf_counter()17 if json.loads(line).get("done"):18 break19 n += 120 end = time.perf_counter()21 return {"ttft": first - t0,22 "tok_s": (n - 1) / (end - first) if n > 1 else 0.0,23 "e2e": end - t0, "out": n}242526def measure(label, prompt, model, n_out=200, num_ctx=4096, runs=5, warmup=2):27 for _ in range(warmup):28 one_run(prompt, model, 32, num_ctx) # discard: load + cache warm-up29 rs = [one_run(prompt, model, n_out, num_ctx) for _ in range(runs)]30 out = {"label": label, "model": model,31 "prompt_words": len(prompt.split()),32 "ttft_median": round(statistics.median(r["ttft"] for r in rs), 3),33 "ttft_max": round(max(r["ttft"] for r in rs), 3),34 "tok_s_median": round(statistics.median(r["tok_s"] for r in rs), 1),35 "e2e_median": round(statistics.median(r["e2e"] for r in rs), 2)}36 print(out)37 return out383940if __name__ == "__main__":41 filler = "The system processes requests in order. " * 142 prompts = {43 "short": "Define memory bandwidth in one sentence.",44 "medium": "Summarise the following.\n" + filler * 120,45 "long": "Summarise the following.\n" + filler * 500,46 }47 results = []48 for model in ["llama3.1:8b", "llama3.1:8b-instruct-q8_0"]:49 for label, p in prompts.items():50 try:51 results.append(measure(f"{label}", p, model))52 except requests.HTTPError as e:53 print("skipped", model, e)54 json.dump(results, open("results.json", "w"), indent=2)The warm-up runs are not optional. The first request pays model loading and kernel setup, and including it inflates TTFT by seconds and turns every comparison into noise. Temperature zero and a fixed seed keep output lengths stable so you are measuring the machine, not sampling luck.
Now check your decode rate against the physics. Generating one token requires reading every weight once, so:
ceiling in tokens/second = memory bandwidth ÷ model file size
On a card with 504 GB/s running a 4.92 GB file, that is 102 tokens per second. On dual-channel DDR5-5600 at 89.6 GB/s, it is 18. Real systems reach 60–80% of the ceiling. If your measurement is far below that, something is wrong — and it is almost always partial offload, single-channel memory, or thermal throttling.
Divide your memory bandwidth by your model file size before you touch a single flag. That one number tells you whether the machine is healthy or broken.
| Reference point | Bandwidth | Ceiling | Typical measured |
|---|---|---|---|
| DDR5-5600 dual channel, CPU only | 89.6 GB/s | 18 tok/s | 10–12 |
| Apple M3 Max | 400 GB/s | 81 tok/s | 55–65 |
| RTX 3060 12 GB | 360 GB/s | 73 tok/s | 50–60 |
| RTX 4070 | 504 GB/s | 102 tok/s | 75–95 |
| RTX 4090 | 1008 GB/s | 205 tok/s | 130–160 |
Two comparisons complete this phase. First, Q4_K_M against Q8_0: the 8-bit file is 1.7 times larger, so it should decode roughly 1.7 times slower, and it does. Confirming that on your own machine is the moment the bandwidth model stops being a claim in a table and becomes something you have observed.
Second, run twenty prompts of your own through both and check a quality property you care about — valid JSON, correct arithmetic, adherence to a word limit. Perplexity tables will not tell you whether the model you are about to deploy can still count.
Phase six: a browser interface
Optional, and worth doing if you want to hand this to someone who does not use a terminal.
1from flask import Flask, request, Response, render_template_string2from client import stream_chat # reuse phase four34app = Flask(__name__)56PAGE = """7<h2>Local assistant</h2>8<textarea id="q" rows="4" style="width:100%"></textarea>9<button onclick="go()">Ask</button>10<pre id="out" style="white-space:pre-wrap"></pre>11<script>12async function go() {13 const out = document.getElementById('out'); out.textContent = '';14 const res = await fetch('/ask', {method:'POST',15 headers:{'Content-Type':'application/json'},16 body: JSON.stringify({q: document.getElementById('q').value})});17 const reader = res.body.getReader(), dec = new TextDecoder();18 while (true) {19 const {done, value} = await reader.read();20 if (done) break;21 out.textContent += dec.decode(value, {stream:true});22 }23}24</script>25"""2627@app.route("/")28def index():29 return render_template_string(PAGE)3031@app.route("/ask", methods=["POST"])32def ask():33 q = request.json["q"]34 msgs = [{"role": "system", "content": "Be concise."},35 {"role": "user", "content": q}]36 return Response(stream_chat(msgs), mimetype="text/plain")3738if __name__ == "__main__":39 app.run(port=5000, threaded=True)Append to a text node rather than rewriting innerHTML on every token — at 95 tokens per second, reassigning the whole document body ninety-five times a second visibly stutters. And note threaded=True: without it, Flask's development server blocks completely for the duration of one generation.
When it does not work
| Symptom | Cause | Fix |
|---|---|---|
| 4–8 tok/s on a machine with a capable GPU | Model running on the CPU | ollama ps and check the PROCESSOR column; reduce context until it reaches 100% GPU |
| Out of memory on start | Context too large; KV cache allocated up front | Halve num_ctx, or set OLLAMA_KV_CACHE_TYPE=q8_0 to halve the cache |
| First request each hour takes 10–15 s | Model unloaded after idle timeout | OLLAMA_KEEP_ALIVE=-1, or send "keep_alive": -1 per request |
| Long prompts hang for 30+ seconds before output | Prefill on CPU, which is 50–100× slower than GPU | Full offload, shorter prompts, or accept it |
| Client raises a timeout mid-answer | Default HTTP timeout shorter than generation | Set timeout=600 and stream |
| Model forgets its instructions after a while | System prompt pushed out of the context window | Hold the system message separately and trim only the turns |
| Speed degrades within one long conversation | Attention over a growing KV cache | Expected; trim history or accept the taper |
| Answers are subtly worse than expected | Quantisation too aggressive for the task | Test Q4_K_M against Q6_K on your own cases before blaming the model |
What to write down, and what to do with it
The report is short and it is the actual deliverable. Record your hardware including memory bandwidth, the exact model tag and context length, the three TTFT and decode figures from phase five, the measured ratio against the bandwidth ceiling, and the quantisation comparison. Then one paragraph of recommendation with a reason attached: what this machine is suitable for and what it is not.
That last paragraph is where the exercise pays off. A setup delivering 0.3-second TTFT and 62 tokens per second at 4K context is genuinely good for interactive chat and code assistance. The same setup facing a 12,000-token document is a different machine — the KV cache alone would want 1.5 GiB, and prefill on that prompt takes four seconds on a GPU and six minutes on a CPU. Knowing which of those you have, with numbers, is the difference between deploying something and hoping.
The natural extension is to keep the harness and re-run it. Backends and runtimes change monthly, models are re-quantised, and a configuration that lost by fifteen percent last quarter may win now. A dated results.json and a script that regenerates it turns every future upgrade decision into ten minutes of measurement instead of an afternoon of reading forum threads written by people with different hardware and a different bottleneck.