Local LLM Deployment and Quantization

Mini Project: Deploy an 8B Model Locally with Ollama


Four hours from now you should have an 8-billion-parameter model answering questions on your own machine, a Python client that streams tokens and manages conversation history, a benchmark file containing your machine's real time-to-first-token and decode rate at three prompt lengths, and enough evidence to tell someone whether this hardware is good enough for a real workload.

The single most common way this build fails is skipping the arithmetic at the start. Someone downloads the largest quantisation their disk will hold, sets a 32,768-token context because larger sounds better, and gets CUDA out of memory — or worse, no error at all, just an inexplicable four tokens per second because the runtime quietly fell back to the CPU. That failure is entirely predictable from three numbers you can compute in two minutes.

So the build starts with sizing, not downloading.

Sizing 16 GB before you download anythingKept by the OS— about 4 GiBQ4_K_Mweights — 4.58 GiBKV cache, 8Kctx — 1.0 GiBComputebuffers — 0.5 GiBtopbottom6.1 GiB of the roughly 11 GiB usable; at 32K context the KV cache alone would be 4 GiB.
The weights are only part of the budget — the KV cache grows with context and is what actually pushes you over.

What you are building

ArtefactDone when
A running modelollama run gives a sensible answer in under 10 seconds cold, under 3 warm
client.pyStreams tokens, keeps multi-turn context, trims history without dropping the system prompt
benchmark.py and results.jsonTTFT and decode rate at three prompt lengths, median of at least five timed runs after warm-up
A quantisation comparisonTwo quantisations measured on both speed and one quality check of your choosing
A short written reportYour hardware, your numbers, and a recommendation with a reason

The success criterion that matters is the last one. Anyone can make a model produce text. The point of the exercise is being able to say "on this machine, an 8B at Q4_K_M gives 0.3-second TTFT and 62 tokens per second at 4K context, which is fine for chat and unusable for 20-page documents" — and to have the measurements behind it.

Phase one: size the machine before downloading anything

Three things consume memory during inference, and you can compute all three in advance.

Weights. An 8B model at Q4_K_M is a 4.92 GB file, which is 4.58 GiB in memory. Note that "4-bit" does not mean half a byte per parameter — the real figure is about 4.9 bits, because every block of weights carries a scale factor and a few sensitive tensors, such as the output layer, are kept at higher precision.

KV cache. This is the term people forget, and it is the one that kills builds. For Llama 3 8B — 32 layers, 8 key/value heads, head dimension 128, 16-bit cache — one token of context costs:

2 (K and V) × 32 × 8 × 128 × 2 bytes = 131,072 bytes = 128 KiB per token

So 4,096 tokens is 0.5 GiB, 8,192 tokens is exactly 1 GiB, and 32,768 tokens is 4 GiB — nearly as much as the model itself. It is allocated when the server starts, not when you send a long prompt.

Compute buffers. Budget 0.4–0.6 GiB.

Now find your usable memory and check the sum:

Bash
nvidia-smi --query-gpu=name,memory.total --format=csv   # NVIDIA VRAMrocm-smi --showmeminfo vram                             # AMD VRAMsysctl hw.memsize                                       # macOS unified memoryfree -g                                                 # Linux system RAM

Subtract roughly 10% for driver and display overhead on a GPU, or 4–5 GiB for the operating system on a CPU-only machine.

Your hardwareUsablePull thisContextThe sum
8 GB VRAM~7.2 GiBllama3.1:8b (Q4_K_M)40964.58 + 0.5 + 0.4 = 5.5 ✓
12 GB VRAM~11 GiBllama3.1:8b-instruct-q6_K81926.15 + 1.0 + 0.5 = 7.7 ✓
24 GB VRAM~23 GiBllama3.1:8b-instruct-q8_0327687.95 + 4.0 + 0.6 = 12.6 ✓
6 GB VRAM~5.3 GiBllama3.1:8b-instruct-q3_K_M40963.74 + 0.5 + 0.4 = 4.6 ✓
16 GB RAM, no GPU~11 GiBllama3.1:8b (Q4_K_M)40965.5 ✓, but expect 10 tok/s
8 GB RAM, no GPU~4 GiBllama3.2:3b20481.9 + 0.2 + 0.3 = 2.4 ✓
Apple Silicon 16 GB~11 GiBllama3.1:8b81924.58 + 1.0 + 0.5 = 6.1 ✓

Write your own sum down before continuing. If it does not fit, the fix in order of preference is: reduce the context, quantise the KV cache to 8-bit, then drop a quantisation level. Dropping context is nearly free; dropping weight precision costs real quality.

Weights plus KV cache plus buffers must fit in memory you actually have. Every "why is my local model broken" question I have ever seen starts with someone not doing this sum.

Phase two: install and verify

Bash
curl -fsSL https://ollama.com/install.sh | sh      # Linux# macOS / Windows: download the installer from ollama.comollama --versioncurl -s http://localhost:11434/api/version         # is the daemon up?sudo systemctl status ollama                       # Linux service check

If the daemon is not running, start it in a terminal you can watch with ollama serve. The log output during the first model load is genuinely informative — it tells you which backend was selected and how many layers were offloaded, and if it says offloaded 0/33 layers to GPU, you have found your problem before it becomes mysterious.

Set the environment before starting the daemon, not after:

Bash
export OLLAMA_KEEP_ALIVE=-1        # never unload; removes the cold-start stallexport OLLAMA_MAX_LOADED_MODELS=1  # stop a second model evicting the firstexport OLLAMA_FLASH_ATTENTION=1    # force flash attention on (recent versions enable it automatically where supported)ollama serve

Phase three: pull the model and read its tag

Bash
ollama pull llama3.1:8bollama listollama ps       # confirms it is resident, and its actual memory footprint

Read the tag properly, because it encodes the arithmetic from phase one. In llama3.1:8b-instruct-q4_K_M: the family, the parameter count, the instruction-tuned variant rather than the raw base model, and the quantisation. A bare llama3.1:8b resolves to exactly that, which is why the download is 4.9 GB rather than 16.

The project uses llama3.1:8b because every number in this lesson is computed for it. Once the pipeline works, swap in a current model from the Ollama library and redo the phase-one sum with its layer count, key/value heads and file size.

Pull a second quantisation now, while you are here — you will compare them later:

Bash
ollama pull llama3.1:8b-instruct-q8_0     # only if your sum from phase one allows itollama run llama3.1:8b "Give three reasons a local model might run slowly."

Then verify the offload immediately, because this is where silent failure lives:

Bash
ollama ps# Look at the PROCESSOR column. "100% GPU" is what you want.# "48%/52% CPU/GPU" means partial offload and roughly a third of full speed.nvidia-smi --query-gpu=memory.used --format=csv

Partial offload is far worse than it sounds. With 28 of 32 layers on the GPU, every token still waits for four CPU layers, and throughput lands nearer a third of full-offload speed than seven-eighths. If you see it, reduce context or drop one quantisation level until you reach 100% GPU.

A model that half-fits on the GPU is slower than a smaller one that fits entirely. The last four layers are worth more than the first twenty-eight.

Phase four: a client worth keeping

Python
import json, requestsBASE = "http://localhost:11434"def stream_chat(messages, model="llama3.1:8b", temperature=0.3, num_ctx=4096):    """Yield content pieces as they are generated."""    with requests.post(f"{BASE}/api/chat", stream=True, timeout=600,            json={"model": model, "messages": messages, "stream": True,                  "keep_alive": -1,                  "options": {"temperature": temperature, "num_ctx": num_ctx}}) as r:        r.raise_for_status()        for line in r.iter_lines():            if not line:                continue            obj = json.loads(line)            if obj.get("done"):                return            yield obj["message"]["content"]class Conversation:    """The server is stateless. History management is your job."""    def __init__(self, system, model="llama3.1:8b", max_pairs=10):        self.system = {"role": "system", "content": system}        self.model, self.max_pairs = model, max_pairs        self.turns = []    def ask(self, text):        self.turns.append({"role": "user", "content": text})        parts = []        for piece in stream_chat([self.system] + self.turns, model=self.model):            parts.append(piece)            print(piece, end="", flush=True)        print()        self.turns.append({"role": "assistant", "content": "".join(parts)})        if len(self.turns) > self.max_pairs * 2:      # never drop the system message            self.turns = self.turns[-self.max_pairs * 2:]        return self.turns[-1]["content"]if __name__ == "__main__":    convo = Conversation("You are a concise technical assistant. Answer in under 120 words.")    while True:        try:            q = input("\n> ").strip()        except (EOFError, KeyboardInterrupt):            break        if q in {"exit", "quit"}:            break        if q == "reset":            convo.turns.clear(); print("(history cleared)"); continue        convo.ask(q)

Two details are deliberate. The timeout is 600 seconds because local generation legitimately takes minutes and a default 30-second client timeout produces truncated answers that look like model failures. And history is trimmed from the front while the system message is held separately — the failure mode otherwise is that a long conversation silently pushes the system prompt out of the context window and the model's behaviour changes for no visible reason.

Test three things: a single question, a follow-up that requires memory of the first ("what did I just ask?"), and twelve turns of chatter to confirm trimming works without the model forgetting its instructions.

Phase five: measure your machine

This is the phase that produces something worth keeping. The key idea is that a request has two phases with different bottlenecks: prefill, reading the prompt, which is compute-bound; and decode, writing tokens, which is memory-bandwidth-bound. A single average tokens-per-second number averages the two and describes neither.

Python
import json, time, statistics, requestsBASE = "http://localhost:11434"def one_run(prompt, model, n_out, num_ctx):    t0 = time.perf_counter(); first = None; n = 0    with requests.post(f"{BASE}/api/generate", stream=True, timeout=900,            json={"model": model, "prompt": prompt, "stream": True,                  "keep_alive": -1,                  "options": {"num_predict": n_out, "temperature": 0,                              "seed": 42, "num_ctx": num_ctx}}) as r:        for line in r.iter_lines():            if not line:                continue            if first is None:                first = time.perf_counter()            if json.loads(line).get("done"):                break            n += 1    end = time.perf_counter()    return {"ttft": first - t0,            "tok_s": (n - 1) / (end - first) if n > 1 else 0.0,            "e2e": end - t0, "out": n}def measure(label, prompt, model, n_out=200, num_ctx=4096, runs=5, warmup=2):    for _ in range(warmup):        one_run(prompt, model, 32, num_ctx)        # discard: load + cache warm-up    rs = [one_run(prompt, model, n_out, num_ctx) for _ in range(runs)]    out = {"label": label, "model": model,           "prompt_words": len(prompt.split()),           "ttft_median": round(statistics.median(r["ttft"] for r in rs), 3),           "ttft_max":    round(max(r["ttft"] for r in rs), 3),           "tok_s_median": round(statistics.median(r["tok_s"] for r in rs), 1),           "e2e_median":  round(statistics.median(r["e2e"] for r in rs), 2)}    print(out)    return outif __name__ == "__main__":    filler = "The system processes requests in order. " * 1    prompts = {        "short":  "Define memory bandwidth in one sentence.",        "medium": "Summarise the following.\n" + filler * 120,        "long":   "Summarise the following.\n" + filler * 500,    }    results = []    for model in ["llama3.1:8b", "llama3.1:8b-instruct-q8_0"]:        for label, p in prompts.items():            try:                results.append(measure(f"{label}", p, model))            except requests.HTTPError as e:                print("skipped", model, e)    json.dump(results, open("results.json", "w"), indent=2)

The warm-up runs are not optional. The first request pays model loading and kernel setup, and including it inflates TTFT by seconds and turns every comparison into noise. Temperature zero and a fixed seed keep output lengths stable so you are measuring the machine, not sampling luck.

Now check your decode rate against the physics. Generating one token requires reading every weight once, so:

ceiling in tokens/second = memory bandwidth ÷ model file size

On a card with 504 GB/s running a 4.92 GB file, that is 102 tokens per second. On dual-channel DDR5-5600 at 89.6 GB/s, it is 18. Real systems reach 60–80% of the ceiling. If your measurement is far below that, something is wrong — and it is almost always partial offload, single-channel memory, or thermal throttling.

Divide your memory bandwidth by your model file size before you touch a single flag. That one number tells you whether the machine is healthy or broken.

Reference pointBandwidthCeilingTypical measured
DDR5-5600 dual channel, CPU only89.6 GB/s18 tok/s10–12
Apple M3 Max400 GB/s81 tok/s55–65
RTX 3060 12 GB360 GB/s73 tok/s50–60
RTX 4070504 GB/s102 tok/s75–95
RTX 40901008 GB/s205 tok/s130–160

Two comparisons complete this phase. First, Q4_K_M against Q8_0: the 8-bit file is 1.7 times larger, so it should decode roughly 1.7 times slower, and it does. Confirming that on your own machine is the moment the bandwidth model stops being a claim in a table and becomes something you have observed.

Second, run twenty prompts of your own through both and check a quality property you care about — valid JSON, correct arithmetic, adherence to a word limit. Perplexity tables will not tell you whether the model you are about to deploy can still count.

Phase six: a browser interface

Optional, and worth doing if you want to hand this to someone who does not use a terminal.

Python
from flask import Flask, request, Response, render_template_stringfrom client import stream_chat        # reuse phase fourapp = Flask(__name__)PAGE = """<h2>Local assistant</h2><textarea id="q" rows="4" style="width:100%"></textarea><button onclick="go()">Ask</button><pre id="out" style="white-space:pre-wrap"></pre><script>async function go() {  const out = document.getElementById('out'); out.textContent = '';  const res = await fetch('/ask', {method:'POST',      headers:{'Content-Type':'application/json'},      body: JSON.stringify({q: document.getElementById('q').value})});  const reader = res.body.getReader(), dec = new TextDecoder();  while (true) {    const {done, value} = await reader.read();    if (done) break;    out.textContent += dec.decode(value, {stream:true});  }}</script>"""@app.route("/")def index():    return render_template_string(PAGE)@app.route("/ask", methods=["POST"])def ask():    q = request.json["q"]    msgs = [{"role": "system", "content": "Be concise."},            {"role": "user", "content": q}]    return Response(stream_chat(msgs), mimetype="text/plain")if __name__ == "__main__":    app.run(port=5000, threaded=True)

Append to a text node rather than rewriting innerHTML on every token — at 95 tokens per second, reassigning the whole document body ninety-five times a second visibly stutters. And note threaded=True: without it, Flask's development server blocks completely for the duration of one generation.

When it does not work

SymptomCauseFix
4–8 tok/s on a machine with a capable GPUModel running on the CPUollama ps and check the PROCESSOR column; reduce context until it reaches 100% GPU
Out of memory on startContext too large; KV cache allocated up frontHalve num_ctx, or set OLLAMA_KV_CACHE_TYPE=q8_0 to halve the cache
First request each hour takes 10–15 sModel unloaded after idle timeoutOLLAMA_KEEP_ALIVE=-1, or send "keep_alive": -1 per request
Long prompts hang for 30+ seconds before outputPrefill on CPU, which is 50–100× slower than GPUFull offload, shorter prompts, or accept it
Client raises a timeout mid-answerDefault HTTP timeout shorter than generationSet timeout=600 and stream
Model forgets its instructions after a whileSystem prompt pushed out of the context windowHold the system message separately and trim only the turns
Speed degrades within one long conversationAttention over a growing KV cacheExpected; trim history or accept the taper
Answers are subtly worse than expectedQuantisation too aggressive for the taskTest Q4_K_M against Q6_K on your own cases before blaming the model

What to write down, and what to do with it

The report is short and it is the actual deliverable. Record your hardware including memory bandwidth, the exact model tag and context length, the three TTFT and decode figures from phase five, the measured ratio against the bandwidth ceiling, and the quantisation comparison. Then one paragraph of recommendation with a reason attached: what this machine is suitable for and what it is not.

That last paragraph is where the exercise pays off. A setup delivering 0.3-second TTFT and 62 tokens per second at 4K context is genuinely good for interactive chat and code assistance. The same setup facing a 12,000-token document is a different machine — the KV cache alone would want 1.5 GiB, and prefill on that prompt takes four seconds on a GPU and six minutes on a CPU. Knowing which of those you have, with numbers, is the difference between deploying something and hoping.

The natural extension is to keep the harness and re-run it. Backends and runtimes change monthly, models are re-quantised, and a configuration that lost by fifteen percent last quarter may win now. A dated results.json and a script that regenerates it turns every future upgrade decision into ten minutes of measurement instead of an afternoon of reading forum threads written by people with different hardware and a different bottleneck.