Course Content
Local LLM Deployment and Quantization
3 sections · 7 lessons
Evaluating Latency and Resource Trade-offs
A team reports their local deployment at 45 tokens per second and calls it fast. Users call it slow. Both are telling the truth.
The benchmark used a nine-token prompt. Real users paste in a 1,200-token document. Prompt processing runs at 38 tokens per second on that machine, so the model spends 31 seconds reading before it emits a single character. Then it generates at 45 tokens per second, which is genuinely quick. The benchmark's single number, measured on a prompt with nothing to read, is just that decode rate, and it hides the entire problem.
This is the characteristic measurement failure in local deployment. A single tokens-per-second number averages two operations with completely different bottlenecks and completely different user consequences. Separating them is the first thing to fix, and everything else about evaluating a local setup follows from that separation.
The metrics that describe a request
Every generation has exactly two phases, and you need a number for each.
| Metric | What it measures | Bound by | User experiences it as |
|---|---|---|---|
| TTFT — time to first token | Prefill: reading the whole prompt | Compute (FLOP/s) | The dead pause after pressing enter |
| TPOT — time per output token | Decode: producing each subsequent token | Memory bandwidth | How fast words appear |
| End-to-end latency | TTFT + (n−1) × TPOT | Both | Total wait for a complete answer |
| Per-user throughput | 1 / TPOT | Bandwidth | Reading speed |
| Aggregate throughput | Tokens/s across all concurrent requests | Batching efficiency | Nothing, directly — it is a capacity metric |
Work the arithmetic for one realistic request — an 800-token prompt and a 300-token answer — on two machines.
| GPU: 12 GB card, Q4_K_M | CPU: 16-core, DDR5-5600 | |
|---|---|---|
| Prefill rate | 3,000 tok/s | 30 tok/s |
| TTFT = 800 / rate | 0.27 s | 26.7 s |
| Decode rate | 95 tok/s | 11 tok/s |
| TPOT | 10.5 ms | 91 ms |
| Decode time = 299 × TPOT | 3.14 s | 27.2 s |
| End-to-end | 3.41 s | 53.9 s |
The GPU is 8.6 times faster at decoding and 99 times faster at prefill. Anyone reporting a single "tokens per second" figure for these two systems will report a gap of roughly ten, and will have hidden the factor of a hundred that users actually feel.
One anchor makes TPOT interpretable. An adult reads at about 240 words per minute, which at roughly 1.3 tokens per word is 5.2 tokens per second. Any streaming rate above about 7 tokens per second outpaces reading, and further gains stop being perceptible in an interactive chat. This is why an 11 tok/s CPU setup feels acceptable once text starts flowing — and why the 26.7-second silence before it does is the only number that matters there.
Report TTFT and TPOT separately or you are not measuring. An average tokens-per-second figure is the arithmetic mean of a compute-bound phase and a bandwidth-bound phase, and it describes neither.
Resource metrics
Speed is one axis. What the system consumes is the other, and three figures matter.
Peak memory, not steady-state. The number that determines whether you crash is the maximum during a long-context request, which is weights plus a full KV cache plus compute buffers. Measure it by running your longest realistic prompt, not your typical one.
Energy per token, which produces a genuinely counterintuitive result. A GPU drawing 200 W at 95 tok/s spends 200 / 95 = 2.1 joules per token. A CPU drawing 90 W at 11 tok/s spends 90 / 11 = 8.2 joules per token. The GPU pulls more than twice the power and is four times more efficient per unit of work, because it finishes so much sooner. "Use the CPU to save energy" is wrong in both directions.
Cost per million tokens, computed honestly. A 600-dollar GPU amortised over three years is USD 16.70 per month. Running it four hours a day at 200 W is 0.8 kWh per day, 24 kWh per month, about USD 7.20 at 30 cents per kWh. Total: roughly 24 dollars per month. At 95 tok/s for four hours a day that is 41 million tokens per month, or 0.58 dollars per million tokens — competitive with hosted small-model pricing, but only at that utilisation. Run it one hour a day and the same hardware costs about USD 1.80 per million, because the fixed cost dominates.
The conclusion is worth stating plainly: local inference is not reliably cheaper than an API. It is chosen for data residency, offline operation, freedom from rate limits, and a latency floor you control. If the pitch is cost savings, do this calculation first — at low volume it usually loses.
Quality metrics, and why perplexity is not enough
Perplexity measures how surprised a model is by held-out text: the exponential of the average negative log-likelihood per token. Lower is better, and it is the standard way to compare quantisations because it is cheap, deterministic, and sensitive.
1./build/bin/llama-perplexity -m model-Q4_K_M.gguf -f wikitext-2-raw/wiki.test.raw -c 409623# Better: compare the quantised model's output distribution against the F16 original4# Step 1 records the F16 logits to kl.dat; step 2 compares the quantised model against them5./build/bin/llama-perplexity -m model-F16.gguf -f calib.txt --kl-divergence-base kl.dat6./build/bin/llama-perplexity -m model-Q4_K_M.gguf -f calib.txt --kl-divergence-base kl.dat --kl-divergenceKL divergence is the sharper instrument. Perplexity can stay almost unchanged while the model's full distribution over next tokens shifts — the top choice survives, but the ordering below it scrambles, which shows up as degraded behaviour under sampling and in constrained generation. KL divergence against the F16 reference catches exactly that.
Both are proxies. Neither tells you whether the model still emits valid JSON. Task-level evaluation on your own data is the only measurement that answers the question you actually have:
1import json, re23def valid_json_rate(client, model, cases):4 ok = 05 for c in cases:6 out = client.chat.completions.create(7 model=model, messages=c["messages"], temperature=0, max_tokens=4008 ).choices[0].message.content9 block = re.search(r"\{.*\}", out, re.S)10 try:11 parsed = json.loads(block.group(0)) if block else None12 ok += int(parsed is not None and set(c["required"]) <= set(parsed))13 except json.JSONDecodeError:14 pass15 return ok / len(cases)1617# Run the same 200 cases against each quantisation. This takes twenty minutes18# and answers a question no perplexity table can.In production you would switch on the server's JSON-schema constrained output (Ollama's format field, or response_format on the OpenAI-compatible endpoints), which guarantees parseable JSON. This test deliberately leaves it off: it measures how well the quantised model follows the format on its own, which is what separates one quantisation from another — and constrained decoding only fixes the syntax, not whether the values inside are right.
Building a harness you can trust
Four details separate a useful benchmark from a misleading one, and all four are commonly skipped.
1import time, statistics, requests, json23def measure(prompt, n_out=200, model="llama3.1:8b", runs=10, warmup=2):4 lat = []5 for i in range(runs + warmup):6 t0 = time.perf_counter()7 first = None8 count = 09 with requests.post("http://localhost:11434/api/generate",10 json={"model": model, "prompt": prompt, "stream": True,11 "options": {"num_predict": n_out, "temperature": 0,12 "seed": 42, "num_ctx": 4096}},13 stream=True, timeout=600) as r:14 for line in r.iter_lines():15 if not line:16 continue17 if first is None:18 first = time.perf_counter()19 if json.loads(line).get("done"):20 break21 count += 122 end = time.perf_counter()23 if i >= warmup: # discard cold-start runs24 lat.append({"ttft": first - t0,25 "tpot": (end - first) / max(count - 1, 1),26 "e2e": end - t0})2728 def pct(key, p):29 xs = sorted(x[key] for x in lat)30 return xs[min(int(len(xs) * p), len(xs) - 1)]3132 return {33 "ttft_p50": pct("ttft", 0.50), "ttft_p95": pct("ttft", 0.95),34 "tpot_p50": pct("tpot", 0.50),35 "tok_s": 1 / statistics.median(x["tpot"] for x in lat),36 "e2e_p95": pct("e2e", 0.95),37 }3839for name, prompt in [("short", "Define latency."),40 ("medium", "Summarise this: " + "word " * 500),41 ("long", "Summarise this: " + "word " * 2000)]:42 print(name, measure(prompt))Warm up. The first request pays model load, kernel compilation, and cache population. Including it inflates TTFT by seconds and makes every comparison noise.
Fix the seed and set temperature to zero. Otherwise output lengths vary between runs and you are measuring sampling luck.
Sweep prompt length. One prompt length gives one point on a curve whose slope is the entire story. Three lengths reveal whether you are prefill-limited.
Report percentiles, not means. A p50 of 3 s with a p95 of 14 s describes a system where one user in twenty has a bad time. The mean of those, around 4 s, describes nothing real.
Comparing quantisations properly
Because decode speed is bandwidth divided by model size, a smaller quantisation is faster as well as smaller — the two axes move together, which makes the comparison unusually clean.
| Format | Size | Weights + 8K cache | Decode (12 GB card) | Perplexity increase | Verdict |
|---|---|---|---|---|---|
| F16 | 16.07 GB | 16.0 GiB | — (will not fit) | baseline | Reference only |
| Q8_0 | 8.54 GB | 9.0 GiB | ~48 tok/s | < 0.1% | Fits, but Q6_K is faster for no visible cost |
| Q6_K | 6.60 GB | 7.2 GiB | ~62 tok/s | ~0.3% | Best quality that fits comfortably |
| Q5_K_M | 5.73 GB | 6.3 GiB | ~72 tok/s | ~0.9% | Excellent balance |
| Q4_K_M | 4.92 GB | 5.6 GiB | ~95 tok/s | ~2.8% | The default for good reason |
| Q3_K_M | 4.02 GB | 4.7 GiB | ~112 tok/s | ~10% | Only when Q4 genuinely will not fit |
| Q2_K | 3.18 GB | 4.0 GiB | ~135 tok/s | ~56% | Formatting and instruction-following fail |
Read the quality column as a curve with a cliff. From 8 bits to 5, nothing meaningful happens. From 5 to 4, a little. Below 4, quality falls quickly and it falls first in the behaviours that are hardest to notice in casual testing: consistency over long outputs, arithmetic, and adherence to a schema.
There is one more comparison that matters more than any row above. Given a fixed memory budget of about 5 GiB, you can run an 8B at Q4_K_M or a 3B at Q8_0. The 8B at Q4 wins comfortably on essentially every task. Spend memory on parameters before you spend it on precision — down to about 4 bits, below which the rule reverses.
Turning measurements into a decision
The mistake is optimising a metric instead of satisfying a constraint. Start from the constraint that cannot move.
| Use case | Binding constraint | Target | Configuration |
|---|---|---|---|
| Interactive chat / IDE completion | TTFT | < 500 ms; TPOT < 100 ms | Full GPU offload, short context, prefix caching, low parallelism |
| Batch document processing | Aggregate throughput | Maximise tokens/s total | High parallelism, continuous batching, accept 3× worse per-request latency |
| Edge device | Peak memory | Fit in a fixed budget | Smallest model that passes your task eval; Q4_K_M; 8-bit KV cache |
| Regulated / offline | Data must not leave | Any workable speed | Largest model the hardware holds; users will wait if they must |
| Shared team server | p95 latency under load | p95 < 5× p50 | Cap parallelism, queue with a depth limit, reject rather than degrade |
For interactive use the single highest-leverage change is usually not the model at all. If your prompt carries a 900-token system preamble, prefix caching removes that from TTFT on every subsequent call — provided the preamble is byte-identical. Interpolating a timestamp or a session ID near the top of the prompt invalidates the cache and costs you the full prefill each time, invisibly.
For batch work, exploit the fact that decoding reads every weight once per step regardless of how many sequences are in the batch. Four concurrent sequences roughly double aggregate throughput while halving per-request speed. Eight might reach 2.5 times aggregate. Push parallelism until KV cache memory runs out or aggregate throughput stops climbing, whichever comes first.
Optimise the constraint that binds. Doubling decode speed on a system where 88% of the wait is prefill improves the user's experience by six percent.
What to watch once it is running
Log four numbers per request: TTFT, TPOT, prompt tokens, and completion tokens. Everything useful derives from those.
1import time, logging2log = logging.getLogger("llm")34def instrumented(stream_fn, *args, **kw):5 t0 = time.perf_counter(); first = None; n = 06 for tok in stream_fn(*args, **kw):7 if first is None:8 first = time.perf_counter()9 n += 110 yield tok11 end = time.perf_counter()12 log.info("ttft=%.3f tpot=%.4f out=%d e2e=%.3f",13 first - t0, (end - first) / max(n - 1, 1), n, end - t0)Then watch for four specific signatures, because each has one likely cause:
| Symptom | Most likely cause | Check |
|---|---|---|
| TTFT occasionally 10–30× normal | Model was unloaded and reloaded from disk | load_duration in the response; raise keep_alive |
| TTFT crept up over weeks | Prompt template grew; prefix cache misses | Log prompt token count over time |
| TPOT doubled at a fixed time of day | Concurrency, or a second model evicting the first | GET /api/ps; cap MAX_LOADED_MODELS |
| TPOT degrades within a single long conversation | Attention over a growing KV cache | Plot TPOT against context length; trim history |
That last one surprises people. Decode cost is dominated by reading the weights, which is constant — but attention must also read the entire KV cache each step, and that grows linearly. At 2K of context the cache is 0.25 GiB against 4.58 GiB of weights and contributes about 5% of the memory traffic. At 32K it is 4 GiB against 4.58 and contributes nearly half. A conversation that starts at 95 tok/s and drifts to 60 by turn forty is behaving exactly as designed.
Deciding, and knowing when to stop
The practical sequence is short and rarely followed. Write down the constraint before you measure anything: the TTFT budget, the memory ceiling, the accuracy floor on your own task. Then measure a baseline at Q4_K_M with full offload, because that configuration is right often enough to be the sensible starting point. Then change one thing.
If the baseline meets the constraint, stop. This is the hardest part. There is always another quantisation to try and another flag to sweep, and the returns after full GPU offload and a sensible context size are small enough to be indistinguishable from measurement noise on a busy machine.
If it does not meet the constraint, the fix is determined by which phase is failing. Failing TTFT means prefill, which means either the prompt is too long, the prefix cache is being invalidated, or layers are still on the CPU. Failing TPOT means bandwidth, which means a smaller model, a smaller quantisation, or better hardware — and no amount of thread tuning. Failing on memory means the KV cache before it means the weights: quantising the cache to 8 bits halves the context term for a quality cost far below what dropping from Q4_K_M to Q3_K_M would inflict.
Keep the results in a file, dated, with the exact model file, flags, and hardware. Backends move quickly and a configuration that lost by 15% six months ago may win now. Without a written baseline you will re-run the same experiments every quarter and reach conclusions you have already reached — which is, in the end, the only genuinely wasted time in this whole exercise.