Local LLM Deployment and Quantization

Evaluating Latency and Resource Trade-offs


A team reports their local deployment at 45 tokens per second and calls it fast. Users call it slow. Both are telling the truth.

The benchmark used a nine-token prompt. Real users paste in a 1,200-token document. Prompt processing runs at 38 tokens per second on that machine, so the model spends 31 seconds reading before it emits a single character. Then it generates at 45 tokens per second, which is genuinely quick. The benchmark's single number, measured on a prompt with nothing to read, is just that decode rate, and it hides the entire problem.

This is the characteristic measurement failure in local deployment. A single tokens-per-second number averages two operations with completely different bottlenecks and completely different user consequences. Separating them is the first thing to fix, and everything else about evaluating a local setup follows from that separation.

45 tokens per second, and still slowWhat the benchmark measured• Decode rate — 45 tokens per second• Averaged over a 400-token answer• From a nine-token prompt• One request, model already warmWhat the user actually waited for• Time to first token — 31 seconds• A 1,200-token document to prefill• A cold model reloaded from disk• Two other people ahead in the queue
Throughput and time-to-first-token are different numbers, and only one of them is what waiting feels like.

The metrics that describe a request

Every generation has exactly two phases, and you need a number for each.

MetricWhat it measuresBound byUser experiences it as
TTFT — time to first tokenPrefill: reading the whole promptCompute (FLOP/s)The dead pause after pressing enter
TPOT — time per output tokenDecode: producing each subsequent tokenMemory bandwidthHow fast words appear
End-to-end latencyTTFT + (n−1) × TPOTBothTotal wait for a complete answer
Per-user throughput1 / TPOTBandwidthReading speed
Aggregate throughputTokens/s across all concurrent requestsBatching efficiencyNothing, directly — it is a capacity metric

Work the arithmetic for one realistic request — an 800-token prompt and a 300-token answer — on two machines.

GPU: 12 GB card, Q4_K_MCPU: 16-core, DDR5-5600
Prefill rate3,000 tok/s30 tok/s
TTFT = 800 / rate0.27 s26.7 s
Decode rate95 tok/s11 tok/s
TPOT10.5 ms91 ms
Decode time = 299 × TPOT3.14 s27.2 s
End-to-end3.41 s53.9 s

The GPU is 8.6 times faster at decoding and 99 times faster at prefill. Anyone reporting a single "tokens per second" figure for these two systems will report a gap of roughly ten, and will have hidden the factor of a hundred that users actually feel.

One anchor makes TPOT interpretable. An adult reads at about 240 words per minute, which at roughly 1.3 tokens per word is 5.2 tokens per second. Any streaming rate above about 7 tokens per second outpaces reading, and further gains stop being perceptible in an interactive chat. This is why an 11 tok/s CPU setup feels acceptable once text starts flowing — and why the 26.7-second silence before it does is the only number that matters there.

Report TTFT and TPOT separately or you are not measuring. An average tokens-per-second figure is the arithmetic mean of a compute-bound phase and a bandwidth-bound phase, and it describes neither.

Resource metrics

Speed is one axis. What the system consumes is the other, and three figures matter.

Peak memory, not steady-state. The number that determines whether you crash is the maximum during a long-context request, which is weights plus a full KV cache plus compute buffers. Measure it by running your longest realistic prompt, not your typical one.

Energy per token, which produces a genuinely counterintuitive result. A GPU drawing 200 W at 95 tok/s spends 200 / 95 = 2.1 joules per token. A CPU drawing 90 W at 11 tok/s spends 90 / 11 = 8.2 joules per token. The GPU pulls more than twice the power and is four times more efficient per unit of work, because it finishes so much sooner. "Use the CPU to save energy" is wrong in both directions.

Cost per million tokens, computed honestly. A 600-dollar GPU amortised over three years is USD 16.70 per month. Running it four hours a day at 200 W is 0.8 kWh per day, 24 kWh per month, about USD 7.20 at 30 cents per kWh. Total: roughly 24 dollars per month. At 95 tok/s for four hours a day that is 41 million tokens per month, or 0.58 dollars per million tokens — competitive with hosted small-model pricing, but only at that utilisation. Run it one hour a day and the same hardware costs about USD 1.80 per million, because the fixed cost dominates.

The conclusion is worth stating plainly: local inference is not reliably cheaper than an API. It is chosen for data residency, offline operation, freedom from rate limits, and a latency floor you control. If the pitch is cost savings, do this calculation first — at low volume it usually loses.

Quality metrics, and why perplexity is not enough

Perplexity measures how surprised a model is by held-out text: the exponential of the average negative log-likelihood per token. Lower is better, and it is the standard way to compare quantisations because it is cheap, deterministic, and sensitive.

Bash
./build/bin/llama-perplexity -m model-Q4_K_M.gguf -f wikitext-2-raw/wiki.test.raw -c 4096# Better: compare the quantised model's output distribution against the F16 original# Step 1 records the F16 logits to kl.dat; step 2 compares the quantised model against them./build/bin/llama-perplexity -m model-F16.gguf -f calib.txt --kl-divergence-base kl.dat./build/bin/llama-perplexity -m model-Q4_K_M.gguf -f calib.txt --kl-divergence-base kl.dat --kl-divergence

KL divergence is the sharper instrument. Perplexity can stay almost unchanged while the model's full distribution over next tokens shifts — the top choice survives, but the ordering below it scrambles, which shows up as degraded behaviour under sampling and in constrained generation. KL divergence against the F16 reference catches exactly that.

Both are proxies. Neither tells you whether the model still emits valid JSON. Task-level evaluation on your own data is the only measurement that answers the question you actually have:

Python
import json, redef valid_json_rate(client, model, cases):    ok = 0    for c in cases:        out = client.chat.completions.create(            model=model, messages=c["messages"], temperature=0, max_tokens=400        ).choices[0].message.content        block = re.search(r"\{.*\}", out, re.S)        try:            parsed = json.loads(block.group(0)) if block else None            ok += int(parsed is not None and set(c["required"]) <= set(parsed))        except json.JSONDecodeError:            pass    return ok / len(cases)# Run the same 200 cases against each quantisation. This takes twenty minutes# and answers a question no perplexity table can.

In production you would switch on the server's JSON-schema constrained output (Ollama's format field, or response_format on the OpenAI-compatible endpoints), which guarantees parseable JSON. This test deliberately leaves it off: it measures how well the quantised model follows the format on its own, which is what separates one quantisation from another — and constrained decoding only fixes the syntax, not whether the values inside are right.

Building a harness you can trust

Four details separate a useful benchmark from a misleading one, and all four are commonly skipped.

Python
import time, statistics, requests, jsondef measure(prompt, n_out=200, model="llama3.1:8b", runs=10, warmup=2):    lat = []    for i in range(runs + warmup):        t0 = time.perf_counter()        first = None        count = 0        with requests.post("http://localhost:11434/api/generate",                json={"model": model, "prompt": prompt, "stream": True,                      "options": {"num_predict": n_out, "temperature": 0,                                  "seed": 42, "num_ctx": 4096}},                stream=True, timeout=600) as r:            for line in r.iter_lines():                if not line:                    continue                if first is None:                    first = time.perf_counter()                if json.loads(line).get("done"):                    break                count += 1        end = time.perf_counter()        if i >= warmup:                       # discard cold-start runs            lat.append({"ttft": first - t0,                        "tpot": (end - first) / max(count - 1, 1),                        "e2e": end - t0})    def pct(key, p):        xs = sorted(x[key] for x in lat)        return xs[min(int(len(xs) * p), len(xs) - 1)]    return {        "ttft_p50": pct("ttft", 0.50), "ttft_p95": pct("ttft", 0.95),        "tpot_p50": pct("tpot", 0.50),        "tok_s":    1 / statistics.median(x["tpot"] for x in lat),        "e2e_p95":  pct("e2e", 0.95),    }for name, prompt in [("short", "Define latency."),                     ("medium", "Summarise this: " + "word " * 500),                     ("long",   "Summarise this: " + "word " * 2000)]:    print(name, measure(prompt))

Warm up. The first request pays model load, kernel compilation, and cache population. Including it inflates TTFT by seconds and makes every comparison noise.

Fix the seed and set temperature to zero. Otherwise output lengths vary between runs and you are measuring sampling luck.

Sweep prompt length. One prompt length gives one point on a curve whose slope is the entire story. Three lengths reveal whether you are prefill-limited.

Report percentiles, not means. A p50 of 3 s with a p95 of 14 s describes a system where one user in twenty has a bad time. The mean of those, around 4 s, describes nothing real.

Comparing quantisations properly

Because decode speed is bandwidth divided by model size, a smaller quantisation is faster as well as smaller — the two axes move together, which makes the comparison unusually clean.

FormatSizeWeights + 8K cacheDecode (12 GB card)Perplexity increaseVerdict
F1616.07 GB16.0 GiB— (will not fit)baselineReference only
Q8_08.54 GB9.0 GiB~48 tok/s< 0.1%Fits, but Q6_K is faster for no visible cost
Q6_K6.60 GB7.2 GiB~62 tok/s~0.3%Best quality that fits comfortably
Q5_K_M5.73 GB6.3 GiB~72 tok/s~0.9%Excellent balance
Q4_K_M4.92 GB5.6 GiB~95 tok/s~2.8%The default for good reason
Q3_K_M4.02 GB4.7 GiB~112 tok/s~10%Only when Q4 genuinely will not fit
Q2_K3.18 GB4.0 GiB~135 tok/s~56%Formatting and instruction-following fail

Read the quality column as a curve with a cliff. From 8 bits to 5, nothing meaningful happens. From 5 to 4, a little. Below 4, quality falls quickly and it falls first in the behaviours that are hardest to notice in casual testing: consistency over long outputs, arithmetic, and adherence to a schema.

There is one more comparison that matters more than any row above. Given a fixed memory budget of about 5 GiB, you can run an 8B at Q4_K_M or a 3B at Q8_0. The 8B at Q4 wins comfortably on essentially every task. Spend memory on parameters before you spend it on precision — down to about 4 bits, below which the rule reverses.

Turning measurements into a decision

The mistake is optimising a metric instead of satisfying a constraint. Start from the constraint that cannot move.

Use caseBinding constraintTargetConfiguration
Interactive chat / IDE completionTTFT< 500 ms; TPOT < 100 msFull GPU offload, short context, prefix caching, low parallelism
Batch document processingAggregate throughputMaximise tokens/s totalHigh parallelism, continuous batching, accept 3× worse per-request latency
Edge devicePeak memoryFit in a fixed budgetSmallest model that passes your task eval; Q4_K_M; 8-bit KV cache
Regulated / offlineData must not leaveAny workable speedLargest model the hardware holds; users will wait if they must
Shared team serverp95 latency under loadp95 < 5× p50Cap parallelism, queue with a depth limit, reject rather than degrade

For interactive use the single highest-leverage change is usually not the model at all. If your prompt carries a 900-token system preamble, prefix caching removes that from TTFT on every subsequent call — provided the preamble is byte-identical. Interpolating a timestamp or a session ID near the top of the prompt invalidates the cache and costs you the full prefill each time, invisibly.

For batch work, exploit the fact that decoding reads every weight once per step regardless of how many sequences are in the batch. Four concurrent sequences roughly double aggregate throughput while halving per-request speed. Eight might reach 2.5 times aggregate. Push parallelism until KV cache memory runs out or aggregate throughput stops climbing, whichever comes first.

Optimise the constraint that binds. Doubling decode speed on a system where 88% of the wait is prefill improves the user's experience by six percent.

What to watch once it is running

Log four numbers per request: TTFT, TPOT, prompt tokens, and completion tokens. Everything useful derives from those.

Python
import time, logginglog = logging.getLogger("llm")def instrumented(stream_fn, *args, **kw):    t0 = time.perf_counter(); first = None; n = 0    for tok in stream_fn(*args, **kw):        if first is None:            first = time.perf_counter()        n += 1        yield tok    end = time.perf_counter()    log.info("ttft=%.3f tpot=%.4f out=%d e2e=%.3f",             first - t0, (end - first) / max(n - 1, 1), n, end - t0)

Then watch for four specific signatures, because each has one likely cause:

SymptomMost likely causeCheck
TTFT occasionally 10–30× normalModel was unloaded and reloaded from diskload_duration in the response; raise keep_alive
TTFT crept up over weeksPrompt template grew; prefix cache missesLog prompt token count over time
TPOT doubled at a fixed time of dayConcurrency, or a second model evicting the firstGET /api/ps; cap MAX_LOADED_MODELS
TPOT degrades within a single long conversationAttention over a growing KV cachePlot TPOT against context length; trim history

That last one surprises people. Decode cost is dominated by reading the weights, which is constant — but attention must also read the entire KV cache each step, and that grows linearly. At 2K of context the cache is 0.25 GiB against 4.58 GiB of weights and contributes about 5% of the memory traffic. At 32K it is 4 GiB against 4.58 and contributes nearly half. A conversation that starts at 95 tok/s and drifts to 60 by turn forty is behaving exactly as designed.

Deciding, and knowing when to stop

The practical sequence is short and rarely followed. Write down the constraint before you measure anything: the TTFT budget, the memory ceiling, the accuracy floor on your own task. Then measure a baseline at Q4_K_M with full offload, because that configuration is right often enough to be the sensible starting point. Then change one thing.

If the baseline meets the constraint, stop. This is the hardest part. There is always another quantisation to try and another flag to sweep, and the returns after full GPU offload and a sensible context size are small enough to be indistinguishable from measurement noise on a busy machine.

If it does not meet the constraint, the fix is determined by which phase is failing. Failing TTFT means prefill, which means either the prompt is too long, the prefix cache is being invalidated, or layers are still on the CPU. Failing TPOT means bandwidth, which means a smaller model, a smaller quantisation, or better hardware — and no amount of thread tuning. Failing on memory means the KV cache before it means the weights: quantising the cache to 8 bits halves the context term for a quality cost far below what dropping from Q4_K_M to Q3_K_M would inflict.

Keep the results in a file, dated, with the exact model file, flags, and hardware. Backends move quickly and a configuration that lost by 15% six months ago may win now. Without a written baseline you will re-run the same experiments every quarter and reach conclusions you have already reached — which is, in the end, the only genuinely wasted time in this whole exercise.