Local LLM Deployment and Quantization

Integrating Local Models with APIs


The demo works beautifully. A Flask route calls the local model, the answer streams back, everyone is impressed. Then two people use it at once, and the second request hangs for thirty-eight seconds with no output at all before it starts. A third arrives and the first user's stream stutters.

Nothing crashed. There is one model in memory, and by default it processes one request at a time. Request two waits for request one to finish generating all 300 of its tokens. Meanwhile a colleague reports that the very first request each morning takes eleven seconds before anything happens — because the runtime unloaded the model after five idle minutes and is reading 4.92 GB back off disk.

Neither problem is a bug in your code. Both are consequences of what a local model server actually is: a long-lived process holding several gigabytes of weights, whose concurrency and lifetime you must configure deliberately. The integration layer is where local deployment stops being about models and starts being about systems.

Why the second user waits 38 seconds for nothingRequest arrivesQueue — one slotper loaded modelPrefillthe promptDecode,streaming tokensKV cache freed,slot releasedA second request cannot start prefill until the first finishes decoding its last token.
Time-to-first-token for request two is request one's total time — the fix is a queue you control, not more workers.

The shape of a local deployment

Almost every local setup has the same three tiers, and confusion about which tier owns what causes most of the trouble.

Text
Client            Your application            Inference server         Model(browser, CLI) -> (routing, auth, RAG,   ->  (Ollama / llama-server  -> weights in                   prompt assembly)           / LM Studio, port N)      RAM or VRAM  stateless         stateless, restartable      stateful, expensive       4.92 GB                                                to start, holds the                                                KV cache

The inference server is the only stateful, expensive part. It takes seconds to start, holds gigabytes, and can serve a fixed number of concurrent requests. Your application should be cheap to restart and hold nothing. Merging the two — importing the model directly into your web process — is the decision that causes the most pain later, and it is worth being explicit about the trade-off.

Separate server (Ollama, llama-server, LM Studio)Embedded library (llama-cpp-python in-process)
Restarting your appModel stays loaded; restart is instantReloads several GB every time
Multiple apps sharing a modelNaturalEach process loads its own copy
ConcurrencyHandled by the server's slot schedulerYou write the locking yourself
DeploymentTwo processes to superviseOne binary, no daemon
Control over sampling internalsWhatever the API exposesTotal — logits processors, custom stopping
Best forAnything with more than one user or more than one clientDesktop apps, CLI tools, air-gapped single binaries

Keep the model in a process you do not restart, and your application in a process you restart freely. Almost every integration problem people hit is a violation of that line.

The Ollama API

Ollama's daemon listens on 127.0.0.1:11434. It exposes its own API and, separately, an OpenAI-compatible surface under /v1.

EndpointPurpose
POST /api/generateRaw completion from a single prompt string
POST /api/chatMulti-turn conversation; applies the model's chat template
POST /api/embedEmbedding vectors, for retrieval
GET /api/tagsModels available on disk
GET /api/psModels currently loaded, with size and expiry — the diagnostic endpoint
POST /v1/chat/completionsOpenAI-compatible, so existing SDKs work unchanged
Bash
curl http://localhost:11434/api/chat -d '{  "model": "llama3.1:8b",  "messages": [{"role": "user", "content": "Name two causes of tail latency."}],  "stream": false,  "keep_alive": "30m",  "options": {"temperature": 0.3, "num_ctx": 4096, "num_predict": 300}}'

Two fields there are load-bearing. keep_alive controls how long the model stays resident after the request; "30m" or -1 eliminates the cold-start problem at the cost of holding memory. num_ctx sizes the KV cache — for Llama 3 8B each token of context costs 128 KiB, so num_ctx: 4096 reserves 0.5 GiB and num_ctx: 32768 reserves 4 GiB whether you use it or not.

The response carries the timings you need for diagnosis, in nanoseconds:

Python
import requests, jsonr = requests.post("http://localhost:11434/api/chat", json={    "model": "llama3.1:8b",    "messages": [{"role": "user", "content": "Explain KV caching briefly."}],    "stream": False,}).json()load_s   = r["load_duration"] / 1e9          # 0 if already residentprefill  = r["prompt_eval_duration"] / 1e9   # time to first token, essentiallydecode   = r["eval_duration"] / 1e9print(f"load {load_s:.2f}s  prefill {prefill:.2f}s  "      f"decode {r['eval_count']/decode:.1f} tok/s")

A non-zero load_duration means you paid a cold start. A large prompt_eval_duration means your prompt is long or your prefill is running on the CPU. These two numbers answer most "why is it slow" questions without any further instrumentation.

Streaming

Streaming is not a nicety. Time-to-first-token is typically 0.2 s while a 300-token reply takes 3 s; streaming converts a three-second wait into a two-tenths-of-a-second wait. Ollama streams newline-delimited JSON:

Python
import json, requestsdef stream_chat(messages, model="llama3.1:8b"):    with requests.post("http://localhost:11434/api/chat",                       json={"model": model, "messages": messages, "stream": True},                       stream=True, timeout=300) as r:        r.raise_for_status()        for line in r.iter_lines():            if not line:                continue            obj = json.loads(line)            if obj.get("done"):                return            yield obj["message"]["content"]for piece in stream_chat([{"role": "user", "content": "Count to ten slowly."}]):    print(piece, end="", flush=True)

Set a generous timeout. The default in most HTTP clients is far shorter than a long local generation, and a timeout mid-stream produces a truncated answer that looks like a model failure.

Conversation state is yours

The server is stateless between requests. It does not remember your conversation; you resend the whole message list every time. That means you own the trimming, and you must do it, because exceeding num_ctx silently drops the oldest tokens — usually including your system prompt.

Python
class Conversation:    def __init__(self, system, max_turns=12):        self.system = {"role": "system", "content": system}        self.turns = []        self.max_turns = max_turns    def ask(self, text):        self.turns.append({"role": "user", "content": text})        reply = "".join(stream_chat([self.system] + self.turns))        self.turns.append({"role": "assistant", "content": reply})        # Trim oldest pairs, never the system message        if len(self.turns) > self.max_turns * 2:            self.turns = self.turns[-self.max_turns * 2:]        return reply

OpenAI-compatible endpoints

LM Studio serves on 127.0.0.1:1234/v1, Ollama mirrors one at 11434/v1, and llama-server provides one too. All three accept the OpenAI SDK, which is the single most useful fact in this whole area: your application can be written once and pointed anywhere.

Python
import osfrom openai import OpenAI# One line of configuration switches between local runtimes or a hosted providerclient = OpenAI(    base_url=os.getenv("LLM_BASE_URL", "http://localhost:11434/v1"),    api_key=os.getenv("LLM_API_KEY", "not-needed-locally"),)stream = client.chat.completions.create(    model=os.getenv("LLM_MODEL", "llama3.1:8b"),    messages=[{"role": "user", "content": "Explain continuous batching."}],    temperature=0.3, max_tokens=400, stream=True,)for chunk in stream:    delta = chunk.choices[0].delta.content    if delta:        print(delta, end="", flush=True)

Compatibility is close but not complete. Function calling works on some models and silently does nothing on others; logprobs, n>1, and seed support varies by server. Treat the OpenAI shape as a transport convention, and test any parameter beyond messages, temperature, and max tokens against the server you are actually running.

Building your own server

You write your own when you need something the standard servers do not offer: model-specific preprocessing, retrieval baked into the endpoint, custom authentication, or per-tenant routing. The pattern is FastAPI in front of llama-cpp-python, with one critical detail.

Python
import asyncio, jsonfrom contextlib import asynccontextmanagerfrom fastapi import FastAPI, HTTPExceptionfrom fastapi.responses import StreamingResponsefrom pydantic import BaseModelfrom llama_cpp import LlamaSTATE = {}@asynccontextmanagerasync def lifespan(app: FastAPI):    STATE["llm"] = Llama(model_path="models/llama-3.1-8b-Q4_K_M.gguf",                         n_ctx=8192, n_gpu_layers=-1, n_threads=8, verbose=False)    STATE["lock"] = asyncio.Lock()      # the model is NOT thread-safe    yield    STATE.clear()app = FastAPI(lifespan=lifespan)class ChatRequest(BaseModel):    messages: list[dict]    temperature: float = 0.3    max_tokens: int = 512    stream: bool = False@app.post("/chat")async def chat(req: ChatRequest):    if not req.messages:        raise HTTPException(400, "messages must not be empty")    llm, lock = STATE["llm"], STATE["lock"]    if not req.stream:        async with lock:            out = await asyncio.to_thread(                llm.create_chat_completion,                messages=req.messages, temperature=req.temperature,                max_tokens=req.max_tokens)        return {"content": out["choices"][0]["message"]["content"],                "usage": out["usage"]}    async def gen():        async with lock:            it = llm.create_chat_completion(                messages=req.messages, temperature=req.temperature,                max_tokens=req.max_tokens, stream=True)            # each next() is a blocking C call: run it off the event loop            while (chunk := await asyncio.to_thread(next, it, None)) is not None:                delta = chunk["choices"][0]["delta"].get("content")                if delta:                    yield f"data: {json.dumps({'t': delta})}\n\n"        yield "data: [DONE]\n\n"    return StreamingResponse(gen(), media_type="text/event-stream")@app.get("/health")def health():    return {"ok": "llm" in STATE}

The lock is the critical detail. A Llama object holds one KV cache and one set of scratch buffers; two concurrent calls corrupt each other's state and produce interleaved nonsense or a segfault. The asyncio.to_thread calls matter too, in both the plain and the streaming path — inference is a long blocking C call, and running it directly on the event loop freezes every other request, including /health.

Note what this buys and what it does not. You now have exactly one request in flight at a time. That is correct and safe, and it is strictly worse than what Ollama or llama-server give you for free, which is the next section.

Concurrency, and what it costs in memory

Serving several users at once is not a matter of threads. It requires the server to hold a separate KV cache per in-flight request, because each conversation has its own history.

For Llama 3 8B at 128 KiB per token, four parallel slots of 4,096 tokens each cost:

4 × 4096 × 128 KiB = 2,147,483,648 bytes = 2 GiB

Add 4.58 GiB of weights and roughly 0.6 GiB of buffers and you need 7.2 GiB — comfortable on a 12 GB card, impossible on an 8 GB one.

Bash
# llama-server: with an explicit --parallel, --ctx-size is the TOTAL, divided among slots./build/bin/llama-server -m model-Q4_K_M.gguf \  --ctx-size 16384 --parallel 4 --cont-batching -ngl 99 --port 8080# 4 slots of 4096 tokens each# Ollama: per-model settings via environmentexport OLLAMA_NUM_PARALLEL=4export OLLAMA_MAX_LOADED_MODELS=1export OLLAMA_KEEP_ALIVE=-1ollama serve

The payoff is better than it looks, because of how decoding works. Generating a token requires reading every weight once; with four sequences batched together, one pass over the weights produces four tokens. The expensive part is amortised.

Concurrent requestsPer-request rateAggregateExtra KV cache
195 tok/s95 tok/s0.5 GiB
2~70 tok/s~140 tok/s1.0 GiB
4~50 tok/s~200 tok/s2.0 GiB
8~30 tok/s~240 tok/s4.0 GiB

Aggregate throughput more than doubles while per-user speed halves. Whether that is a good trade depends entirely on what you are building: an interactive assistant should keep parallelism low so individual replies stay fast, while a batch document processor should push it as high as memory allows.

Continuous batching (--cont-batching, on by default in current llama-server builds) is worth having wherever it exists. Without it, a new request waits for the current batch to finish entirely; with it, a finished sequence is evicted from the batch and a queued one takes its slot mid-flight. On mixed workloads where some replies are 20 tokens and others 800, it roughly halves the average wait.

Caching, and the one that matters most

Two very different caches both get called "caching", and confusing them wastes effort.

Response caching is what you would expect: hash the request, store the answer. It helps only on exact repeats, which are rarer than people assume in conversational systems, but it is nearly free to add.

Python
import hashlib, jsonfrom functools import lru_cachedef key(messages, temperature, max_tokens):    return hashlib.sha256(json.dumps(        [messages, temperature, max_tokens], sort_keys=True).encode()).hexdigest()@lru_cache(maxsize=512)def cached_completion(k, payload_json):    payload = json.loads(payload_json)    return "".join(stream_chat(payload["messages"]))# Only cache deterministic calls. Caching temperature=0.9 output makes it deterministic# by accident, which users notice and report as a bug.

Prefix caching is the one that actually moves latency. If every request begins with the same 900-token system prompt and few-shot examples, the server can retain that prefix's KV cache and skip re-prefilling it. On a CPU-bound setup where prefill runs at 30 tokens per second, that is 30 seconds saved on every single call.

Both llama-server and Ollama do this automatically when consecutive requests share a prefix. The practical consequence for your code is a design rule: put everything static at the front of the prompt and everything variable at the end. A prompt that interpolates the current timestamp into the system message defeats prefix caching completely, and the cost is invisible until you measure prefill time.

Structure prompts so the unchanging part comes first. A single variable token near the top of a 900-token prefix throws away the entire cached prefill.

Streaming to a browser

Server-sent events are the right transport: one-directional, plain HTTP, and supported everywhere. The server code above already emits them; the client is short.

JavaScript
async function ask(messages, onToken) {  const res = await fetch("/chat", {    method: "POST",    headers: { "Content-Type": "application/json" },    body: JSON.stringify({ messages, stream: true }),  });  const reader = res.body.getReader();  const decoder = new TextDecoder();  let buffer = "";  while (true) {    const { done, value } = await reader.read();    if (done) break;    buffer += decoder.decode(value, { stream: true });    const lines = buffer.split("\n\n");    buffer = lines.pop();                       // keep the incomplete tail    for (const line of lines) {      if (!line.startsWith("data: ")) continue;      const body = line.slice(6);      if (body === "[DONE]") return;      onToken(JSON.parse(body).t);    }  }}

The buffering detail catches people out. A network packet can split an SSE frame in half, so you must accumulate and only process complete \n\n-terminated frames. Parsing each chunk as it arrives works in development, where responses arrive in one piece, and produces intermittent JSON errors in production.

Two more things worth building in from the start. Render tokens into a text node rather than reassigning innerHTML on every token — at 95 tokens per second that is 95 full re-parses per second and it will visibly stutter. And wire up an AbortController so a user can stop a long generation; on a single-GPU deployment, one abandoned 2,000-token response blocks the queue for everyone.

Making it survive contact with users

The habits that separate a working demo from something you can leave running come down to a short list, and each maps to a failure described above.

Configure lifetime explicitly. Set keep_alive to -1 or a long duration for anything interactive. The default five-minute unload turns the first request after lunch into an eleven-second stall, and it is the single most common complaint about self-hosted assistants.

Size parallelism against memory, not optimism. Multiply your slot count by your context length by the per-token cache cost, add the weights, and check it fits. Getting this wrong produces an out-of-memory error under load — which is to say, in front of users, not in testing.

Health-check the model, not the port. A server that has lost its GPU allocation still accepts TCP connections. A useful health check sends a two-token generation and asserts it returns within a threshold. Poll /api/ps to confirm the model you expect is resident and has not been evicted by a second model.

Put something in front of it. None of these servers authenticate. Binding OLLAMA_HOST to 0.0.0.0 without a reverse proxy publishes an unauthenticated inference endpoint to your network. Terminate TLS, add an API key, and rate-limit at the proxy.

Keep the base URL in configuration. The reason to write against the OpenAI-compatible interface even when you are certain you will always use Ollama is that you will not always use Ollama. Moving from a laptop to a GPU box, or falling back to a hosted API when the local machine is down, should be an environment variable — and if it is, you can run both and compare them on real traffic, which is how you find out whether the local model is actually good enough.