Course Content
Local LLM Deployment and Quantization
3 sections · 7 lessons
Integrating Local Models with APIs
The demo works beautifully. A Flask route calls the local model, the answer streams back, everyone is impressed. Then two people use it at once, and the second request hangs for thirty-eight seconds with no output at all before it starts. A third arrives and the first user's stream stutters.
Nothing crashed. There is one model in memory, and by default it processes one request at a time. Request two waits for request one to finish generating all 300 of its tokens. Meanwhile a colleague reports that the very first request each morning takes eleven seconds before anything happens — because the runtime unloaded the model after five idle minutes and is reading 4.92 GB back off disk.
Neither problem is a bug in your code. Both are consequences of what a local model server actually is: a long-lived process holding several gigabytes of weights, whose concurrency and lifetime you must configure deliberately. The integration layer is where local deployment stops being about models and starts being about systems.
The shape of a local deployment
Almost every local setup has the same three tiers, and confusion about which tier owns what causes most of the trouble.
Client Your application Inference server Model(browser, CLI) -> (routing, auth, RAG, -> (Ollama / llama-server -> weights in prompt assembly) / LM Studio, port N) RAM or VRAM stateless stateless, restartable stateful, expensive 4.92 GB to start, holds the KV cacheThe inference server is the only stateful, expensive part. It takes seconds to start, holds gigabytes, and can serve a fixed number of concurrent requests. Your application should be cheap to restart and hold nothing. Merging the two — importing the model directly into your web process — is the decision that causes the most pain later, and it is worth being explicit about the trade-off.
| Separate server (Ollama, llama-server, LM Studio) | Embedded library (llama-cpp-python in-process) | |
|---|---|---|
| Restarting your app | Model stays loaded; restart is instant | Reloads several GB every time |
| Multiple apps sharing a model | Natural | Each process loads its own copy |
| Concurrency | Handled by the server's slot scheduler | You write the locking yourself |
| Deployment | Two processes to supervise | One binary, no daemon |
| Control over sampling internals | Whatever the API exposes | Total — logits processors, custom stopping |
| Best for | Anything with more than one user or more than one client | Desktop apps, CLI tools, air-gapped single binaries |
Keep the model in a process you do not restart, and your application in a process you restart freely. Almost every integration problem people hit is a violation of that line.
The Ollama API
Ollama's daemon listens on 127.0.0.1:11434. It exposes its own API and, separately, an OpenAI-compatible surface under /v1.
| Endpoint | Purpose |
|---|---|
POST /api/generate | Raw completion from a single prompt string |
POST /api/chat | Multi-turn conversation; applies the model's chat template |
POST /api/embed | Embedding vectors, for retrieval |
GET /api/tags | Models available on disk |
GET /api/ps | Models currently loaded, with size and expiry — the diagnostic endpoint |
POST /v1/chat/completions | OpenAI-compatible, so existing SDKs work unchanged |
1curl http://localhost:11434/api/chat -d '{2 "model": "llama3.1:8b",3 "messages": [{"role": "user", "content": "Name two causes of tail latency."}],4 "stream": false,5 "keep_alive": "30m",6 "options": {"temperature": 0.3, "num_ctx": 4096, "num_predict": 300}7}'Two fields there are load-bearing. keep_alive controls how long the model stays resident after the request; "30m" or -1 eliminates the cold-start problem at the cost of holding memory. num_ctx sizes the KV cache — for Llama 3 8B each token of context costs 128 KiB, so num_ctx: 4096 reserves 0.5 GiB and num_ctx: 32768 reserves 4 GiB whether you use it or not.
The response carries the timings you need for diagnosis, in nanoseconds:
1import requests, json23r = requests.post("http://localhost:11434/api/chat", json={4 "model": "llama3.1:8b",5 "messages": [{"role": "user", "content": "Explain KV caching briefly."}],6 "stream": False,7}).json()89load_s = r["load_duration"] / 1e9 # 0 if already resident10prefill = r["prompt_eval_duration"] / 1e9 # time to first token, essentially11decode = r["eval_duration"] / 1e912print(f"load {load_s:.2f}s prefill {prefill:.2f}s "13 f"decode {r['eval_count']/decode:.1f} tok/s")A non-zero load_duration means you paid a cold start. A large prompt_eval_duration means your prompt is long or your prefill is running on the CPU. These two numbers answer most "why is it slow" questions without any further instrumentation.
Streaming
Streaming is not a nicety. Time-to-first-token is typically 0.2 s while a 300-token reply takes 3 s; streaming converts a three-second wait into a two-tenths-of-a-second wait. Ollama streams newline-delimited JSON:
1import json, requests23def stream_chat(messages, model="llama3.1:8b"):4 with requests.post("http://localhost:11434/api/chat",5 json={"model": model, "messages": messages, "stream": True},6 stream=True, timeout=300) as r:7 r.raise_for_status()8 for line in r.iter_lines():9 if not line:10 continue11 obj = json.loads(line)12 if obj.get("done"):13 return14 yield obj["message"]["content"]1516for piece in stream_chat([{"role": "user", "content": "Count to ten slowly."}]):17 print(piece, end="", flush=True)Set a generous timeout. The default in most HTTP clients is far shorter than a long local generation, and a timeout mid-stream produces a truncated answer that looks like a model failure.
Conversation state is yours
The server is stateless between requests. It does not remember your conversation; you resend the whole message list every time. That means you own the trimming, and you must do it, because exceeding num_ctx silently drops the oldest tokens — usually including your system prompt.
1class Conversation:2 def __init__(self, system, max_turns=12):3 self.system = {"role": "system", "content": system}4 self.turns = []5 self.max_turns = max_turns67 def ask(self, text):8 self.turns.append({"role": "user", "content": text})9 reply = "".join(stream_chat([self.system] + self.turns))10 self.turns.append({"role": "assistant", "content": reply})11 # Trim oldest pairs, never the system message12 if len(self.turns) > self.max_turns * 2:13 self.turns = self.turns[-self.max_turns * 2:]14 return replyOpenAI-compatible endpoints
LM Studio serves on 127.0.0.1:1234/v1, Ollama mirrors one at 11434/v1, and llama-server provides one too. All three accept the OpenAI SDK, which is the single most useful fact in this whole area: your application can be written once and pointed anywhere.
1import os2from openai import OpenAI34# One line of configuration switches between local runtimes or a hosted provider5client = OpenAI(6 base_url=os.getenv("LLM_BASE_URL", "http://localhost:11434/v1"),7 api_key=os.getenv("LLM_API_KEY", "not-needed-locally"),8)910stream = client.chat.completions.create(11 model=os.getenv("LLM_MODEL", "llama3.1:8b"),12 messages=[{"role": "user", "content": "Explain continuous batching."}],13 temperature=0.3, max_tokens=400, stream=True,14)15for chunk in stream:16 delta = chunk.choices[0].delta.content17 if delta:18 print(delta, end="", flush=True)Compatibility is close but not complete. Function calling works on some models and silently does nothing on others; logprobs, n>1, and seed support varies by server. Treat the OpenAI shape as a transport convention, and test any parameter beyond messages, temperature, and max tokens against the server you are actually running.
Building your own server
You write your own when you need something the standard servers do not offer: model-specific preprocessing, retrieval baked into the endpoint, custom authentication, or per-tenant routing. The pattern is FastAPI in front of llama-cpp-python, with one critical detail.
1import asyncio, json2from contextlib import asynccontextmanager3from fastapi import FastAPI, HTTPException4from fastapi.responses import StreamingResponse5from pydantic import BaseModel6from llama_cpp import Llama78STATE = {}910@asynccontextmanager11async def lifespan(app: FastAPI):12 STATE["llm"] = Llama(model_path="models/llama-3.1-8b-Q4_K_M.gguf",13 n_ctx=8192, n_gpu_layers=-1, n_threads=8, verbose=False)14 STATE["lock"] = asyncio.Lock() # the model is NOT thread-safe15 yield16 STATE.clear()1718app = FastAPI(lifespan=lifespan)1920class ChatRequest(BaseModel):21 messages: list[dict]22 temperature: float = 0.323 max_tokens: int = 51224 stream: bool = False2526@app.post("/chat")27async def chat(req: ChatRequest):28 if not req.messages:29 raise HTTPException(400, "messages must not be empty")30 llm, lock = STATE["llm"], STATE["lock"]3132 if not req.stream:33 async with lock:34 out = await asyncio.to_thread(35 llm.create_chat_completion,36 messages=req.messages, temperature=req.temperature,37 max_tokens=req.max_tokens)38 return {"content": out["choices"][0]["message"]["content"],39 "usage": out["usage"]}4041 async def gen():42 async with lock:43 it = llm.create_chat_completion(44 messages=req.messages, temperature=req.temperature,45 max_tokens=req.max_tokens, stream=True)46 # each next() is a blocking C call: run it off the event loop47 while (chunk := await asyncio.to_thread(next, it, None)) is not None:48 delta = chunk["choices"][0]["delta"].get("content")49 if delta:50 yield f"data: {json.dumps({'t': delta})}\n\n"51 yield "data: [DONE]\n\n"5253 return StreamingResponse(gen(), media_type="text/event-stream")5455@app.get("/health")56def health():57 return {"ok": "llm" in STATE}The lock is the critical detail. A Llama object holds one KV cache and one set of scratch buffers; two concurrent calls corrupt each other's state and produce interleaved nonsense or a segfault. The asyncio.to_thread calls matter too, in both the plain and the streaming path — inference is a long blocking C call, and running it directly on the event loop freezes every other request, including /health.
Note what this buys and what it does not. You now have exactly one request in flight at a time. That is correct and safe, and it is strictly worse than what Ollama or llama-server give you for free, which is the next section.
Concurrency, and what it costs in memory
Serving several users at once is not a matter of threads. It requires the server to hold a separate KV cache per in-flight request, because each conversation has its own history.
For Llama 3 8B at 128 KiB per token, four parallel slots of 4,096 tokens each cost:
4 × 4096 × 128 KiB = 2,147,483,648 bytes = 2 GiB
Add 4.58 GiB of weights and roughly 0.6 GiB of buffers and you need 7.2 GiB — comfortable on a 12 GB card, impossible on an 8 GB one.
1# llama-server: with an explicit --parallel, --ctx-size is the TOTAL, divided among slots2./build/bin/llama-server -m model-Q4_K_M.gguf \3 --ctx-size 16384 --parallel 4 --cont-batching -ngl 99 --port 80804# 4 slots of 4096 tokens each56# Ollama: per-model settings via environment7export OLLAMA_NUM_PARALLEL=48export OLLAMA_MAX_LOADED_MODELS=19export OLLAMA_KEEP_ALIVE=-110ollama serveThe payoff is better than it looks, because of how decoding works. Generating a token requires reading every weight once; with four sequences batched together, one pass over the weights produces four tokens. The expensive part is amortised.
| Concurrent requests | Per-request rate | Aggregate | Extra KV cache |
|---|---|---|---|
| 1 | 95 tok/s | 95 tok/s | 0.5 GiB |
| 2 | ~70 tok/s | ~140 tok/s | 1.0 GiB |
| 4 | ~50 tok/s | ~200 tok/s | 2.0 GiB |
| 8 | ~30 tok/s | ~240 tok/s | 4.0 GiB |
Aggregate throughput more than doubles while per-user speed halves. Whether that is a good trade depends entirely on what you are building: an interactive assistant should keep parallelism low so individual replies stay fast, while a batch document processor should push it as high as memory allows.
Continuous batching (--cont-batching, on by default in current llama-server builds) is worth having wherever it exists. Without it, a new request waits for the current batch to finish entirely; with it, a finished sequence is evicted from the batch and a queued one takes its slot mid-flight. On mixed workloads where some replies are 20 tokens and others 800, it roughly halves the average wait.
Caching, and the one that matters most
Two very different caches both get called "caching", and confusing them wastes effort.
Response caching is what you would expect: hash the request, store the answer. It helps only on exact repeats, which are rarer than people assume in conversational systems, but it is nearly free to add.
1import hashlib, json2from functools import lru_cache34def key(messages, temperature, max_tokens):5 return hashlib.sha256(json.dumps(6 [messages, temperature, max_tokens], sort_keys=True).encode()).hexdigest()78@lru_cache(maxsize=512)9def cached_completion(k, payload_json):10 payload = json.loads(payload_json)11 return "".join(stream_chat(payload["messages"]))12# Only cache deterministic calls. Caching temperature=0.9 output makes it deterministic13# by accident, which users notice and report as a bug.Prefix caching is the one that actually moves latency. If every request begins with the same 900-token system prompt and few-shot examples, the server can retain that prefix's KV cache and skip re-prefilling it. On a CPU-bound setup where prefill runs at 30 tokens per second, that is 30 seconds saved on every single call.
Both llama-server and Ollama do this automatically when consecutive requests share a prefix. The practical consequence for your code is a design rule: put everything static at the front of the prompt and everything variable at the end. A prompt that interpolates the current timestamp into the system message defeats prefix caching completely, and the cost is invisible until you measure prefill time.
Structure prompts so the unchanging part comes first. A single variable token near the top of a 900-token prefix throws away the entire cached prefill.
Streaming to a browser
Server-sent events are the right transport: one-directional, plain HTTP, and supported everywhere. The server code above already emits them; the client is short.
1async function ask(messages, onToken) {2 const res = await fetch("/chat", {3 method: "POST",4 headers: { "Content-Type": "application/json" },5 body: JSON.stringify({ messages, stream: true }),6 });7 const reader = res.body.getReader();8 const decoder = new TextDecoder();9 let buffer = "";1011 while (true) {12 const { done, value } = await reader.read();13 if (done) break;14 buffer += decoder.decode(value, { stream: true });15 const lines = buffer.split("\n\n");16 buffer = lines.pop(); // keep the incomplete tail17 for (const line of lines) {18 if (!line.startsWith("data: ")) continue;19 const body = line.slice(6);20 if (body === "[DONE]") return;21 onToken(JSON.parse(body).t);22 }23 }24}The buffering detail catches people out. A network packet can split an SSE frame in half, so you must accumulate and only process complete \n\n-terminated frames. Parsing each chunk as it arrives works in development, where responses arrive in one piece, and produces intermittent JSON errors in production.
Two more things worth building in from the start. Render tokens into a text node rather than reassigning innerHTML on every token — at 95 tokens per second that is 95 full re-parses per second and it will visibly stutter. And wire up an AbortController so a user can stop a long generation; on a single-GPU deployment, one abandoned 2,000-token response blocks the queue for everyone.
Making it survive contact with users
The habits that separate a working demo from something you can leave running come down to a short list, and each maps to a failure described above.
Configure lifetime explicitly. Set keep_alive to -1 or a long duration for anything interactive. The default five-minute unload turns the first request after lunch into an eleven-second stall, and it is the single most common complaint about self-hosted assistants.
Size parallelism against memory, not optimism. Multiply your slot count by your context length by the per-token cache cost, add the weights, and check it fits. Getting this wrong produces an out-of-memory error under load — which is to say, in front of users, not in testing.
Health-check the model, not the port. A server that has lost its GPU allocation still accepts TCP connections. A useful health check sends a two-token generation and asserts it returns within a threshold. Poll /api/ps to confirm the model you expect is resident and has not been evicted by a second model.
Put something in front of it. None of these servers authenticate. Binding OLLAMA_HOST to 0.0.0.0 without a reverse proxy publishes an unauthenticated inference endpoint to your network. Terminate TLS, add an API key, and rate-limit at the proxy.
Keep the base URL in configuration. The reason to write against the OpenAI-compatible interface even when you are certain you will always use Ollama is that you will not always use Ollama. Moving from a laptop to a GPU box, or falling back to a hosted API when the local machine is down, should be an environment variable — and if it is, you can run both and compare them on real traffic, which is how you find out whether the local model is actually good enough.