Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Design a full backend architecture for an LLM app (API + DB + model).
What you need to know
A system-design answer is judged on three things: clarifying questions that change the design, clear boundaries between components, and failure handling. Five questions come first, because each changes the architecture:
| question | why it changes the design |
|---|---|
| Requests per second and latency target? | streaming, caching, how many workers, which model tier |
| Static or constantly updated knowledge? | batch re-index vs streaming ingestion; freshness metrics |
| Must conversations persist? | session store, retention policy |
| Personal or regulated data? | redaction, region of storage and inference, audit logs |
| Cost ceiling per conversation? | model choice, caching, context budget |
The layers:
client | HTTPS + server-sent eventsAPI layer (FastAPI, async) auth, per-tenant rate limit, validation, streaming |orchestrator memory -> query rewrite -> retrieve -> prompt (versioned) | -> model adapter -> output guardrails -> stream +-- Redis response cache, rate-limit buckets, session state +-- Postgres + pgvector users, conversations, messages, documents, chunks +-- object storage raw uploaded files +-- queue + workers ingestion: parse -> chunk -> embed -> upsert +-- model adapter retries, backoff, circuit breaker, provider fallback |observability traces, token and cost metrics, nightly eval jobStateless API servers hold no conversation state in memory, so any server can handle any request and you scale by adding servers. State lives in Redis and Postgres.
Ingestion is asynchronous. Uploading a 400-page PDF returns immediately with a job id; a worker parses, chunks, embeds and upserts it. The request path never waits for embedding.
One database for a long time. Postgres with pgvector holds relational data and vectors together, which means one backup, one permission model and transactional updates. A separate vector database is justified only when the vector count or query load outgrows it.
The request path, as code with every component injected so each can be swapped or faked:
1import time2from collections.abc import Callable, Iterator3from dataclasses import dataclass45@dataclass6class Components:7 load_memory: Callable[[str], list[dict]] # Redis / Postgres8 rewrite: Callable[[list[dict], str], str] # small, cheap model9 retrieve: Callable[[str, int], list[str]] # pgvector + full-text, fused10 render: Callable[[str, list[str], list[dict]], str] # versioned prompt template11 generate: Callable[[str], Iterator[str]] # adapter: retries, fallback, stream12 check: Callable[[str], bool] # output guardrails13 save: Callable[[str, str, str, dict], None] # persist turn + usage, async in prod1415def handle_chat(c: Components, session_id: str, question: str) -> Iterator[str]:16 """One request: every stage is timed so the trace shows where the latency went."""17 timings, t = {}, time.perf_counter()18 def lap(name: str) -> None:19 nonlocal t20 now = time.perf_counter()21 timings[name], t = round((now - t) * 1000, 1), now2223 history = c.load_memory(session_id); lap("memory")24 query = c.rewrite(history, question) if history else question; lap("rewrite")25 chunks = c.retrieve(query, 5); lap("retrieve")26 prompt = c.render(question, chunks, history); lap("render")27 answer = []28 for delta in c.generate(prompt):29 answer.append(delta)30 yield delta # streamed to the client as it arrives31 lap("generate")32 text = "".join(answer)33 if not c.check(text):34 yield "\n[This answer was withdrawn by a safety check.]"35 c.save(session_id, question, text, timings)The tricky parts:
- A generator for the whole request means the API layer can stream deltas straight into server-sent events; the guardrail result is appended as a final event because earlier tokens are already on screen.
- The rewrite is skipped on the first turn — there is no history to resolve pronouns against, so it would be a wasted model call.
lap()records per-stage latency into the saved record. When p95 latency rises, you can see whether retrieval or generation moved.- Every dependency is a parameter. The same function runs in production with real clients and in tests with lambdas.
Complexity per request: one small model call for the rewrite (after the first turn), one retrieval (O(log n) with an ANN index, plus full-text search), one main model call, and a few Redis and Postgres round trips. Generation dominates latency and cost.
A real-life example
Wiring the flow with fakes shows the order of operations and what gets saved:
1saved = {}2fake = Components(3 load_memory=lambda sid: [{"role": "user", "content": "I ordered a phone"}],4 rewrite=lambda hist, q: "refund time for a phone order",5 retrieve=lambda q, k: ["Refunds for electronics take 7 days."],6 render=lambda q, chunks, hist: f"Context: {chunks[0]}\nQuestion: {q}",7 generate=lambda prompt: iter(["Your refund ", "takes 7 days."]),8 check=lambda text: "days" in text,9 save=lambda sid, q, a, timings: saved.update(session=sid, answer=a, stages=list(timings)),10)11print("".join(handle_chat(fake, "s-42", "how long for the refund?")))12# Your refund takes 7 days.13print(saved)14# {'session': 's-42', 'answer': 'Your refund takes 7 days.',15# 'stages': ['memory', 'rewrite', 'retrieve', 'render', 'generate']}A latency budget for a realistic version of the same request (illustrative targets, not measurements):
| stage | target p95 | notes |
|---|---|---|
| auth + rate limit | 5 ms | Redis |
| load memory | 10 ms | Redis |
| query rewrite | 400 ms | small model, short output |
| hybrid retrieval | 60 ms | pgvector HNSW + full-text, fused |
| render prompt | 1 ms | in memory |
| time to first token | 800 ms | main model, streaming |
| full answer | 4–6 s | depends on answer length |
The user sees text after roughly 1.3 seconds; everything else happens while they read.
An insurance company's policy assistant, serving a few thousand agents, fits comfortably in this design with a handful of API servers, one Postgres primary with a read replica, and Redis.
Follow-up questions to expect
- "What would you ship first?" — One FastAPI service, Postgres with pgvector, a synchronous ingestion script, one provider, and tracing with cost logging. Add the queue, hybrid search, re-ranker, semantic cache and fallback provider when a metric shows they are needed.
- "The provider has an outage — what happens?" — The adapter's circuit breaker opens, traffic goes to the fallback provider, the fallback rate alarm fires, and responses are logged with which provider served them.
- "How do you control cost?" — Per-tenant budgets enforced before the call, response and prompt caching, a context budget, cheaper models for rewrite and classification steps, and a cost-per-conversation dashboard.