Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Design a full backend architecture for an LLM app (API + DB + model).


What you need to know

A system-design answer is judged on three things: clarifying questions that change the design, clear boundaries between components, and failure handling. Five questions come first, because each changes the architecture:

questionwhy it changes the design
Requests per second and latency target?streaming, caching, how many workers, which model tier
Static or constantly updated knowledge?batch re-index vs streaming ingestion; freshness metrics
Must conversations persist?session store, retention policy
Personal or regulated data?redaction, region of storage and inference, audit logs
Cost ceiling per conversation?model choice, caching, context budget

The layers:

Text
client  |  HTTPS + server-sent eventsAPI layer (FastAPI, async)       auth, per-tenant rate limit, validation, streaming  |orchestrator                     memory -> query rewrite -> retrieve -> prompt (versioned)  |                              -> model adapter -> output guardrails -> stream  +-- Redis                      response cache, rate-limit buckets, session state  +-- Postgres + pgvector        users, conversations, messages, documents, chunks  +-- object storage             raw uploaded files  +-- queue + workers            ingestion: parse -> chunk -> embed -> upsert  +-- model adapter              retries, backoff, circuit breaker, provider fallback  |observability                    traces, token and cost metrics, nightly eval job

Stateless API servers hold no conversation state in memory, so any server can handle any request and you scale by adding servers. State lives in Redis and Postgres.

Ingestion is asynchronous. Uploading a 400-page PDF returns immediately with a job id; a worker parses, chunks, embeds and upserts it. The request path never waits for embedding.

One database for a long time. Postgres with pgvector holds relational data and vectors together, which means one backup, one permission model and transactional updates. A separate vector database is justified only when the vector count or query load outgrows it.

The request path, as code with every component injected so each can be swapped or faked:

Python
import timefrom collections.abc import Callable, Iteratorfrom dataclasses import dataclass@dataclassclass Components:    load_memory: Callable[[str], list[dict]]          # Redis / Postgres    rewrite: Callable[[list[dict], str], str]         # small, cheap model    retrieve: Callable[[str, int], list[str]]         # pgvector + full-text, fused    render: Callable[[str, list[str], list[dict]], str]   # versioned prompt template    generate: Callable[[str], Iterator[str]]          # adapter: retries, fallback, stream    check: Callable[[str], bool]                      # output guardrails    save: Callable[[str, str, str, dict], None]       # persist turn + usage, async in proddef handle_chat(c: Components, session_id: str, question: str) -> Iterator[str]:    """One request: every stage is timed so the trace shows where the latency went."""    timings, t = {}, time.perf_counter()    def lap(name: str) -> None:        nonlocal t        now = time.perf_counter()        timings[name], t = round((now - t) * 1000, 1), now    history = c.load_memory(session_id); lap("memory")    query = c.rewrite(history, question) if history else question; lap("rewrite")    chunks = c.retrieve(query, 5); lap("retrieve")    prompt = c.render(question, chunks, history); lap("render")    answer = []    for delta in c.generate(prompt):        answer.append(delta)        yield delta                                    # streamed to the client as it arrives    lap("generate")    text = "".join(answer)    if not c.check(text):        yield "\n[This answer was withdrawn by a safety check.]"    c.save(session_id, question, text, timings)

The tricky parts:

  • A generator for the whole request means the API layer can stream deltas straight into server-sent events; the guardrail result is appended as a final event because earlier tokens are already on screen.
  • The rewrite is skipped on the first turn — there is no history to resolve pronouns against, so it would be a wasted model call.
  • lap() records per-stage latency into the saved record. When p95 latency rises, you can see whether retrieval or generation moved.
  • Every dependency is a parameter. The same function runs in production with real clients and in tests with lambdas.

Complexity per request: one small model call for the rewrite (after the first turn), one retrieval (O(log n) with an ANN index, plus full-text search), one main model call, and a few Redis and Postgres round trips. Generation dominates latency and cost.

A real-life example

Wiring the flow with fakes shows the order of operations and what gets saved:

Python
saved = {}fake = Components(    load_memory=lambda sid: [{"role": "user", "content": "I ordered a phone"}],    rewrite=lambda hist, q: "refund time for a phone order",    retrieve=lambda q, k: ["Refunds for electronics take 7 days."],    render=lambda q, chunks, hist: f"Context: {chunks[0]}\nQuestion: {q}",    generate=lambda prompt: iter(["Your refund ", "takes 7 days."]),    check=lambda text: "days" in text,    save=lambda sid, q, a, timings: saved.update(session=sid, answer=a, stages=list(timings)),)print("".join(handle_chat(fake, "s-42", "how long for the refund?")))# Your refund takes 7 days.print(saved)# {'session': 's-42', 'answer': 'Your refund takes 7 days.',#  'stages': ['memory', 'rewrite', 'retrieve', 'render', 'generate']}

A latency budget for a realistic version of the same request (illustrative targets, not measurements):

stagetarget p95notes
auth + rate limit5 msRedis
load memory10 msRedis
query rewrite400 mssmall model, short output
hybrid retrieval60 mspgvector HNSW + full-text, fused
render prompt1 msin memory
time to first token800 msmain model, streaming
full answer4–6 sdepends on answer length

The user sees text after roughly 1.3 seconds; everything else happens while they read.

An insurance company's policy assistant, serving a few thousand agents, fits comfortably in this design with a handful of API servers, one Postgres primary with a read replica, and Redis.

Follow-up questions to expect

  • "What would you ship first?" — One FastAPI service, Postgres with pgvector, a synchronous ingestion script, one provider, and tracing with cost logging. Add the queue, hybrid search, re-ranker, semantic cache and fallback provider when a metric shows they are needed.
  • "The provider has an outage — what happens?" — The adapter's circuit breaker opens, traffic goes to the fallback provider, the fallback rate alarm fires, and responses are logged with which provider served them.
  • "How do you control cost?" — Per-tenant budgets enforced before the call, response and prompt caching, a context budget, cheaper models for rewrite and classification steps, and a cost-per-conversation dashboard.