Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

ChatGPT assistant: context, tools and serving at scale


The assistant is framed, its metrics are set, and the model — built or bought — is understood. What remains is everything around the model that decides whether the product works: what goes into the context on each turn, how the model acts on the world through tools, and how 50 million conversations a day are served at a bill someone will sign.

Managing context

The constraint that shapes every real chatbot, and the component candidates most often leave out. A model has no memory between calls. Every turn, the entire conversation is re-sent.

An 8K window, filled from the bottomSystemprompt — 400Tool schemas — 900Retrievedmemory — 1,200Conversation— 5,800Left for thereply — 700topbottomEvery turn re-sends the whole stack; the model remembers nothing on its own.
Summarising the history frees tokens but invalidates the prefix cache you were relying on.

The budget, which fills faster than people expect

Take a 16,000-token window and a realistic assistant:

ComponentTokens
System prompt (role, policy, style, formatting rules)900
Tool definitions (six tools with schemas)1,400
Retrieved documents (four passages)3,000
Reserved for the response1,000
Left for conversation history9,700

At roughly 400 tokens per exchange, that is about 24 turns. A long support conversation exceeds it, and a conversation that spans days certainly does.

The four strategies, and what each loses

1. Sliding window. Keep the most recent N turns, drop the oldest. Trivial, fast, and it silently discards the beginning of the conversation — which is frequently where the user stated what they were trying to do.

2. Rolling summarisation. When history exceeds a threshold, summarise the oldest turns into a paragraph and keep that in place of them. Preserves the thread of the conversation at a fraction of the tokens. Costs an extra model call, and loses specifics — exact numbers, exact names, exact wording — which is precisely what users later refer back to.

3. Retrieval over the conversation. Embed every past turn, and at each new message retrieve the few most relevant ones. Excellent for long-running conversations where the relevant history is old and specific. It breaks conversational flow, because retrieved turns arrive without their neighbours, and it needs the same machinery as Section 5 (Retrieval-Augmented Generation).

4. Structured memory. Extract durable facts as key-value entries — user prefers metric units, user's account tier is Enterprise, user is debugging a webhook failure — and carry those forward instead of raw text. Compact and durable. It requires an extraction step that can be wrong, and wrong facts persist and compound.

The recommendation: combine. Recent turns verbatim, a rolling summary of the middle, structured facts pinned at the top, and retrieval over the far past when the conversation is long enough to warrant it.

The cache interaction nobody mentions

Here is a consequence that separates people who have run these systems from people who have read about them.

Prefix caching (see The inference cost and latency budget and the Smart Compose serving lesson) reuses the key-value cache for a prompt prefix that has not changed. It is the largest single cost saving available in a chatbot, because the system prompt, tool definitions, and earlier turns repeat on every call.

It only works while the prefix is byte-identical. Rolling summarisation rewrites the front of the context, which invalidates the cache for the entire conversation from that point on. So a summarisation step that saves 4,000 tokens of context can cost more than it saves by throwing away cache reuse on every subsequent turn.

Two design responses. Summarise rarely and in large chunks rather than continuously, so you pay the invalidation once. And order the context so the stable parts come first — system prompt, then tool definitions, then structured memory, then summary, then recent turns, then the new message. Everything volatile at the end, everything stable at the front. That ordering is worth real money and costs nothing.

Tool use and agentic behaviour

A model that can only produce text is limited to what it learned during training. Give it the ability to call a function — search, a calculator, an order-lookup API — and it can act on current, private, and computed information. This is also where the failure modes get interesting and where the safety surface expands most.

The loop

The mechanism is small. The model does not execute anything; it emits a request that your code executes.

  1. Describe the tools. Names, purposes, and parameter schemas go into the context.
  2. Decide. The model produces either a normal response or a structured tool call: lookup_order(order_id="A-40219").
  3. Call. Your code validates the arguments, checks permissions, and executes.
  4. Observe. The result is appended to the context as a tool-result message.
  5. Repeat from step 2, until the model produces a normal response or a limit is hit.
Decidedoes this need a tool?Callstructured argumentsObserveresult, or an errorReasonanswer, or call againloop until doneor until the step budget runs outwhat makes this safe to runa hard cap on iterationstools are allow-listed and argument-validatedevery call and result is logged for replaythe failure mode to nameWithout a step cap the model loops: it calls a tool, misreads the error, and calls itagain with the same arguments — indefinitely, and expensively.
The model does not execute anything — it emits a request, and your code decides whether to honour it.

The four failure modes, and their controls

Wrong tool. The model picks search when it should have called the order lookup. Usually a tool-description problem rather than a model problem: descriptions should say when not to use the tool as well as when to. Keep the tool count low — quality degrades noticeably as the list grows, and a dozen tools is already a lot.

Invented arguments. The model calls lookup_order(order_id="A-12345") with an order number it made up because the user did not provide one. This is hallucination applied to function parameters, and it is common. Controls: strict schema validation, required fields with no defaults, and returning a clear error to the model — no order_id was provided by the user — rather than a generic failure, so it can recover by asking.

Loops that never terminate. The model searches, is unsatisfied, searches again with a slightly different query, and repeats. Controls: a hard maximum iteration count (six is a reasonable default), a wall-clock budget, a token budget for the whole conversation, and duplicate-call detection that returns you already called this with these arguments instead of running it again.

Tool errors read as content. A tool returns {"error": "rate limited"} and the model reports to the user that their order is rate limited. Controls: format errors distinctly, and give explicit instructions for how to handle them.

The cost multiplication

This is the part that surprises teams. Each loop iteration is a fresh generation over a longer context, because every observation is appended.

An illustrative six-iteration loop starting from a 2,000-token context, with each observation adding 800 tokens: contexts of 2,000, 2,800, 3,600, 4,400, 5,200, 6,000 — a total of 24,000 prefill tokens against 2,000 for a single-shot answer. Twelve times the prefill cost, plus six decodes instead of one, plus the tool latency itself.

Prefix caching helps a great deal here, because each iteration extends the previous context rather than rewriting it — one more reason for the stable-prefix ordering described above.

Budget explicitly: a per-conversation token cap, a per-conversation tool-call cap, and a degradation path that returns the best available answer when a cap is hit rather than an error.

Side effects and least privilege

The controls above are about correctness. One control is about safety, and it is the one that matters most: a tool that changes the world needs a confirmation step outside the model.

Reading an order is safe to automate. Issuing a refund, sending an email, or deleting a record is not, because as Safety as architecture showed, an attacker who can get text in front of the model can influence what it decides to call. Give every tool the narrowest permissions that work, scope them to the current user, and require an explicit human confirmation for anything with a side effect. This is the strongest available mitigation for prompt injection and the safety lesson in this section returns to it.

Serving at scale

Fifty million conversations a day at twelve turns each is 600 million model calls a day. This is where that becomes a capacity plan and a bill.

Where the 600 million daily calls goRequest arrivesAdmission queueContinuousbatchingStreamfirst tokenTokensstream outStreaming turns a ten-second wait into a 400 ms wait plus reading time.
Time to first token is the metric users feel; total generation time is the metric that costs money.

Streaming changes the latency problem

A 400-token response at 25 tokens per second takes 16 seconds. That is unacceptable as a wait and entirely acceptable as a stream, because the user starts reading immediately.

So the metric is time to first token, targeted at 500 ms p90, and the secondary metric is inter-token latency, which needs to stay above reading speed — roughly 15 tokens per second keeps ahead of most readers. Time to completion barely matters, which is a genuinely unusual property.

Two consequences. Prefill latency is now the user-visible number, so long prompts and large retrieval sets hurt more than their token cost suggests. And output filtering (see Safety, and why it is layered) becomes awkward, because you cannot classify text you have already sent.

Continuous batching

Naive batching groups N requests, runs them together, and returns when the slowest finishes. With generation lengths varying from 20 tokens to 800, most of the batch sits idle waiting.

Continuous batching (also called in-flight batching) works at the step level instead: at every decoding step the server takes whatever requests are currently active, runs one step for all of them, evicts the ones that finished, and admits waiting requests into the freed slots. No request waits for another to complete.

The gain on realistic mixed-length traffic is large — commonly reported as several-fold higher throughput at the same latency compared with static batching. It is the single most important serving technique in this case study, and it is table stakes in modern inference servers rather than something you build.

Capacity planning

Work it through. Fifty million conversations, twelve turns, 250 output tokens per turn is 150 billion output tokens a day, or about 1.74 million tokens per second averaged. Peak is higher; assume a 2.5× peak-to-average ratio, so 4.3 million tokens per second at peak.

At an illustrative 900 output tokens per second per accelerator (batch 32, mid-size model), that is roughly 4,800 accelerators at peak, before redundancy, before regional distribution, and before any headroom. Add 30% and you are near 6,200.

That number is the point of the exercise. It tells you why routing (see Monitoring and follow-ups) and output-length control are not micro-optimisations — a 20% reduction in average response length removes about 1,200 accelerators from the plan.

When demand exceeds capacity

It will. Design the degradation rather than discovering it:

  • Admission control with priority tiers. Paying users, then free users, then background jobs. Shed from the bottom.
  • Queue with a visible position or a wait estimate. A queue the user can see is tolerable; a silent 40-second delay is not.
  • Degrade to a smaller model rather than failing. A slightly worse answer beats an error.
  • Reduce maximum output length under load. Cheap, immediate, and largely invisible.
  • Reject early. If the queue is longer than the timeout, fail fast with a clear message instead of accepting work you cannot finish.

Cost per conversation, computed

Illustrative, at $0.00069 per GPU-second, 12 turns, 250 output tokens per turn, context growing from 2,000 to about 6,400 tokens, prefix caching on:

  • Decode: 12 × 250 = 3,000 tokens ÷ 900 tokens/s = 3.33 GPU-s
  • Prefill with caching: 2,000 initial + 11 × ~450 new tokens = 6,950 tokens ÷ 12,000 tokens/s = 0.58 GPU-s
  • Total: 3.91 GPU-s × $0.00069 = $0.0027 per conversation

Without prefix caching, prefill would be the sum of all twelve full contexts — about 50,000 tokens, or 4.2 GPU-s — taking the total to $0.0052. Caching halves the bill.

At 50 million conversations a day: about $135,000 a day, roughly $49 million a year, for inference alone.