Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

ChatGPT assistant: framing, metrics and the training pipeline


Design a conversational assistant: a user types a message, the system replies, and the conversation continues across many turns. It should be able to answer questions, help with tasks, and use tools where needed.

This is the most-asked question in the generative design round, and the one candidates most often answer badly — not because the components are hard, but because they answer a different question than the one asked.

Building the model or consuming oneBuilding your own• Pretraining plusalignment, months of work• Full control over safety and behaviour• Capital cost measured in millionsConsuming an API• The interestingdesign starts at the prompt• Safety and quality are partly outsourced• Cost is per token, forever, and visible
Almost every real answer is the second, so the design question is what you build around the model.

Clarifying questions

  • General-purpose or domain-specific? A customer-support assistant over your own documentation is a fundamentally different system from a general assistant, and the answer changes the model choice, the safety surface, and the evaluation strategy.
  • Multi-turn? Almost always yes, and that means context management (Managing context) is a first-class component rather than a detail.
  • Tools and API access? Can it search, calculate, look up an order, send an email? Each capability added multiplies the safety surface, and the ones with side effects multiply it most.
  • What are the safety requirements? Consumer product, enterprise internal tool, and regulated-industry deployment have very different bars.
  • Are we building the model or consuming one? The single most important question in this case study, and the one candidates forget to ask.
  • Scale. Assume 50 million conversations a day, averaging 12 turns.
  • Latency expectation? First token within about 500 ms; full response can take several seconds because it streams.

The framing that decides the answer

A system around a generative model, not the model itself.

Draw the boundary explicitly. Almost everything you will be scored on lives outside the model:

Text
 user ──► input safety ──► context assembly ──► retrieval ──► [ MODEL ] ──► tool loop                                    ▲                             │                                    └──── conversation store ◄────┘                                                                  ▼                                              output safety ──► streaming ──► user

Six boxes; one of them is the model. A candidate who spends thirty minutes on transformer architecture has designed one box and left five empty.

Building versus consuming, and why it changes everything

If you are building the modelIf you are consuming one
Where effort goesPretraining data, alignment, evaluation, serving infrastructureContext, retrieval, tools, safety, routing, cost
Cost shapeEnormous fixed cost, low marginalZero fixed, meaningful marginal — per token, forever
Quality leverTraining data and alignmentPrompt, retrieval, adaptation, model selection
Realistic forA handful of organisationsNearly everyone
Biggest riskAlignment and capabilityVendor dependency, cost at scale, model updates changing behaviour

This section covers the training pipeline (at the end of this lesson) because you will be asked about it, and because understanding it is what lets you reason about what a consumed model can and cannot be made to do. But say clearly which case you are designing for. An interviewer asking this question is usually testing whether you know the difference.

Metrics

A chatbot has no correct output, no reference answer, and a user population that will not tell you when it was wrong. Everything from Evaluating output with no correct answer applies here at its hardest.

Three levels that disagree with each otherTurn: was this reply goodSession: was the task doneUser: did they come back
A reply can be rated excellent, in a session that failed, for a user who never returns.

Three levels, because they disagree

Turn level — was this response good? Two axes that trade against each other:

  • Helpfulness. Did it answer the question, at the right length, in a usable form?
  • Harmlessness. Did it avoid harmful, false, or inappropriate content, including by refusing when refusal was right?

They pull in opposite directions, which is exactly why they are measured separately. A system that refuses everything is perfectly harmless and useless. Track over-refusal rate — the share of benign requests declined — as an explicit metric, because it is the cost side of the safety ledger and teams routinely fail to measure it.

Conversation level — did the user get what they came for? Task completion rate, turns to resolution, regeneration rate (how often the user asks for another attempt), and abandonment. These are the metrics that matter most and the hardest to label, since determining whether a task completed usually requires reading the conversation.

Product level — retention. Weekly return rate and long-run engagement. Slow, confounded, and the only thing that eventually matters.

The primary instrument: pairwise human preference

Show a rater the same conversation with two candidate responses and ask which is better. Aggregate into a win rate.

Why pairwise rather than absolute scoring: humans are far more consistent at comparison than at calibration. Asked to rate responses 1 to 5 they anchor differently, drift over a session, and cluster on 3 and 4. Asked which of two is better they agree with each other much more often.

Practicalities: a written rubric, a calibration set every rater completes first, both orderings shown to cancel position bias, and a tie option. Report a confidence interval — a 54% win rate over 200 comparisons is not a result.

Model-as-judge, for scale

Human comparison at a few dollars per item cannot run on every pull request. A model-as-judge scores thousands of items for a few dollars total, so it becomes the gate.

The biases from the evaluation lesson all apply, and one deserves repeating here because chatbot evaluation is where it does the most damage:

Online metrics, and the trap in the obvious one

Available signals: thumbs up and down (very sparse, typically a fraction of a percent of turns, and skewed towards annoyed users), regeneration rate, copy-to-clipboard, conversation continuation, and explicit reports.

The trap is session length. It is easy to measure, it looks like engagement, and it is ambiguous in the worst possible way: a long session means the assistant is either very useful or failing repeatedly. Optimising it directly rewards an assistant that takes six turns to do what should take two. Pair it with turns-to-resolution and treat a rise in both together as a regression, not a win.

The training pipeline, at the level this round needs

You need this to answer "how does a model learn to be a helpful assistant?" and, more usefully, to reason about what a model can be made to do without retraining it. Three stages, each contributing something the previous one could not.

Stage 1 — pretraining

Next-token prediction (see Autoregressive generation, explained plainly) over an enormous, broad corpus. Months of compute on thousands of accelerators.

What it contributes: language, world knowledge, reasoning patterns, code, and the ability to continue text coherently. Essentially all of the model's capability arrives here.

What it does not contribute: any notion of being helpful. A purely pretrained model asked "How do I reset my password?" is as likely to continue with three more questions in the same style — because that is what documents containing that sentence look like — as to answer it. The knowledge is present; the behaviour is not.

Stage 2 — supervised fine-tuning on demonstrations

Continue training on a curated set of (instruction, ideal response) pairs written or edited by people. Typically thousands to tens of thousands of examples. Hours to days of compute.

What it contributes: the format of assistance. The model learns that a question gets an answer, that an instruction gets a completion, that a response ends, and roughly what a good answer looks like structurally.

What it cannot contribute: ranking. There are many acceptable answers to most requests and demonstration data shows one of them. It cannot express that answer A is better than answer B when both are fine.

Stage 3 — preference-based alignment

Collect comparisons: for the same prompt, show people two model responses and record which they prefer. Use those comparisons to adjust the model towards the preferred behaviour — either by training a separate reward model that scores responses and optimising against it, or by optimising the model directly on the preference pairs, which several modern methods do.

What it contributes: the difference between an acceptable answer and a good one. Tone, length, hedging where appropriate, structure, and — importantly — refusal behaviour, since "refuses politely and explains" is learned as a preferred response to certain prompts.

Stated without mathematics, the objective rewards one thing: producing responses that people choose over the alternatives, while not drifting far from what the fine-tuned model would have said. That second half exists because unconstrained optimisation against a preference signal degenerates — the model finds phrasings that score well and read badly.

1. Pretrainingdatatrillions of tokens of web textteacheslanguage, facts, reasoning patternscostmonths, thousands of GPUsa model that continues text but does not follow instructions2. Supervised fine-tuningdatatens of thousands of written demonstrationsteachesthe shape of a helpful answercostdays, tens of GPUsa model that answers, but cannot tell a good answer from amediocre one3. Preference tuningdatahuman rankings of pairs of answersteacheswhich of two acceptable answers people prefercostdays, plus continuous human labellinga model aligned to a preference distribution, with its biasesAlmost all the capability comes from stage 1 and almost all the usability from stages 2 and 3 — which is why a small, well-tuned model often beats a larger untuned one on a realproduct.
Capability and helpfulness come from different stages — conflating them is the mistake behind "just use a bigger model".

The honest note: you are probably consuming, not training

For the overwhelming majority of teams, all three stages are somebody else's. That changes the answer, and saying so is a strength rather than a dodge:

DeficiencyAvailable lever when consuming
Wrong format or toneSystem prompt; few-shot examples
Missing factsRetrieval (Section 5)
Domain style and vocabularyParameter-efficient fine-tuning (see The build, fine-tune, or prompt decision)
Refuses too much or too littlePrompt and post-filtering; limited
Fundamental capability gapChange models — nothing else will fix it

The last row is the honest one. If the model cannot do the reasoning your task requires, no amount of prompting fixes it, and recognising that boundary quickly is a practitioner skill.