Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

What makes generative design different, and the eight-step framework


A generative system produces an artifact that did not exist before — a sentence, an image, a video clip — rather than picking a label from a fixed list. That single change breaks four assumptions that ordinary machine learning systems are built on, and every design decision in this course follows from those four breaks.

The four assumptions that breakGenerative,not predictiveOutput is not a labelNo ground truth existsInference is expensiveFailures look correct
Every later design decision is a response to one of these four properties.

Property 1 — the output is non-deterministic

Ask a classifier the same question twice and you get the same answer twice. Ask a generative model the same question twice and you get two different answers, both of which may be fine.

That is not a bug you can configure away. Sampling is what makes the output interesting rather than repetitive (the lesson on decoding strategies shows why). Even with sampling turned off, batching and floating-point arithmetic on an accelerator make bit-identical output across runs unreliable.

The engineering consequence is immediate. You cannot write assert output == expected. Your test suite has to check properties — does it parse as valid JSON, does it cite a source, is it under 200 words — rather than exact strings.

Property 2 — there is no ground truth to compare against

For a spam classifier, an email is spam or it is not, and a human can say which. For "write a polite decline to this meeting invite", there are thousands of good answers and no list of them. So there is nothing to compute accuracy against.

This is the hardest problem in the course, and a whole part of this section, Evaluating output with no correct answer, is devoted to it. It changes what an evaluation strategy even looks like: not one number, but a combination of a small human-judged set, cheap automatic tripwires, and online behaviour.

Property 3 — inference is expensive enough to shape the architecture

A gradient-boosted-tree classifier answers in a few milliseconds on a shared CPU, costing so little per call that nobody budgets it. A 300-token generation takes seconds on a dedicated accelerator. The gap is roughly three orders of magnitude in both time and money.

At that scale, cost stops being an operational detail and becomes an architectural constraint. It decides model size, whether you can afford a re-ranking pass, whether the feature is synchronous or a background job, and sometimes whether the product is viable at its price at all. The inference cost and latency budget makes that budget concrete.

Property 4 — failures look like confident correct answers

A conventional service fails loudly: a 500, a timeout, a null. A generative model fails by producing fluent, well-formatted, plausible text that is wrong. There is no exception to catch and no confidence score that reliably separates the two.

This is why safety and grounding are architecture (Safety as architecture), not a wrapper you add later.

The framework this course uses

This course uses the seven-step machine-learning design framework from Machine Learning System Design Interview with one step added — safety — and two steps given more weight. If you have not read its lesson on the seven-step framework, read it: this course assumes the framework and does not re-teach it. What follows is how it changes when the output is generated rather than predicted.

The eight steps

  1. The problem, and framing it as a generative task — clarifying questions, then the explicit input, the explicit output, and what "good" means.
  2. Metrics — automatic, human, and online, plus where each one lies to you.
  3. Data — sources, licensing, curation, filtering.
  4. Model choice and adaptation — which family, and the build / fine-tune / prompt decision (the last part of this lesson).
  5. Training or adaptation — the objective, and the compute reality.
  6. Serving — the inference path inside a latency and cost budget.
  7. Safety and failure modes — hallucination, unsafe output, injection, and the guardrail architecture.
  8. Monitoring and follow-ups — quality drift, feedback collection, and the deep-dive questions this problem attracts.
Clarify4mChoose the modelfamily5mand justify itData + adaptation6mTraining or fine-tuning6mInference architecture8mEvaluation8mthe hardest partSafety + cost8m45 minutesgenerative output has no single correct answer, so"how would you measure this" is the question thatseparates candidatesInference, evaluation and safety take more than half the time — because in generative systems those are the parts that are actually hard.
Training gets six minutes; almost nobody in this round is training a foundation model, and the interviewer knows it.

The time budget for a 45-minute round

StepMinutesWhat the interviewer is listening for
1 Framing5Did you ask before designing? Is the output specified precisely?
2 Metrics5Do you know that one automatic number is never enough?
3 Data4Sources, licensing, and filtering — not "we'd get some data"
4 Model and adaptation6The build/fine-tune/prompt decision, with a reason
5 Training or adaptation4The objective named, and the compute cost admitted
6 Serving8A latency budget with numbers in it
7 Safety6An architecture, not a promise to "add guardrails"
8 Monitoring4How you find out quality dropped before users tell you
Buffer3

Serving and safety are wide on purpose. They are where this round separates candidates.

What this round scores that the ML system design round does not

Four things.

  • The adaptation decision. The ML system design round asks which model. This round asks whether you should train anything at all, and most of the time the answer is no.
  • Evaluation without ground truth. ML system design metrics have correct answers behind them. Here you have to construct the notion of correct.
  • An explicit cost and latency budget. In ML system design it is good practice. Here a design without one is incomplete, because cost decides feasibility.
  • Safety as a component. Not a closing sentence — a labelled box in your diagram with a latency cost attached.

The build, fine-tune, or prompt decision

Step 4 of the framework deserves its own treatment now, because it is the most practical decision in this section and the one that most often separates a candidate who has shipped a generative feature from one who has read about them. Interviewers ask it directly — "would you train a model for this?" — and the strong answer is usually no, with a reason.

The four options

1. Pretrain from scratch. Train a new foundation model on a large corpus. Buys you full control over the data, the licence, and the capability profile.

2. Full fine-tuning. Take a pretrained model and continue training all of its weights on your data. Buys the strongest adaptation available short of pretraining.

3. Parameter-efficient fine-tuning. Freeze the pretrained weights and train a small number of new ones — low-rank adapter matrices are the common form. Buys most of the quality of full fine-tuning at a fraction of the cost, and produces a small artifact (megabytes, not gigabytes) that can be swapped per customer. Section 10 (Personalized Headshot Generation) is built on this.

4. Prompting, with retrieval. Change no weights. Write instructions, supply examples, and put the relevant facts into the context at request time. Section 5 (Retrieval-Augmented Generation) is the full treatment.

The comparison

Figures below are illustrative orders of magnitude for a mid-sized model in 2026, not quotes. They will drift; the ratios between rows are the durable part.

PretrainFull fine-tuneParameter-efficientPrompt + retrieval
Examples neededBillions of tokens10k–100k200–10k0–50
One-off compute cost$1M+$500–$8,000$20–$400$0
Wall-clock to first resultMonthsDaysHoursMinutes
Artifact size per variantFull modelFull model10–200 MBA text file
Cost to updateRetrainRetrainRetrain the adapterEdit the prompt or reindex
Serving costYoursYoursShared base, cheap swapHigher per request — long prompts
One-off setup cost against training examples (illustrative)100100010k02001,0005,00020,000100,000log scalePrompt + retrieval (USD)Parameter-efficient fine-tune (USD)Full fine-tune (USD)
One-off setup cost against training examples (illustrative)

Read two things off that chart. First, prompting is flat at zero because no training happens — its cost is per request, not up front. Second, pretraining is not plotted, because at roughly a million dollars and up it would compress everything else onto the axis. That gap is the point.

What each option is actually good at

This is where most wrong answers come from. Match the tool to the deficiency:

The problemThe fix
Output format is wrongPrompt — instructions and two examples
The model lacks your factsRetrieval, not fine-tuning
Facts change weeklyRetrieval — nothing else can keep up
Tone and house style are wrongFine-tune (parameter-efficient first)
Domain vocabulary is mishandledFine-tune
A consistent structured output at high volumeFine-tune a smaller model — cheaper per request
The model must render one specific personParameter-efficient fine-tune — Section 10
An unserved language or modalityPretraining, or continued pretraining

The honest default

Start at the cheap end and move up only against a measured gap.

  1. Build the evaluation set first — 100 to 300 realistic items with human judgements.
  2. Prompt the best model you can call. Measure.
  3. Add retrieval if the failures are factual. Measure.
  4. Parameter-efficient fine-tune if the failures are stylistic or structural. Measure.
  5. Full fine-tune only if step 4 plateaued below the bar and the volume justifies it.
  6. Pretrain if you are one of the small number of organisations for which that sentence makes sense.

Most teams should stop at step 3 or 4. Saying that out loud, and naming what would push you further, is a strong interview answer.