Course Content
Generative AI System Design Interview
11 sections · 27 lessons
What makes generative design different, and the eight-step framework
A generative system produces an artifact that did not exist before — a sentence, an image, a video clip — rather than picking a label from a fixed list. That single change breaks four assumptions that ordinary machine learning systems are built on, and every design decision in this course follows from those four breaks.
Property 1 — the output is non-deterministic
Ask a classifier the same question twice and you get the same answer twice. Ask a generative model the same question twice and you get two different answers, both of which may be fine.
That is not a bug you can configure away. Sampling is what makes the output interesting rather than repetitive (the lesson on decoding strategies shows why). Even with sampling turned off, batching and floating-point arithmetic on an accelerator make bit-identical output across runs unreliable.
The engineering consequence is immediate. You cannot write assert output == expected. Your test suite has to check properties — does it parse as valid JSON, does it cite a source, is it under 200 words — rather than exact strings.
Property 2 — there is no ground truth to compare against
For a spam classifier, an email is spam or it is not, and a human can say which. For "write a polite decline to this meeting invite", there are thousands of good answers and no list of them. So there is nothing to compute accuracy against.
This is the hardest problem in the course, and a whole part of this section, Evaluating output with no correct answer, is devoted to it. It changes what an evaluation strategy even looks like: not one number, but a combination of a small human-judged set, cheap automatic tripwires, and online behaviour.
Property 3 — inference is expensive enough to shape the architecture
A gradient-boosted-tree classifier answers in a few milliseconds on a shared CPU, costing so little per call that nobody budgets it. A 300-token generation takes seconds on a dedicated accelerator. The gap is roughly three orders of magnitude in both time and money.
At that scale, cost stops being an operational detail and becomes an architectural constraint. It decides model size, whether you can afford a re-ranking pass, whether the feature is synchronous or a background job, and sometimes whether the product is viable at its price at all. The inference cost and latency budget makes that budget concrete.
Property 4 — failures look like confident correct answers
A conventional service fails loudly: a 500, a timeout, a null. A generative model fails by producing fluent, well-formatted, plausible text that is wrong. There is no exception to catch and no confidence score that reliably separates the two.
This is why safety and grounding are architecture (Safety as architecture), not a wrapper you add later.
The framework this course uses
This course uses the seven-step machine-learning design framework from Machine Learning System Design Interview with one step added — safety — and two steps given more weight. If you have not read its lesson on the seven-step framework, read it: this course assumes the framework and does not re-teach it. What follows is how it changes when the output is generated rather than predicted.
The eight steps
- The problem, and framing it as a generative task — clarifying questions, then the explicit input, the explicit output, and what "good" means.
- Metrics — automatic, human, and online, plus where each one lies to you.
- Data — sources, licensing, curation, filtering.
- Model choice and adaptation — which family, and the build / fine-tune / prompt decision (the last part of this lesson).
- Training or adaptation — the objective, and the compute reality.
- Serving — the inference path inside a latency and cost budget.
- Safety and failure modes — hallucination, unsafe output, injection, and the guardrail architecture.
- Monitoring and follow-ups — quality drift, feedback collection, and the deep-dive questions this problem attracts.
The time budget for a 45-minute round
| Step | Minutes | What the interviewer is listening for |
|---|---|---|
| 1 Framing | 5 | Did you ask before designing? Is the output specified precisely? |
| 2 Metrics | 5 | Do you know that one automatic number is never enough? |
| 3 Data | 4 | Sources, licensing, and filtering — not "we'd get some data" |
| 4 Model and adaptation | 6 | The build/fine-tune/prompt decision, with a reason |
| 5 Training or adaptation | 4 | The objective named, and the compute cost admitted |
| 6 Serving | 8 | A latency budget with numbers in it |
| 7 Safety | 6 | An architecture, not a promise to "add guardrails" |
| 8 Monitoring | 4 | How you find out quality dropped before users tell you |
| Buffer | 3 |
Serving and safety are wide on purpose. They are where this round separates candidates.
What this round scores that the ML system design round does not
Four things.
- The adaptation decision. The ML system design round asks which model. This round asks whether you should train anything at all, and most of the time the answer is no.
- Evaluation without ground truth. ML system design metrics have correct answers behind them. Here you have to construct the notion of correct.
- An explicit cost and latency budget. In ML system design it is good practice. Here a design without one is incomplete, because cost decides feasibility.
- Safety as a component. Not a closing sentence — a labelled box in your diagram with a latency cost attached.
The build, fine-tune, or prompt decision
Step 4 of the framework deserves its own treatment now, because it is the most practical decision in this section and the one that most often separates a candidate who has shipped a generative feature from one who has read about them. Interviewers ask it directly — "would you train a model for this?" — and the strong answer is usually no, with a reason.
The four options
1. Pretrain from scratch. Train a new foundation model on a large corpus. Buys you full control over the data, the licence, and the capability profile.
2. Full fine-tuning. Take a pretrained model and continue training all of its weights on your data. Buys the strongest adaptation available short of pretraining.
3. Parameter-efficient fine-tuning. Freeze the pretrained weights and train a small number of new ones — low-rank adapter matrices are the common form. Buys most of the quality of full fine-tuning at a fraction of the cost, and produces a small artifact (megabytes, not gigabytes) that can be swapped per customer. Section 10 (Personalized Headshot Generation) is built on this.
4. Prompting, with retrieval. Change no weights. Write instructions, supply examples, and put the relevant facts into the context at request time. Section 5 (Retrieval-Augmented Generation) is the full treatment.
The comparison
Figures below are illustrative orders of magnitude for a mid-sized model in 2026, not quotes. They will drift; the ratios between rows are the durable part.
| Pretrain | Full fine-tune | Parameter-efficient | Prompt + retrieval | |
|---|---|---|---|---|
| Examples needed | Billions of tokens | 10k–100k | 200–10k | 0–50 |
| One-off compute cost | $1M+ | $500–$8,000 | $20–$400 | $0 |
| Wall-clock to first result | Months | Days | Hours | Minutes |
| Artifact size per variant | Full model | Full model | 10–200 MB | A text file |
| Cost to update | Retrain | Retrain | Retrain the adapter | Edit the prompt or reindex |
| Serving cost | Yours | Yours | Shared base, cheap swap | Higher per request — long prompts |
Read two things off that chart. First, prompting is flat at zero because no training happens — its cost is per request, not up front. Second, pretraining is not plotted, because at roughly a million dollars and up it would compress everything else onto the axis. That gap is the point.
What each option is actually good at
This is where most wrong answers come from. Match the tool to the deficiency:
| The problem | The fix |
|---|---|
| Output format is wrong | Prompt — instructions and two examples |
| The model lacks your facts | Retrieval, not fine-tuning |
| Facts change weekly | Retrieval — nothing else can keep up |
| Tone and house style are wrong | Fine-tune (parameter-efficient first) |
| Domain vocabulary is mishandled | Fine-tune |
| A consistent structured output at high volume | Fine-tune a smaller model — cheaper per request |
| The model must render one specific person | Parameter-efficient fine-tune — Section 10 |
| An unserved language or modality | Pretraining, or continued pretraining |
The honest default
Start at the cheap end and move up only against a measured gap.
- Build the evaluation set first — 100 to 300 realistic items with human judgements.
- Prompt the best model you can call. Measure.
- Add retrieval if the failures are factual. Measure.
- Parameter-efficient fine-tune if the failures are stylistic or structural. Measure.
- Full fine-tune only if step 4 plateaued below the bar and the volume justifies it.
- Pretrain if you are one of the small number of organisations for which that sentence makes sense.
Most teams should stop at step 3 or 4. Saying that out loud, and naming what would push you further, is a strong interview answer.