- MantraMindAI
- Courses
- Tech Interview Prep
- Generative AI System Design Interview
Generative AI System Design Interview
For engineers preparing for the generative AI system design interview, whether you come from ML, backend or product engineering. You will be able to design systems like Smart Compose, a ChatGPT-style assistant, RAG, text-to-image and text-to-video end to end, with sound metrics, a cost and latency budget and safety built in; 27 lessons in 11 sections, about 26 hours.
Course Content
Learn the eight-step framework and the generative foundations every case study relies on: autoregressive decoding, transformers, diffusion, evaluation without ground truth, inference cost and safety. 4 lessons, about 200 minutes.
Design Gmail Smart Compose, a short-output completion system where a sub-100 ms budget and the privacy of email decide the model size and the serving design. 2 lessons, about 120 minutes.
Design Google Translate and use it to learn evaluation properly, from BLEU and learned metrics to human judgement, alongside encoder-decoder models and multilingual trade-offs. 2 lessons, about 130 minutes.
Design a ChatGPT-style assistant, covering the training pipeline, context management, tool use, serving at scale with streaming and continuous batching, and layered safety against prompt injection. 3 lessons, about 160 minutes.
Design a retrieval-augmented generation system, from chunking and indexing through retrieval quality, grounded answers with citations, evaluation and per-user access control. 3 lessons, about 150 minutes.
Design an image captioning system, the bridge from text to images, and learn the vision encoders, contrastive image-text pretraining and caption metrics the later sections build on. 2 lessons, about 120 minutes.
Design a realistic face generator with a GAN, covering image metrics such as FID, training stabilisers, latent-space editing and safeguards against deepfakes. 2 lessons, about 120 minutes.
Learn how diffusion and latent diffusion work, and how to make a slow many-step image model fast and cheap enough to serve. 2 lessons, about 140 minutes.
Design a text-to-image system, covering text conditioning and classifier-free guidance, re-captioned training data, prompt-alignment evaluation and prompt and output safety. 3 lessons, about 140 minutes.
Design a personalised headshot product, comparing adaptation techniques such as LoRA and working through the per-user training pipeline, unit economics and consent. 2 lessons, about 120 minutes.
Design a text-to-video system that extends diffusion across time for consistent motion, and bring its data pipeline and compute cost under control. 2 lessons, about 130 minutes.