Generative AI System Design Interview

Course Overview
Advanced
Free Course

For engineers preparing for the generative AI system design interview, whether you come from ML, backend or product engineering. You will be able to design systems like Smart Compose, a ChatGPT-style assistant, RAG, text-to-image and text-to-video end to end, with sound metrics, a cost and latency budget and safety built in; 27 lessons in 11 sections, about 26 hours.

Instructor: MantraMindAI
Sections: 11

Course Content

Section 1: Introduction and Generative Foundations

Learn the eight-step framework and the generative foundations every case study relies on: autoregressive decoding, transformers, diffusion, evaluation without ground truth, inference cost and safety. 4 lessons, about 200 minutes.

Section 2: Gmail Smart Compose

Design Gmail Smart Compose, a short-output completion system where a sub-100 ms budget and the privacy of email decide the model size and the serving design. 2 lessons, about 120 minutes.

Section 3: Google Translate

Design Google Translate and use it to learn evaluation properly, from BLEU and learned metrics to human judgement, alongside encoder-decoder models and multilingual trade-offs. 2 lessons, about 130 minutes.

Section 4: ChatGPT: Personal Assistant Chatbot

Design a ChatGPT-style assistant, covering the training pipeline, context management, tool use, serving at scale with streaming and continuous batching, and layered safety against prompt injection. 3 lessons, about 160 minutes.

Section 5: Retrieval-Augmented Generation

Design a retrieval-augmented generation system, from chunking and indexing through retrieval quality, grounded answers with citations, evaluation and per-user access control. 3 lessons, about 150 minutes.

Section 6: Image Captioning

Design an image captioning system, the bridge from text to images, and learn the vision encoders, contrastive image-text pretraining and caption metrics the later sections build on. 2 lessons, about 120 minutes.

Section 7: Realistic Face Generation

Design a realistic face generator with a GAN, covering image metrics such as FID, training stabilisers, latent-space editing and safeguards against deepfakes. 2 lessons, about 120 minutes.

Section 8: High-Resolution Image Synthesis

Learn how diffusion and latent diffusion work, and how to make a slow many-step image model fast and cheap enough to serve. 2 lessons, about 140 minutes.

Section 9: Text-to-Image Generation

Design a text-to-image system, covering text conditioning and classifier-free guidance, re-captioned training data, prompt-alignment evaluation and prompt and output safety. 3 lessons, about 140 minutes.

Section 10: Personalized Headshot Generation

Design a personalised headshot product, comparing adaptation techniques such as LoRA and working through the per-user training pipeline, unit economics and consent. 2 lessons, about 120 minutes.

Section 11: Text-to-Video Generation

Design a text-to-video system that extends diffusion across time for consistent motion, and bring its data pipeline and compute cost under control. 2 lessons, about 130 minutes.