Course Content
Generative AI System Design Interview
11 sections · 27 lessons
Text-to-video: the model, the compute problem, serving and safety
The first half of this case study fixed the frame — consistency across frames is the requirement — along with the metrics that judge it and a data pipeline that feeds the model latents. This lesson designs the model, prices it, and turns it into a product: an asynchronous service with previews, content filtering and controls against the most convincing deepfakes in the course.
The model
Diffusion extended across time. Three ideas, introduced by showing the obvious approach failing first.
Why frame-by-frame generation does not work
Attempt 1: generate each frame independently from the same prompt. You get 120 unrelated images. Different rooms, different people, different lighting. Not a video.
Attempt 2: use the same random seed for every frame. Better and still unusable. The starting noise is the same, so the images are related — but the prompt gives no reason for frame 47 to differ from frame 46 in any coherent way. You get 120 near-identical images with random small differences, which reads as violent flicker rather than motion.
Attempt 3: condition each frame on the previous one. Now there is a temporal link, and two new problems. Errors accumulate: a small artefact in frame 10 becomes the conditioning for frame 11, and by frame 60 the clip has drifted somewhere unrecognisable. And generation is strictly sequential — 120 dependent steps — so it is slow and cannot be batched across frames.
The problem in all three attempts is the same: consistency is being hoped for rather than enforced.
The answer: generate the clip as one object
Stop thinking of frames. The latent is a volume: frames × height × width × channels. The diffusion process from How diffusion works, slowly runs on that whole volume at once. All 120 frames are denoised together, at every step.
The denoising network gains temporal layers:
- Temporal attention. Each spatial position attends across frames, so the patch of image at position (x, y) in frame 40 can see what was at that position in frames 1 to 120. This is what keeps an object the same object.
- Temporal convolutions across the frame axis, which cheaply enforce local smoothness between neighbouring frames.
A common and economical construction: take a pretrained image diffusion model, insert temporal layers between its existing spatial layers, and train only the new layers on video while keeping the image layers frozen or lightly tuned. The model inherits everything the image model knows about appearance and learns only about time — a large saving, and a direct application of the principle from The build, fine-tune, or prompt decision.
Consistency is now architectural. The model cannot generate frame 40 without reference to frames 39 and 41, because they are denoised in the same pass.
Latent video diffusion
Latent diffusion compressed space. Here you compress space and time: a video autoencoder downsamples spatially by 8× and temporally by 4×, so 120 frames at 1024×576 become 30 latent frames at 128×72.
The saving is the product of both: roughly 64× spatially and 4× temporally, about 256× fewer values than pixel-space video diffusion. Without it, none of this is affordable.
What is lost is the same as before, plus a temporal version: fast motion between adjacent frames is smoothed by temporal compression, so rapid movement can decode with slight blur or ghosting. Run Section 8's round-trip diagnostic on video too — encode and decode real clips with fast motion and look.
The cascade
Even in latent space, generating 30 latent frames at 128×72 directly is expensive and, more importantly, wasteful: the model spends its expensive high-resolution capacity on frames whose temporal structure has not been decided yet.
A widely used architecture is a cascade of separate diffusion models:
- Base model — low resolution, low frame rate. Illustratively 16 frames at 256×144. Cheap, and it decides composition and motion.
- Temporal upsampler — conditioned on the base output, interpolates to 64 frames. It only has to invent the frames in between, which is much easier than inventing motion.
- Spatial upsampler — conditioned on the temporal output, raises resolution to 1024×576.
Each stage is a diffusion model doing a narrower job on cheaper data than a single monolithic model would. The expensive high-resolution work happens last, only once the motion is settled, which is the whole point. Many recent systems instead generate the whole clip with one large diffusion transformer over the video latent, sometimes followed by a separate upsampler; the cascade remains the clearest way to reason about where the compute goes.
What remains unsolved
Be honest, because it is the differentiating answer:
- Long-range consistency. Models are trained on short clips and reliably hold identity for a few seconds. Beyond that, chaining clips together causes visible drift, and there is no complete fix.
- Physical plausibility. The model learns what motion looks like, not what physics is. Contact, weight, and occlusion are frequently wrong, and no amount of temporal attention supplies a physics model.
- Controllable motion. Specifying how something should move — speed, path, timing — through a text prompt is unreliable for the same reasons The known weaknesses gave about spatial relations.
The compute problem
Video is the most expensive thing in this course by a wide margin. This step computes how expensive and then buys it back.
Cost per second of output
Build up from the image figure in the text-to-image serving lesson. Illustrative, at $0.00069 per GPU-second.
A single 1024×576 image at 25 steps with guidance costs about 0.40 GPU-seconds.
A five-second clip at 24 fps is 120 frames. Independently that would be 48 GPU-seconds — and the real figure is higher, because temporal attention couples frames and grows faster than linearly, while the cascade wins some of it back. An illustrative net figure for a five-second clip:
| Cascade stage | GPU-seconds |
|---|---|
| Base: 16 frames at 256×144, 30 steps with guidance | 9 |
| Temporal upsampler: 16 → 64 frames | 14 |
| Spatial upsampler: 64 frames to 1024×576 | 31 |
| Frame interpolation 64 → 120 frames, decode, encode to a video file | 6 |
| Total | 60 GPU-seconds |
That is $0.041 per clip, and 12 GPU-seconds per second of output. Put differently: five seconds of video costs about the same as 150 still images.
At 200,000 clips a day: 12 million GPU-seconds, or about $8,300 a day, roughly $3 million a year — and, at full utilisation, 139 accelerators running continuously.
Cost against clip length
Read the curvature. Cost per second of output rises with length: about 14 GPU-seconds per second of output at one second, 12 at five seconds, 16 at eight, and 28 at thirty-two. Two effects compete — a fixed setup cost amortises over longer clips, which helps; and temporal attention plus the overlap needed to chain segments grows faster than linearly, which eventually wins.
The practical consequence is a product decision: long clips are disproportionately expensive, so price and rate-limit by duration rather than by clip. A user generating one 30-second clip costs fifteen times a user generating one 5-second clip, and a flat per-clip price invites exactly that.
The four architectural responses
1. Lower base resolution, then upsample. Already in the cascade, and the largest single saving. Halving the base resolution quarters the base-stage cost, and the base stage is where motion quality is decided, so push this until motion degrades — not until image quality does.
2. Keyframes plus interpolation. Generate fewer frames with the expensive model and interpolate between them with a cheap dedicated model. Going from 24 generated frames per second to 8 plus interpolation cuts the generation work by about two thirds. It works well for smooth motion and produces visible artefacts on fast or complex motion, so gate it on the motion magnitude you measured in the metrics lesson.
3. Aggressive step reduction. Everything from the diffusion serving lesson applies and is worth considerably more here, because the base is so much larger. A distilled few-step sampler taking 30 steps to 4 is a 7× saving on the most expensive workload in the course. If you can only do one optimisation, do this one.
4. Drop guidance on late steps. Guidance doubles passes. Running it only during the base stage and the early steps of the upsamplers recovers a meaningful fraction at small quality cost.
Applied together, an illustrative five-second clip falls from 60 GPU-seconds to around 12 — from $0.041 to $0.008, and from $3 million a year to under $600,000.
Serving, safety and follow-ups
The final step of the course: everything that has to be true for this to be a product rather than a demonstration.
Asynchronous generation, because there is no alternative
Sixty GPU-seconds of work cannot be a synchronous request. Wall-clock time is a minute or more even with batching, connections drop, retries duplicate expensive work, and a load spike takes the whole service down.
The architecture is a job, not a request:
- Submit. Validate, filter the prompt, and return a job identifier immediately.
- Queue, with a priority tier and a position the user can see.
- Process on a bounded worker pool, checkpointing between cascade stages so a preempted job resumes rather than restarts. As in Section 10, this makes interruptible capacity usable, and it is worth more here because the jobs are longer.
- Progress, reported per cascade stage — "generating motion", "increasing resolution" — which is honest and reads far better than a spinner.
- Notify on completion, and hold the result for a defined window.
Preview frames are the feature that makes the wait acceptable. The base stage finishes first and produces a complete, low-resolution version of the clip. Show it. The user learns within about ten seconds whether the motion and composition are what they wanted, and can cancel before the two expensive upsampling stages run. This is the progressive preview from the text-to-image serving lesson applied to a job that is fifteen times longer, and the saving on cancelled generations is correspondingly larger.
Content filtering, per frame and whole clip
Video needs both, and the distinction is not pedantic:
- Per-frame classification on sampled frames catches content that is visible in a still. Cheap — sample every eighth frame rather than all of them.
- Whole-clip classification using a video model catches what only exists in motion. A sequence of individually innocuous frames can depict something the frames do not, and a per-frame filter will pass every one of them.
Run per-frame filtering on the base-stage output, before the expensive upsamplers. Rejecting a clip after 9 GPU-seconds instead of 60 saves 85% of the compute on a rejected generation, and rejections are not rare.
Deepfake risk, which is worse at video scale
Everything from the safety lessons for face generation and text-to-image generation applies with more force, for a specific reason: video is more convincing than a still. A still can be dismissed; a moving, speaking person is believed. Add synchronised audio and the threshold for belief drops further.
Controls, layered:
- Identity checking per frame and across the clip. Sample frames, embed faces, compare against a public-figure index — and also check consistency, since a clip that is a different person in different frames is a different failure but the same pipeline.
- Refuse photorealistic depictions of named individuals, as in Section 9.
- Watermarking that survives video processing. This is harder than for stills. The watermark must survive re-encoding, resizing, frame-rate conversion, and cropping, and it must be temporally coherent, because a per-frame watermark applied independently can itself flicker. Embed it per frame with temporal coherence, and attach signed provenance metadata to the container.
- Rate limits and account verification for high-volume access, plus retained generation logs.
- A takedown and report path with a real response time.
Say plainly that these bound the harm and do not prevent it, and that provenance built in at generation time is more durable than detection after the fact.
The extensions
- Longer clips. Generate overlapping segments and blend, conditioning each on the tail of the previous. Works to a point and drifts — the unsolved problem from the model design above.
- Image-to-video. Condition on a supplied first frame, which is easier than pure text-to-video because appearance is given and only motion must be generated. It is also the more useful product for many users, and worth proposing.
- Camera control. Condition on a camera trajectory alongside the prompt, so the user specifies a pan or a dolly directly instead of hoping the words produce it. The practical answer to the controllable-motion gap, exactly as layout conditioning was for spatial relations in Section 9.
- Synchronised audio. Either generate audio separately and align, which is simpler and produces loose sync, or generate jointly, which is much harder and is the only route to convincing lip-sync. Naming lip-sync as its own problem is the accurate answer.