Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

Text-to-video: framing, video metrics and the data pipeline


Design a system that turns a written description into a short video clip.

The capstone. Every constraint in the course arrives at once — diffusion, conditioning, alignment, cost, safety — plus one that is new and dominates everything: time.

Why "120 images" is the wrong modelFrames generated separately• Identity changes between frames• Backgrounds flicker and re-shuffle• Motion has no physical continuityOne spatio-temporal sample• Attention runs acrosstime as well as space• Consistency comes from the joint sample• Cost grows with frames times resolution
Temporal consistency is not a post-process — it has to be a property of what the model samples.

Clarifying questions

  • How long a clip? This is the first question, and the answer changes the cost by an order of magnitude. Two seconds, five seconds, and a minute are three different systems. Assume five seconds.
  • Resolution and frame rate? Assume 1024×576 at 24 frames per second, which is 120 frames per clip.
  • Is audio included? Silent video and synchronised audio are separate problems, and lip-sync is a third. Assume silent, with audio as a follow-up.
  • What turnaround? A minute? Ten? This decides synchronous versus asynchronous, and the answer is almost always asynchronous.
  • What control does the user get? Text alone, or also a starting image, camera movement, and a reference for style?
  • What volume? Assume 200,000 clips a day.

The framing

Conditional generation over a temporal sequence, where the dominant requirement is consistency across frames.

The last clause is the whole section. Everything else you already know from Sections 8 and 9.

Why "120 images" is the wrong mental model

The tempting framing is that a five-second clip is 120 images. It is wrong in both directions and being able to say why is a strong opening.

It understates the difficulty. The 120 images must show the same scene, with the same objects, moving plausibly. Generate them independently and you get 120 different scenes. Objects appear and vanish, a person's shirt changes colour every frame, the background is a different room each time. Independence is precisely what must be destroyed.

It understates the cost. Because the frames must be generated jointly, the model attends across time as well as space, and that coupling grows faster than linearly with frame count. A five-second clip costs considerably more than 120 single images, not the same.

And it understates the evaluation problem. A rater can judge an image in two seconds. Judging a clip requires watching it, so evaluation throughput is bounded by real time (see the metrics step below).

The three failure modes that define the problem

Name these early; the rest of the case study is a response to them.

FailureWhat it looks like
FlickerFrame-to-frame instability in texture, lighting, and fine detail. Nothing moves wrongly; everything shimmers.
Identity driftAn object gradually becomes a different object. A face slowly stops being the same face; a red car turns maroon then brown.
Implausible motionEvery frame is individually fine and the sequence does not obey physics. Limbs pass through each other; objects slide rather than step.

The first two are consistency problems and are largely addressed by architecture (see the model lesson). The third is a data and capability problem, and it is the least solved.

Metrics

Video metrics are weaker than image metrics, which are already weak. This step is about being honest about that and building an evaluation that works anyway.

Four axes, all measured weaklyFour axes for videoPer-frame qualityTemporal consistencyMotion realismPrompt alignmentHuman preference
Every automatic video metric is a borrowed image metric, so human raters carry more weight here than anywhere else.

The four things to measure

Per-frame quality. Extract frames and apply the image metrics from the high-resolution metrics lesson. Necessary, insufficient, and it says nothing about motion.

Distributional video quality. The video analogue of Fréchet Inception Distance uses a feature extractor trained on video, so its features encode motion as well as appearance, and it compares the distribution of generated clips against real ones. It carries every FID caveat plus new ones — it is sensitive to the number of frames and the resolution you feed it, so two teams' numbers are not comparable unless the protocol matches exactly.

Temporal consistency. Measurable without references: compute deep features per frame and measure how much they change between consecutive frames, or estimate optical flow between frames, warp one frame into the next, and measure the residual error. Low change means stable.

Prompt alignment. As in the text-to-image metrics lesson, with an extra axis — does the motion described in the prompt occur? Question-based evaluation extends here: "does the person walk left to right?" asked of a vision-language model over sampled frames.

Why human evaluation dominates, and what it costs

Automatic video metrics are weaker than their image counterparts for three reasons: the feature extractors are trained on much smaller and narrower video datasets; the metrics conflate appearance quality with motion quality; and none of them models physical plausibility at all.

So human evaluation is the anchor, and it is expensive in a way image evaluation is not, because a rater has to watch.

Work the arithmetic, because it belongs in your design:

  • 500 clips × 5 seconds = 42 minutes of pure watching, per rater, per pass.
  • Comparative evaluation means watching both clips, often twice: roughly 2.5 hours.
  • With judgement, rest, and quality control: budget 3–4 hours per rater per pass.
  • Three raters per pass for agreement: 9–12 rater-hours per evaluation round.

An image evaluation of the same size is under an hour. That factor of ten changes your process: you run video evaluation less often, so you need better automatic tripwires between rounds, and you need a smaller, sharper evaluation set rather than a bigger one.

What human raters catch that nothing else does

  • Identity drift, which appears gradually over seconds and is invisible frame to frame.
  • Physics violations — a hand passing through a mug, a foot sliding instead of pushing off, a shadow that does not move with its object.
  • Hands and faces in motion, where errors that are tolerable in a still become obvious.
  • Text, which now has to be consistent as well as legible.
  • The "wrong kind of motion" — technically smooth movement that no real object would make.

Data

Video training data is where this section stops being an extension of Section 9 and becomes a distinct engineering problem. The pipeline is a real system before any model is trained.

The ingestion pipeline is a real systemRawlong-form videoCut at shotboundariesDropstatic clipsRe-captionwith a VLMStore as latentsUncut video puts a scene change in the middle of a training clip and teaches the model to cut.
Storing decoded latents rather than video is what makes the training loop affordable to run twice.

Video-text pairs, and why the captions are wrong

Source video with associated text: platform descriptions, subtitles, stock-footage tags, alt-text. All of it describes what the video is about and almost none of it describes what moves.

A stock clip captioned "business meeting office" shows people gesturing, someone standing, a door opening. None of that is in the caption. Train on it and, exactly as in the text-to-image data lesson, the model learns to respond to topic keywords and never learns what "walks towards the camera" looks like.

Re-captioning is therefore mandatory, not optional, and it is harder than for stills. A still caption comes from one model call. A motion caption needs frames sampled across the clip, each described, and then the sequence summarised into a description of what changed — typically a vision-language model over sampled frames followed by a temporal summarisation step. Budget several model calls per clip, and expect the resulting captions to be better at what changed than at how it moved, because the captioner has the same weaknesses the generator will inherit.

Clip segmentation, which is not optional

Raw video contains shot cuts. A training clip that spans a cut teaches the model that scenes change abruptly mid-clip — and it will reproduce that, generating clips that cut to an unrelated scene halfway through.

So the pipeline runs shot-boundary detection and splits on cuts before anything else. Then it filters:

  • Too little motion. Static clips are abundant — locked-off shots, slideshows, talking heads — and training on them produces the static output from the metrics step above. Filter on optical-flow magnitude with a floor.
  • Too much or wrong motion. Camera shake, rapid pans, and jump cuts teach instability.
  • Text overlays, watermarks, logos, and letterboxing, or the model learns to draw them.
  • Aesthetic and resolution filters, as for images.
  • Deduplication, including re-uploads of the same footage, which drives memorisation exactly as in the high-resolution safety lesson.

The storage and preprocessing problem

This is the part that surprises teams, and it is worth being specific.

Consider 100 million five-second training clips at 1024×576, 24 frames per second.

  • As compressed video at roughly 2 Mbps: about 1.25 MB per clip, 125 TB. Manageable.
  • As decoded frames: 120 frames × 1024 × 576 × 3 bytes = 212 MB per clip. Storing that is 21 exabytes. Not an option.

So frames must be decoded on the fly during training — and video decoding is expensive. Decoding 120 frames takes on the order of 100 ms of processor time, so a training batch of 256 clips needs about 25 seconds of decode work per step. Feeding accelerators that want a step every second means a very large processor fleet doing nothing but decoding, and in practice the data loader, not the accelerator, becomes the bottleneck.

The standard fix is elegant and reuses Section 8: encode the clips into latents once, offline, with the video autoencoder, and train on latents.

  • A latent clip at 8× spatial and 4× temporal compression is 30 frames × 128 × 72 × 4 values, or about 2.2 MB in 16-bit.
  • 100 million clips is about 220 TB — larger than the compressed video, far smaller than decoded frames, and it removes decoding from the training loop entirely.
  • At an illustrative $0.02 per GB-month, that is roughly $4,400 a month in storage, plus a one-off encoding pass.

The one-off encoding pass is itself a job: 100 million clips at a few hundred milliseconds each is tens of thousands of accelerator-hours.