Image and Video Generation

Commercial Video Generation — Architecture and Trade-offs


A three-person agency takes a job: a 20-second product teaser, six shots, delivered in four days. They pick a generation platform by reading a leaderboard, take the model at the top, and start rendering. Four days later they have 68 clips and a problem.

Eleven of the 68 are individually beautiful. But the product is a matte-green bottle with a wordmark on the label, and across those eleven clips the green shifts between olive and mint, the wordmark renders as four different unreadable squiggles, and in two of them the bottle grows a second cap during the camera move. The shots are gorgeous and unusable together. The team's actual requirement was never "highest quality clip"; it was the same object, recognisable, across six shots — and no leaderboard measures that.

This is the standard way platform selection goes wrong. The systems on offer differ on axes that a single quality ranking collapses into one number: how long a coherent take can be, whether identity survives across shots, how much of the shot you can actually direct, whether audio is generated with the picture or bolted on afterwards, and what a usable clip costs once you count the failed attempts. Those axes come out of architecture, and the architecture changes far more slowly than the version numbers do.

So this is not a feature list. It is the set of durable questions to ask of any of these systems, the reasons the answers are what they are, and a procedure for testing a platform against your own work in an afternoon rather than trusting anyone's ranking.

What a platform gives you, and what it chargesHow much you can direct• Text only, or text plus a first frame• First and last frame pins the motion• Camera controls on some platforms• Character identity rarely survives a cutWhat it costs to find out• Seconds of video, not tokens, are billed• Generation is async: submit, then poll• Yield matters more than any leaderboard• 68 clips and no usable shot list
Pick on yield per prompt for your shot type, because a leaderboard score says nothing about your sixth attempt.

What the machine is actually doing

Start by separating two jobs that sound similar and are not.

Frame interpolation takes two real frames and invents what is between them. It measures how pixels moved from one to the other, then warps and blends. The answer is bounded on both sides — the model is not deciding what the scene contains, only how it got from A to B. That makes it reliable and cheap, and it is why interpolation is a finishing tool rather than a creative one.

Video generation has no anchors. From a sentence, or from a single still, it must invent every frame and the motion connecting them, keeping objects, materials, lighting and physics consistent the whole way. There is nothing to measure against. Everything is a decision.

Modern systems do this by extending latent image diffusion into time. An image diffusion model compresses a picture into a small latent grid, denoises that grid under text conditioning, and decodes back to pixels. A video model compresses spacetime: a 3D autoencoder squeezes both the frame dimensions and the time dimension, and the denoiser is usually a transformer attending across all of the resulting patches at once — so a patch in frame 40 can see a patch in frame 3. That cross-time attention is what keeps a bottle the same bottle for the length of the take.

The arithmetic that explains everything else

Take an 8-second clip at 24 fps and 1280×720. That is 192 frames, 921,600 pixels each, three channels: 530.8 million values in pixel space.

Compress by 8× in each spatial dimension and 4× in time, into 16 latent channels:

  • Latent frames: 192 / 4 = 48
  • Latent resolution: 1280/8 × 720/8 = 160 × 90
  • Latent values: 48 × 160 × 90 × 16 = 11.06 million — a 48× reduction

Now patch that latent volume into 2×2 tokens for the transformer: 80 × 45 = 3,600 tokens per latent frame, times 48 frames = 172,800 tokens in the sequence. Self-attention is quadratic in sequence length, so one attention layer considers roughly 172,800² ≈ 2.99×10102.99 \times 10^{10} token pairs.

Compare a single 512×512 image under the same scheme: a 64×64 latent, patched to 32×32 = 1,024 tokens, giving 1,024² ≈ 1.05 million pairs. The video is about 28,000× more attention work than the image.

Duration is not a product decision, it is a quadratic cost curve. Doubling clip length roughly quadruples the attention bill, which is why every vendor's coherent-take limit sits in seconds and moves slowly.

That one number explains most of the landscape. It explains why clips are short. It explains why "turbo" tiers exist — they cut resolution, frame count or denoising steps, all of which attack the same quadratic. It explains why longer offerings are usually chained segments rather than one continuous generation, and why chained segments drift. And it explains why practically every system splits attention (spatial within a frame, temporal across frames, or windowed) rather than doing true full 3D attention: full attention at this sequence length does not fit.

The five axes that actually differ

When you compare systems, compare these. Everything else is marketing.

AxisWhat to askArchitectural cause
Coherent durationHow long a single generation, not how long a stitched output?Quadratic attention over spacetime tokens; the training clip length caps what the model has ever learned to keep consistent
Within-shot consistencyDoes an object hold its shape, colour and count through a camera move?How much temporal attention range the model has, and whether the latent autoencoder preserves fine detail through 4× time compression
Cross-shot identityCan the same character or product appear in six separate generations?Requires an explicit conditioning path — a reference image or identity embedding. Text alone cannot specify a face
Control surfaceWhat can you pin down besides words — first frame, last frame, camera path, a driving video?Each control is a separate conditioning channel that had to be trained in; it is not a UI feature
Native audioIs sound generated jointly with the picture, or added afterwards?Joint generation means audio and video latents are denoised together, so lip movement and footfalls land on the right frame. Post-hoc audio cannot do that

Cost and latency ride on top of all five. They are not an independent axis so much as the price of where a system sits on the others.

Named systems — as of September 2026, Google's Veo, Kling, Runway's Gen series, ByteDance's Seedance, and open-weight models such as Wan and LTX — are best understood as different bets across this table rather than as points on a quality line. One family invests in directability and cross-shot consistency because its users are making sequences. Another invests in joint audio because its users are making self-contained social clips. Another invests in speed and price because its users iterate dozens of times per idea. A ranking that says one is "best" has silently chosen which axis matters.

Any table of who currently leads is out of date within months. OpenAI's Sora, one of the best-known names in 2025, was shut down in 2026, with its API closing in September. The axes are stable; the answers are not. Test the axes yourself, against your own shots.

Conditioning: how much of the shot you can actually direct

Text is a weak controller. "A matte-green bottle with a wordmark" specifies a region of possibility space containing millions of bottles, and the model picks one per seed. Every additional conditioning channel narrows that region using something more precise than words.

ModeWhat it fixesWhat it leaves freeUse it when
Text-to-videoSubject, style, rough actionEverything visual — identity, composition, paletteExploring ideas; b-roll where nothing must match
Image-to-video (first frame)Composition, colour, identity, lighting at frame 0All motion, and how far identity drifts by the last frameYou need a specific look. This is the highest-leverage control available
First + last keyframeBoth endpoints, so the shot must land somewhere specificThe path between themA defined beat: closed door to open door, day to night
Reference identityA face, a character or a product across separate generationsPose, framing, contextAny multi-shot sequence with a recurring subject
Camera / motion controlTrajectory, speed, shot typeScene contentThe shot must cut with adjacent footage
Video-to-videoFull motion and timing, from a driving clipAppearance and styleYou have blocking — even a phone reference — and want it restyled

The practical consequence for the agency in the opening story: they were asking text-to-video to hold identity across six independent generations, which it structurally cannot do. The fix is to move the identity decision out of the video model. Generate one still of the bottle with an image model, where you can iterate cheaply and use precise spatial control, approve it, then drive every one of the six shots as image-to-video from that same approved still. Identity is now fixed at frame 0 of each shot by construction, and the only remaining question is how far it drifts within a four-second take — a far smaller problem than inventing the bottle six times.

Why chaining clips is harder than it looks

Once you accept short takes, you will chain them, and chaining has two distinct failure modes people conflate.

Seams. A hard cut between two independent generations is often fine — cuts are normal film grammar. It is the continuous join that hurts: extending a shot by feeding the last frame of clip A as the first frame of clip B. Frame-to-frame that looks continuous, but B's motion starts from rest, so you get a visible hitch at every joint, roughly once per four seconds.

Drift. Chained generation is autoregressive: each segment conditions on the decoded output of the previous one, so every artefact becomes input. Colour creeps, faces soften, backgrounds simplify. This compounds — the fifth segment is conditioned on four generations of accumulated error.

The compounding is worth quantifying, because intuition badly underestimates it. Suppose a given platform holds product identity acceptably in 82% of four-second shots. For a single shot that is a good number. For a six-shot sequence where all shots must match:

P(all six clean)=0.826=0.304P(\text{all six clean}) = 0.82^{6} = 0.304

So there is a 69.6% chance of at least one broken shot per attempt at the full sequence. At 82% per shot, you are re-rolling most sequences. Push per-shot reliability to 95% and the same calculation gives 0.956=0.7350.95^{6} = 0.735, so 26.5% failure — still not comfortable, but a different working experience. This is why cross-shot conditioning is worth more than raw per-clip quality on any multi-shot job: it moves the base of an exponent.

The API shape, and the mistake to avoid around it

Generation takes tens of seconds to several minutes, which does not fit inside an HTTP request. So essentially every commercial system uses the same asynchronous shape: submit a job, get an identifier immediately, then poll or receive a webhook until the job reports success or failure.

Text
client                                   generation service  |-- POST /generate {prompt, image, ...} --->|  (validated, queued)  |<-- 202 { job_id: "j_8812" } -------------|  |                                           |  |-- GET /jobs/j_8812 --------------------->|  |<-- { status: "queued" } -----------------|  |-- GET /jobs/j_8812  (after backoff) ----->|  |<-- { status: "running", progress: 0.4 } -|  |-- GET /jobs/j_8812 --------------------->|  |<-- { status: "succeeded", url: "..." } --|
Python
import time, random, requestsdef wait_for(session, base, job_id, timeout=1800):    """Poll a generation job with exponential backoff and jitter.    Field names and paths below are placeholders. Read the vendor's    current reference and substitute; the *shape* is what transfers.    """    delay, waited = 2.0, 0.0    while waited < timeout:        r = session.get(f"{base}/jobs/{job_id}")        r.raise_for_status()        body = r.json()        state = body["status"]        if state == "succeeded":            return body["url"]        if state in ("failed", "cancelled"):            raise RuntimeError(f"{state}: {body.get('error')}")        time.sleep(delay + random.uniform(0, 0.5))   # jitter        waited += delay        delay = min(delay * 1.6, 20.0)               # back off, cap at 20s    raise TimeoutError(f"job {job_id} exceeded {timeout}s")

Four details separate a toy client from one that survives a render night.

  • Backoff with jitter. A fixed one-second poll on 40 parallel jobs is 40 requests per second of pure overhead and will earn you rate limiting. Jitter stops your own jobs from synchronising into bursts.
  • Idempotency on submit. If the submit request times out you do not know whether the job was created. Retrying without an idempotency key bills you twice for identical renders, and at video prices that is the expensive mistake.
  • Webhooks where offered. Polling is the fallback. A webhook removes the latency-versus-request-count trade-off entirely.
  • Persist the full request with the asset. Prompt, seed, model version, every parameter, and the job id, stored next to the output file. Without it you cannot reproduce or explain a clip a month later.

Never write integration code against an API you have not read. Plausible-looking client code for these services circulates widely, including method names and packages that were guessed before a product shipped and never existed. Copy field names from the vendor's current reference, not from any tutorial.

Prompting for motion

Still-image prompting describes a frame. Video prompting has to describe a frame and what changes, and the second part is where most weak prompts fail. A useful discipline is to write five separable slots rather than one run-on sentence, because they map onto genuinely different decisions the model makes.

SlotWeakStrong
Setting"in a park""a sunlit city park, scattered autumn leaves, low morning sun behind the trees"
Action"a dog""a golden retriever sprints left to right across frame, ears flapping, kicking up leaves"
Camera— (omitted)"low tracking shot moving alongside the dog at knee height, steady"
Mood"cool video""joyful, energetic"
Style— (omitted)"natural light, shallow depth of field, slight motion blur on the legs"

Omitting the camera slot does not give you a neutral camera; it gives you the model's prior, which is usually a slow generic drift. If your clip must cut against footage shot on a locked-off tripod, you have to say so.

Two failure modes are worth naming because they look like model weakness and are actually prompt weakness:

  • The multi-beat prompt. "She opens the door, walks to the desk, sits down and starts typing" packed into one four-second clip. The model has neither the frames nor the training distribution for a three-beat sequence at that length, and it typically averages the beats into an indistinct wander. One clear action per clip; separate generations for separate beats.
  • Conflicting physics. "Handheld documentary look" plus "perfectly smooth glide" plus "120 fps slow motion" pull in different directions. The model resolves the conflict by picking one and ignoring the others, and which one it picks is seed-dependent, which reads as instability.

Iterate on a fast, cheap tier to settle wording and composition, then re-render the winner on the expensive tier with the same seed where the platform supports it. Treating model choice as a cost/quality dial rather than a fixed decision is the single biggest saving available.

Evaluating a platform in an afternoon

Here is a protocol that beats any leaderboard for your specific need, because it measures your shots.

  1. Pick five prompts representative of your real work — not showcase prompts. Include the awkward one you know is hard.
  2. Render three seeds each per candidate platform. Fifteen clips per vendor.
  3. Score every clip against a fixed rubric, blind to which vendor produced it if you can manage it.
  4. Compute cost per usable clip, not cost per clip.
CriterionWhat you are looking forCommon tell
Prompt adherenceIs the requested action and camera move actually present?Camera direction silently ignored
Temporal identitySame object, same count, same colour from first frame to lastExtra fingers appearing mid-move; a logo re-drawing itself
Physical plausibilityLiquids, cloth, contact, weightFeet sliding; poured liquid that does not accumulate
Small-detail stabilityText, hands, faces, patternsLegible text in frame 1 becoming glyph soup by frame 60
Seam behaviourWhat happens where two generations joinMotion hitch at every four-second boundary
Audio sync (if relevant)Do sounds land on the frame that causes them?Footfall audio drifting a few frames late

Now the arithmetic that flips rankings. Suppose vendor A yields 9 usable clips out of 15 at 1.20 dollars per render, and vendor B yields 6 of 15 at 0.40 dollars per render.

  • Vendor A: 15 × 1.20 = 18.00 dollars for 9 usable → 2.00 dollars per usable clip
  • Vendor B: 15 × 0.40 = 6.00 dollars for 6 usable → 1.00 dollar per usable clip

Vendor A wins on quality by 50% and loses on economics by 2×. Which one is right depends entirely on whether the job is 400 social clips or one hero shot, and on whether a human has to watch all 15 either way — because at scale review time, not render cost, is usually the binding constraint. Add reviewer minutes to the calculation and low-yield platforms get worse than the money alone suggests.

A cheap automated pre-screen

You cannot watch everything at scale, but you can cut the review pile. Sudden discontinuities show up as spikes in frame-to-frame optical flow — the same measurement used for interpolation, repurposed as a smoke detector.

Python
import cv2, numpy as npdef flag_discontinuities(path, z=3.0):    """Return frame indices where motion magnitude spikes far above    the clip's own baseline -- candidates for a human look."""    cap = cv2.VideoCapture(path)    prev, mags = None, []    while True:        ok, frame = cap.read()        if not ok:            break        gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)        if prev is not None:            flow = cv2.calcOpticalFlowFarneback(                prev, gray, None, 0.5, 3, 15, 3, 5, 1.2, 0)            mags.append(float(np.mean(np.linalg.norm(flow, axis=2))))        prev = gray    cap.release()    if len(mags) < 8:        return []    m = np.array(mags)    thresh = m.mean() + z * m.std()      # relative to this clip, not absolute    return [i for i, v in enumerate(m) if v > thresh]

The threshold is deliberately relative to the clip's own statistics. A fixed absolute threshold flags every fast-action clip and misses every discontinuity in a slow one. This will not judge beauty — it catches the jumps, teleports and hard content swaps that make a clip obviously broken, which is exactly the tier of failure you want filtered before a person spends time on it.

What this means when you build something

Assume the vendor will change. Put a shot-level interface between your pipeline and any provider: a function that takes a prompt, an optional reference image, a duration and a seed, and returns a file path. Everything vendor-specific — endpoints, field names, polling, retries — lives behind it. When the landscape shifts, and it will, you swap an implementation rather than rewriting a pipeline.

Structure the work so the expensive model does the least possible. Decide identity and composition with cheap, controllable image generation. Use the video model for motion only, driven from approved stills. Fix frame rate and resolution afterwards with interpolation and upscaling, which are deterministic and cost a fraction of a re-render. A four-second shot generated at 12 fps and interpolated to 24 is often indistinguishable from one generated at 24 and costs meaningfully less to produce.

Budget in usable clips. A plan built on "we need six shots" is a plan that ships late; a plan built on "we need six shots at an observed 60% yield, so budget 15 renders and two review passes" is a plan. Record prompt, seed, model version and job id alongside every asset, because the same prompt on a new model version is a different clip — and a client asking for "the same thing but three seconds longer" six weeks later is the moment you discover whether your archive is reproducible.

Finally, re-run the fifteen-clip test whenever a version ships or a renewal comes up. It costs an afternoon and a few tens of dollars, and it is the only comparison that is measured on your work rather than someone else's.