Course Content
Image and Video Generation
4 sections · 7 lessons
Frame Interpolation Basics
You have a two-second clip that came out of a generative video model at 8 frames per second — 16 frames. The client's delivery spec is 24 fps. You need 46 frames, and 30 of them do not exist.
The two obvious fixes both fail visibly. Duplicate each frame three times and the motion becomes a staircase: the subject's hand sits still for 125 ms, jumps, sits still again. Your eye reads it as strobing, and it is unwatchable on anything with real motion. Cross-fade between consecutive frames and you get something worse — the hand appears in both positions at 50% opacity simultaneously. Two transparent hands. This is ghosting, and it is what naive blending always produces, because a 50/50 average of two different positions is not the object halfway between them.
The problem is that both methods operate on pixels at fixed coordinates. A pixel at (100, 50) in frame 0 and a pixel at (100, 50) in frame 1 may be entirely unrelated parts of the scene. To synthesise the frame in between, you must first work out which pixel went where. That correspondence problem is optical flow, and everything else in frame interpolation is built on top of it.
Why anyone needs this
| Application | Typical requirement | Why the naive method fails |
|---|---|---|
| Generative video upsampling | 8 or 12 fps → 24 or 30 fps | duplication reintroduces the judder the model was trying to avoid |
| Slow motion from normal footage | 30 fps → 240 fps (8×) | 7 invented frames per gap; ghosting compounds catastrophically |
| Frame rate conversion for broadcast | 24 fps film → 60 fps display | a 2.5× ratio has no clean duplication pattern; 3:2 pulldown judders |
| Video compression | transmit every third frame, reconstruct the rest | reconstruction must be accurate, not merely plausible |
| Restoring old film | 16 fps archival → 24 fps | the visible speed-up or judder is the whole thing being fixed |
Note what is not on that list: inventing new content. Interpolation fills the gap between two frames you already have. It cannot extend a clip, and it cannot repair something absent from both endpoints. That constraint is also its strength — the answer is bounded on both sides, so a good interpolator is far more reliable than a generative model asked to do the same job.
Optical flow: the correspondence problem
Optical flow assigns every pixel a 2D displacement vector saying where it moved to in the next frame. A 512×512 frame therefore has 262,144 vectors — a dense field, not a handful of tracked points.
The brightness constancy assumption
The whole classical theory rests on one assumption: a physical point keeps the same brightness as it moves.
In plain English: the pixel showing the corner of a table is equally bright in both frames; it has just moved. Take a first-order Taylor expansion of the right-hand side and cancel, and you get the optical flow constraint equation:
where Ix, Iy are the spatial image gradients, It is the change in brightness at that pixel between frames, and (u, v) is the flow you want.
Count the unknowns. Two — u and v — and one equation. The system is underdetermined at every single pixel. This is the aperture problem, and it is not a mathematical technicality; it is a real limit on what is visible. Look at a long vertical edge through a small hole. You can see it moving right. You cannot tell whether it is also moving up, because sliding along its own direction produces no observable change. Only motion perpendicular to the edge is measurable locally.
Every optical flow method is a different answer to the same question: what extra assumption do you add to make an underdetermined problem solvable?
The classical answers: Lucas–Kanade assumes flow is constant across a small window, giving one equation per pixel in that window and a least-squares solution. Horn–Schunck adds a global smoothness penalty. Farnebäck approximates each neighbourhood with a quadratic polynomial and solves for the displacement that maps one polynomial to the other. Learned methods replace the assumption with training data.
Where brightness constancy breaks
| Violation | Real-world cause | Visible result |
|---|---|---|
| Illumination change | a cloud passes, auto-exposure adjusts | large spurious flow across the whole frame |
| Specular highlights | reflection on metal or water moves with the camera, not the object | flow tracks the highlight instead of the surface |
| Transparency | glass, smoke, motion blur | two motions at one pixel; no single vector is correct |
| Textureless regions | a blank wall, clear sky | gradients are zero, so the equation gives no information at all |
| Large displacement | fast motion, low frame rate | Taylor expansion invalid; needs a coarse-to-fine pyramid |
Classical: Farnebäck
1import cv223prev_gray = cv2.cvtColor(frame0, cv2.COLOR_BGR2GRAY)4next_gray = cv2.cvtColor(frame1, cv2.COLOR_BGR2GRAY)56flow = cv2.calcOpticalFlowFarneback(7 prev_gray, next_gray, None,8 pyr_scale=0.5, # each pyramid level is half the previous9 levels=3, # 3 levels handles displacement up to ~8x the window10 winsize=15, # averaging window: bigger = smoother, less accurate11 iterations=3,12 poly_n=5, poly_sigma=1.2, flags=0)1314print(flow.shape) # (H, W, 2) -- dx and dy per pixelThe pyramid parameters are the ones that matter. Farnebäck's polynomial expansion only handles displacements comparable to winsize. A 40-pixel motion with winsize=15 and one level is simply not found — the algorithm reports near-zero flow, confidently. Downscaling by 2 three times turns that 40-pixel motion into 5 pixels at the coarsest level, where it is easily measured, and the estimate is then refined upward. If your interpolation ghosts only on fast-moving objects, insufficient pyramid levels is the first thing to check.
Learned: RAFT
RAFT replaces the hand-designed assumption with a learned one, and its structure is worth understanding because it explains the accuracy gap.
- Feature extraction. Both frames go through a shared CNN, producing per-pixel feature vectors at 1/8 resolution.
- All-pairs correlation. Every feature in frame 0 is dotted with every feature in frame 1, building a 4D correlation volume. For a 512×512 input at 1/8 resolution that is 64×64 × 64×64 = about 16.8 million similarity scores. This is the crucial part: matches at any displacement are computed up front, so large motion is not a special case requiring a pyramid.
- Iterative refinement. A GRU starts from zero flow and updates it 12–32 times, each iteration looking up correlation values around the current estimate.
On the Sintel benchmark, RAFT's published average endpoint error is about 1.6 pixels on the clean pass and under 3 on the harder final pass, several times lower than classical methods like Farnebäck. On fast motion and thin structures the gap is larger still. The cost is a GPU and roughly 100 ms per 512×512 pair versus Farnebäck's ~30 ms on CPU.
1import torch2from torchvision.models.optical_flow import raft_large, Raft_Large_Weights34weights = Raft_Large_Weights.DEFAULT5model = raft_large(weights=weights).eval().to("cuda")6tf = weights.transforms()78img0, img1 = tf(frame0_tensor, frame1_tensor) # scales to [-1, 1]; H and W must be divisible by 89with torch.no_grad():10 flows = model(img0.to("cuda"), img1.to("cuda"))11flow = flows[-1] # the list is every refinement iteration; take the lastTwo traps. RAFT returns a list of intermediate estimates — using flows[0] instead of flows[-1] gives you the barely-refined first guess, which looks like the model is broken. And input height and width must be divisible by 8. The supplied transform only converts and normalises; it does not resize or pad, so resize or pad the frames yourself first (and crop the flow back afterwards), or the model fails on the odd shape.
Warping: turning flow into a frame
With flow in hand, synthesising the midpoint requires one more idea. You want I0.5, but you only have flow from frame 0 to frame 1 — and to fill a pixel in the intermediate frame you need to know where that pixel came from.
Assume motion is locally linear over the gap. Then flow from the intermediate frame back to frame 0 is the negative fraction of the full flow, and forward to frame 1 is the remaining fraction:
In plain English: for each pixel of the frame you are building, step backwards along the flow to find where it was in frame 0, step forwards to find where it will be in frame 1, sample both, and blend them by how far through the gap you are.
Work one pixel. Suppose the flow at a point is (12, −4): the object moves 12 pixels right and 4 up between frames. We are synthesising t = 0.5 and filling the intermediate pixel at (106, 48).
- Ft→0 = −0.5 × (12, −4) = (−6, 2). Sample frame 0 at (106 − 6, 48 + 2) = (100, 50).
- Ft→1 = 0.5 × (12, −4) = (6, −2). Sample frame 1 at (106 + 6, 48 − 2) = (112, 46).
- Blend: 0.5 × I0(100, 50) + 0.5 × I1(112, 46).
Both samples land on the same physical point — (112, 46) is exactly (100, 50) displaced by the flow. That is why the blend is sharp rather than ghosted: you are averaging the same object with itself, not two positions of it.
Real flow values are fractional, so sampling uses bilinear interpolation over the four surrounding pixels. This matters: nearest-neighbour sampling quantises every motion to whole pixels, and sub-pixel motion — a slow pan — then alternates between moving 0 and 1 pixels, producing visible stutter in what should be a smooth glide.
1import torch.nn.functional as F23def warp(img, flow):4 """img: (1,3,H,W), flow: (1,2,H,W). Backward warp via grid_sample."""5 B, _, H, W = img.shape6 ys, xs = torch.meshgrid(torch.arange(H), torch.arange(W), indexing="ij")7 grid = torch.stack((xs, ys), 0).float().to(img.device)[None]8 src = grid + flow # where to read from9 # grid_sample wants coordinates normalised to [-1, 1]10 src_x = 2.0 * src[:, 0] / max(W - 1, 1) - 1.011 src_y = 2.0 * src[:, 1] / max(H - 1, 1) - 1.012 norm = torch.stack((src_x, src_y), dim=-1)13 return F.grid_sample(img, norm, mode="bilinear",14 padding_mode="border", align_corners=True)1516def interpolate(img0, img1, flow01, t=0.5):17 a = warp(img0, -t * flow01)18 b = warp(img1, (1.0 - t) * flow01)19 return (1 - t) * a + t * bNote this is a backward warp — for each output pixel, gather from the input. The alternative, forward warping, pushes each input pixel to its destination, and leaves holes wherever no input pixel happens to land plus collisions wherever several do. Backward warping produces a complete image by construction, which is why every practical system uses it.
Where it goes wrong: occlusion
The blend above assumes both samples are valid. Often one is not.
A person walks in front of a doorway. Part of the doorway is visible in frame 0 and hidden in frame 1 — occlusion. Part is hidden in frame 0 and revealed in frame 1 — disocclusion. In both cases there is no correspondence to find, because the content genuinely does not exist in one of the frames. The flow field will still report a vector there, because the algorithm always reports something, and that vector is meaningless. Blending it in produces a smeared, torn region around every moving object's boundary.
Detecting it: the forward–backward consistency check
Compute flow in both directions. If the flow is trustworthy, following it forward and then back should return you to where you started:
with α = 0.01 and β = 0.5 as standard values. The threshold scales with motion magnitude because fast motion tolerates more absolute error.
Two pixels, worked through. Both have F0→1 = (12, −4), so ||F0→1||² = 144 + 16 = 160.
| Well-matched pixel | Occluded pixel | |
|---|---|---|
| Backward flow at the destination | (−11.6, 3.8) | (2, 1) |
| Round-trip residual | (0.4, −0.2), squared length 0.20 | (14, −3), squared length 205 |
| Threshold | 0.01 × (160 + 149.0) + 0.5 = 3.59 | 0.01 × (160 + 5) + 0.5 = 2.15 |
| Verdict | 0.20 < 3.59 → trust it | 205 > 2.15 → reject |
The occluded pixel's backward flow points nowhere near home, because the surface it came from is hidden in frame 1 and the flow algorithm matched it against whatever happened to be there instead.
Acting on it
Once you have a per-pixel confidence mask, replace the fixed 50/50 blend with a weighted one. Where forward flow is trustworthy and backward flow is not, take frame 0 alone. Where the reverse holds, take frame 1 alone. Where both are trustworthy, blend as normal.
1conf0 = consistency_mask(flow01, flow10) # 1 where flow01 is reliable2conf1 = consistency_mask(flow10, flow01)34w0 = (1 - t) * conf05w1 = t * conf16total = w0 + w1 + 1e-67frame_t = (w0 * warp(img0, -t * flow01) +8 w1 * warp(img1, (1 - t) * flow01)) / totalThat epsilon is not decoration. Where both confidences are zero — a region occluded from both sides, which happens with fast rotation — the denominator is zero and you get NaN propagating through the frame. Handle it explicitly: fall back to the nearer frame, or inpaint the region.
Learned interpolation: RIFE
The flow-then-warp pipeline has an architectural weakness. It estimates flow between the two input frames, then derives the intermediate flow by assuming linear motion. But what you actually need is flow from the intermediate frame — which does not exist — to each input. The linear approximation is exactly wrong for rotation, acceleration and anything following a curve.
RIFE (Real-Time Intermediate Flow Estimation) attacks that directly. Its IFNet estimates Ft→0 and Ft→1 as the primary output, coarse-to-fine, and simultaneously predicts a per-pixel fusion map deciding how much of each warped frame to use — learning the occlusion handling rather than computing it from a consistency rule. A small refinement network then fixes residual artefacts.
| Farnebäck + warp | RAFT + warp | RIFE | |
|---|---|---|---|
| Vimeo90K PSNR | lowest of the three | in between | ~35.6 dB (RIFE paper) |
| Speed, 480p on a mid GPU | ~30 ms (CPU) | ~100 ms | ~30 ms |
| Occlusion handling | manual consistency mask | manual consistency mask | learned fusion map |
| Non-linear motion | poor | poor | handled directly |
| Arbitrary t | yes | yes | yes (recursive or conditioned) |
| Needs training data | no | yes | yes |
The classical pipeline is still worth knowing, and not only for teaching. Flow is an independently useful signal — for stabilisation, for temporal-consistency losses, for masking moving regions — and Farnebäck runs on CPU with no model download.
Measuring quality
The convenient thing about interpolation is that ground truth exists: take a real 60 fps clip, drop every second frame, reconstruct it, and compare against the frame you removed.
PSNR converts mean squared error into decibels:
Because it is logarithmic, small MSE differences move it a lot. An MSE of 10 gives 38.13 dB; an MSE of 20 gives 35.12 dB; an MSE of 50 gives 31.14 dB. Roughly, every doubling of error costs 3 dB. Above ~40 dB differences are usually invisible; below ~30 dB they are obvious.
PSNR's weakness is that it rewards blur. A slightly blurry frame has lower squared error than a sharp frame that is one pixel off — so an interpolator that hedges by averaging scores better than one that commits to a position and is nearly right. This is why PSNR alone systematically favours the ghosting you were trying to eliminate.
| Metric | Measures | Blind spot |
|---|---|---|
| PSNR | pixel-wise error, in dB | rewards blur; ignores structure |
| SSIM | local luminance, contrast and structure agreement | still a per-frame measure |
| LPIPS | distance in a learned perceptual feature space | needs a network; not comparable across backbones |
| Warping error / tOF | temporal stability across the sequence | says nothing about per-frame fidelity |
The last row is the one people forget. Every metric above scores frames in isolation, but interpolation artefacts are fundamentally temporal: a sequence where each frame scores 36 dB but the artefacts jitter between frames looks far worse than one at 34 dB that is stable. Always play the result at full speed. A frame-by-frame inspection will pass footage that flickers unwatchably in motion.
What this means when you build something
Return to the 8 fps clip. The practical sequence is: run RIFE recursively — 8 → 16 → 32 fps by repeated halving, then resample to 24 — rather than asking for arbitrary t in one shot, because recursive halving keeps each interpolation between temporally close frames where motion is smallest and flow is most reliable.
The failures you will actually hit, and what causes them:
| Symptom | Cause | Fix |
|---|---|---|
| Ghosting on fast objects only | displacement exceeds what the flow estimator can find | more pyramid levels, or switch to RAFT/RIFE; or interpolate at lower resolution and upscale the flow |
| Torn halos around moving edges | occlusion regions blended as if valid | forward–backward consistency mask, or a model with a learned fusion map |
| Whole-frame smear at one moment | a hard cut — flow between unrelated shots is meaningless | detect scene cuts and duplicate across them instead of interpolating |
| Subtle stutter on slow pans | nearest-neighbour sampling quantising sub-pixel motion | bilinear sampling in the warp |
| NaN pixels | both confidence weights zero, division by zero | epsilon in the denominator plus an explicit fallback |
| Good stills, bad playback | artefacts are temporal, evaluation was per-frame | watch at speed; measure warping error |
And know the boundary. Interpolation cannot rescue footage where the motion between frames exceeds what correspondence can resolve — a 4 fps clip of a running figure has no reliable correspondence between adjacent frames, and no interpolator will invent one. At that point you need a generative model that synthesises motion rather than one that measures it. Recognising which of the two problems you have, before you spend a day tuning flow parameters, is most of the skill.