Image and Video Generation

Frame Interpolation Basics


You have a two-second clip that came out of a generative video model at 8 frames per second — 16 frames. The client's delivery spec is 24 fps. You need 46 frames, and 30 of them do not exist.

The two obvious fixes both fail visibly. Duplicate each frame three times and the motion becomes a staircase: the subject's hand sits still for 125 ms, jumps, sits still again. Your eye reads it as strobing, and it is unwatchable on anything with real motion. Cross-fade between consecutive frames and you get something worse — the hand appears in both positions at 50% opacity simultaneously. Two transparent hands. This is ghosting, and it is what naive blending always produces, because a 50/50 average of two different positions is not the object halfway between them.

The problem is that both methods operate on pixels at fixed coordinates. A pixel at (100, 50) in frame 0 and a pixel at (100, 50) in frame 1 may be entirely unrelated parts of the scene. To synthesise the frame in between, you must first work out which pixel went where. That correspondence problem is optical flow, and everything else in frame interpolation is built on top of it.

16 real frames, 30 that have to be inventedf0newnewf1newnewf2new012345678 fps sourcewarp from bothnextreal frameEach synthetic frame warps f0 forward and f1 backward, then blends where the two agree.
The forward-backward consistency check is what flags occluded pixels, where no correspondence exists to warp from.

Why anyone needs this

ApplicationTypical requirementWhy the naive method fails
Generative video upsampling8 or 12 fps → 24 or 30 fpsduplication reintroduces the judder the model was trying to avoid
Slow motion from normal footage30 fps → 240 fps (8×)7 invented frames per gap; ghosting compounds catastrophically
Frame rate conversion for broadcast24 fps film → 60 fps displaya 2.5× ratio has no clean duplication pattern; 3:2 pulldown judders
Video compressiontransmit every third frame, reconstruct the restreconstruction must be accurate, not merely plausible
Restoring old film16 fps archival → 24 fpsthe visible speed-up or judder is the whole thing being fixed

Note what is not on that list: inventing new content. Interpolation fills the gap between two frames you already have. It cannot extend a clip, and it cannot repair something absent from both endpoints. That constraint is also its strength — the answer is bounded on both sides, so a good interpolator is far more reliable than a generative model asked to do the same job.

Optical flow: the correspondence problem

Optical flow assigns every pixel a 2D displacement vector saying where it moved to in the next frame. A 512×512 frame therefore has 262,144 vectors — a dense field, not a handful of tracked points.

The brightness constancy assumption

The whole classical theory rests on one assumption: a physical point keeps the same brightness as it moves.

I(x,y,t)=I(x+Δx, y+Δy, t+Δt)I(x, y, t) = I(x + \Delta x,\ y + \Delta y,\ t + \Delta t)

In plain English: the pixel showing the corner of a table is equally bright in both frames; it has just moved. Take a first-order Taylor expansion of the right-hand side and cancel, and you get the optical flow constraint equation:

Ixu+Iyv+It=0I_x u + I_y v + I_t = 0

where Ix, Iy are the spatial image gradients, It is the change in brightness at that pixel between frames, and (u, v) is the flow you want.

Count the unknowns. Two — u and v — and one equation. The system is underdetermined at every single pixel. This is the aperture problem, and it is not a mathematical technicality; it is a real limit on what is visible. Look at a long vertical edge through a small hole. You can see it moving right. You cannot tell whether it is also moving up, because sliding along its own direction produces no observable change. Only motion perpendicular to the edge is measurable locally.

Every optical flow method is a different answer to the same question: what extra assumption do you add to make an underdetermined problem solvable?

The classical answers: Lucas–Kanade assumes flow is constant across a small window, giving one equation per pixel in that window and a least-squares solution. Horn–Schunck adds a global smoothness penalty. Farnebäck approximates each neighbourhood with a quadratic polynomial and solves for the displacement that maps one polynomial to the other. Learned methods replace the assumption with training data.

Where brightness constancy breaks

ViolationReal-world causeVisible result
Illumination changea cloud passes, auto-exposure adjustslarge spurious flow across the whole frame
Specular highlightsreflection on metal or water moves with the camera, not the objectflow tracks the highlight instead of the surface
Transparencyglass, smoke, motion blurtwo motions at one pixel; no single vector is correct
Textureless regionsa blank wall, clear skygradients are zero, so the equation gives no information at all
Large displacementfast motion, low frame rateTaylor expansion invalid; needs a coarse-to-fine pyramid

Classical: Farnebäck

Python
import cv2prev_gray = cv2.cvtColor(frame0, cv2.COLOR_BGR2GRAY)next_gray = cv2.cvtColor(frame1, cv2.COLOR_BGR2GRAY)flow = cv2.calcOpticalFlowFarneback(    prev_gray, next_gray, None,    pyr_scale=0.5,   # each pyramid level is half the previous    levels=3,        # 3 levels handles displacement up to ~8x the window    winsize=15,      # averaging window: bigger = smoother, less accurate    iterations=3,    poly_n=5, poly_sigma=1.2, flags=0)print(flow.shape)          # (H, W, 2) -- dx and dy per pixel

The pyramid parameters are the ones that matter. Farnebäck's polynomial expansion only handles displacements comparable to winsize. A 40-pixel motion with winsize=15 and one level is simply not found — the algorithm reports near-zero flow, confidently. Downscaling by 2 three times turns that 40-pixel motion into 5 pixels at the coarsest level, where it is easily measured, and the estimate is then refined upward. If your interpolation ghosts only on fast-moving objects, insufficient pyramid levels is the first thing to check.

Learned: RAFT

RAFT replaces the hand-designed assumption with a learned one, and its structure is worth understanding because it explains the accuracy gap.

  1. Feature extraction. Both frames go through a shared CNN, producing per-pixel feature vectors at 1/8 resolution.
  2. All-pairs correlation. Every feature in frame 0 is dotted with every feature in frame 1, building a 4D correlation volume. For a 512×512 input at 1/8 resolution that is 64×64 × 64×64 = about 16.8 million similarity scores. This is the crucial part: matches at any displacement are computed up front, so large motion is not a special case requiring a pyramid.
  3. Iterative refinement. A GRU starts from zero flow and updates it 12–32 times, each iteration looking up correlation values around the current estimate.

On the Sintel benchmark, RAFT's published average endpoint error is about 1.6 pixels on the clean pass and under 3 on the harder final pass, several times lower than classical methods like Farnebäck. On fast motion and thin structures the gap is larger still. The cost is a GPU and roughly 100 ms per 512×512 pair versus Farnebäck's ~30 ms on CPU.

Python
import torchfrom torchvision.models.optical_flow import raft_large, Raft_Large_Weightsweights = Raft_Large_Weights.DEFAULTmodel = raft_large(weights=weights).eval().to("cuda")tf = weights.transforms()img0, img1 = tf(frame0_tensor, frame1_tensor)   # scales to [-1, 1]; H and W must be divisible by 8with torch.no_grad():    flows = model(img0.to("cuda"), img1.to("cuda"))flow = flows[-1]        # the list is every refinement iteration; take the last

Two traps. RAFT returns a list of intermediate estimates — using flows[0] instead of flows[-1] gives you the barely-refined first guess, which looks like the model is broken. And input height and width must be divisible by 8. The supplied transform only converts and normalises; it does not resize or pad, so resize or pad the frames yourself first (and crop the flow back afterwards), or the model fails on the odd shape.

Warping: turning flow into a frame

With flow in hand, synthesising the midpoint requires one more idea. You want I0.5, but you only have flow from frame 0 to frame 1 — and to fill a pixel in the intermediate frame you need to know where that pixel came from.

Assume motion is locally linear over the gap. Then flow from the intermediate frame back to frame 0 is the negative fraction of the full flow, and forward to frame 1 is the remaining fraction:

Ft→0=−t⋅F0→1,Ft→1=(1−t)⋅F0→1F_{t \to 0} = -t \cdot F_{0 \to 1}, \qquad F_{t \to 1} = (1 - t) \cdot F_{0 \to 1}

It(x)=(1−t) I0 ⁣(x+Ft→0(x))  +  t I1 ⁣(x+Ft→1(x))I_t(\mathbf{x}) = (1 - t)\,I_0\!\left(\mathbf{x} + F_{t\to 0}(\mathbf{x})\right) \; + \; t\,I_1\!\left(\mathbf{x} + F_{t\to 1}(\mathbf{x})\right)

In plain English: for each pixel of the frame you are building, step backwards along the flow to find where it was in frame 0, step forwards to find where it will be in frame 1, sample both, and blend them by how far through the gap you are.

Work one pixel. Suppose the flow at a point is (12, −4): the object moves 12 pixels right and 4 up between frames. We are synthesising t = 0.5 and filling the intermediate pixel at (106, 48).

  • Ft→0 = −0.5 × (12, −4) = (−6, 2). Sample frame 0 at (106 − 6, 48 + 2) = (100, 50).
  • Ft→1 = 0.5 × (12, −4) = (6, −2). Sample frame 1 at (106 + 6, 48 − 2) = (112, 46).
  • Blend: 0.5 × I0(100, 50) + 0.5 × I1(112, 46).

Both samples land on the same physical point — (112, 46) is exactly (100, 50) displaced by the flow. That is why the blend is sharp rather than ghosted: you are averaging the same object with itself, not two positions of it.

Real flow values are fractional, so sampling uses bilinear interpolation over the four surrounding pixels. This matters: nearest-neighbour sampling quantises every motion to whole pixels, and sub-pixel motion — a slow pan — then alternates between moving 0 and 1 pixels, producing visible stutter in what should be a smooth glide.

Python
import torch.nn.functional as Fdef warp(img, flow):    """img: (1,3,H,W), flow: (1,2,H,W). Backward warp via grid_sample."""    B, _, H, W = img.shape    ys, xs = torch.meshgrid(torch.arange(H), torch.arange(W), indexing="ij")    grid = torch.stack((xs, ys), 0).float().to(img.device)[None]    src = grid + flow                              # where to read from    # grid_sample wants coordinates normalised to [-1, 1]    src_x = 2.0 * src[:, 0] / max(W - 1, 1) - 1.0    src_y = 2.0 * src[:, 1] / max(H - 1, 1) - 1.0    norm = torch.stack((src_x, src_y), dim=-1)    return F.grid_sample(img, norm, mode="bilinear",                         padding_mode="border", align_corners=True)def interpolate(img0, img1, flow01, t=0.5):    a = warp(img0, -t * flow01)    b = warp(img1, (1.0 - t) * flow01)    return (1 - t) * a + t * b

Note this is a backward warp — for each output pixel, gather from the input. The alternative, forward warping, pushes each input pixel to its destination, and leaves holes wherever no input pixel happens to land plus collisions wherever several do. Backward warping produces a complete image by construction, which is why every practical system uses it.

Where it goes wrong: occlusion

The blend above assumes both samples are valid. Often one is not.

A person walks in front of a doorway. Part of the doorway is visible in frame 0 and hidden in frame 1 — occlusion. Part is hidden in frame 0 and revealed in frame 1 — disocclusion. In both cases there is no correspondence to find, because the content genuinely does not exist in one of the frames. The flow field will still report a vector there, because the algorithm always reports something, and that vector is meaningless. Blending it in produces a smeared, torn region around every moving object's boundary.

Detecting it: the forward–backward consistency check

Compute flow in both directions. If the flow is trustworthy, following it forward and then back should return you to where you started:

∥F0→1(x)+F1→0 ⁣(x+F0→1(x))∥2  <  α(∥F0→1∥2+∥F1→0∥2)+β\left\lVert F_{0\to1}(\mathbf{x}) + F_{1\to0}\!\left(\mathbf{x} + F_{0\to1}(\mathbf{x})\right) \right\rVert^2 \;\lt\; \alpha\left(\lVert F_{0\to1}\rVert^2 + \lVert F_{1\to0}\rVert^2\right) + \beta

with α = 0.01 and β = 0.5 as standard values. The threshold scales with motion magnitude because fast motion tolerates more absolute error.

Two pixels, worked through. Both have F0→1 = (12, −4), so ||F0→1||² = 144 + 16 = 160.

Well-matched pixelOccluded pixel
Backward flow at the destination(−11.6, 3.8)(2, 1)
Round-trip residual(0.4, −0.2), squared length 0.20(14, −3), squared length 205
Threshold0.01 × (160 + 149.0) + 0.5 = 3.590.01 × (160 + 5) + 0.5 = 2.15
Verdict0.20 < 3.59 → trust it205 > 2.15 → reject

The occluded pixel's backward flow points nowhere near home, because the surface it came from is hidden in frame 1 and the flow algorithm matched it against whatever happened to be there instead.

Acting on it

Once you have a per-pixel confidence mask, replace the fixed 50/50 blend with a weighted one. Where forward flow is trustworthy and backward flow is not, take frame 0 alone. Where the reverse holds, take frame 1 alone. Where both are trustworthy, blend as normal.

Python
conf0 = consistency_mask(flow01, flow10)      # 1 where flow01 is reliableconf1 = consistency_mask(flow10, flow01)w0 = (1 - t) * conf0w1 = t * conf1total = w0 + w1 + 1e-6frame_t = (w0 * warp(img0, -t * flow01) +           w1 * warp(img1, (1 - t) * flow01)) / total

That epsilon is not decoration. Where both confidences are zero — a region occluded from both sides, which happens with fast rotation — the denominator is zero and you get NaN propagating through the frame. Handle it explicitly: fall back to the nearer frame, or inpaint the region.

Learned interpolation: RIFE

The flow-then-warp pipeline has an architectural weakness. It estimates flow between the two input frames, then derives the intermediate flow by assuming linear motion. But what you actually need is flow from the intermediate frame — which does not exist — to each input. The linear approximation is exactly wrong for rotation, acceleration and anything following a curve.

RIFE (Real-Time Intermediate Flow Estimation) attacks that directly. Its IFNet estimates Ft→0 and Ft→1 as the primary output, coarse-to-fine, and simultaneously predicts a per-pixel fusion map deciding how much of each warped frame to use — learning the occlusion handling rather than computing it from a consistency rule. A small refinement network then fixes residual artefacts.

Farnebäck + warpRAFT + warpRIFE
Vimeo90K PSNRlowest of the threein between~35.6 dB (RIFE paper)
Speed, 480p on a mid GPU~30 ms (CPU)~100 ms~30 ms
Occlusion handlingmanual consistency maskmanual consistency masklearned fusion map
Non-linear motionpoorpoorhandled directly
Arbitrary tyesyesyes (recursive or conditioned)
Needs training datanoyesyes

The classical pipeline is still worth knowing, and not only for teaching. Flow is an independently useful signal — for stabilisation, for temporal-consistency losses, for masking moving regions — and Farnebäck runs on CPU with no model download.

Measuring quality

The convenient thing about interpolation is that ground truth exists: take a real 60 fps clip, drop every second frame, reconstruct it, and compare against the frame you removed.

PSNR converts mean squared error into decibels:

PSNR=10log⁡10 ⁣(2552MSE)\text{PSNR} = 10 \log_{10}\!\left(\frac{255^2}{\text{MSE}}\right)

Because it is logarithmic, small MSE differences move it a lot. An MSE of 10 gives 38.13 dB; an MSE of 20 gives 35.12 dB; an MSE of 50 gives 31.14 dB. Roughly, every doubling of error costs 3 dB. Above ~40 dB differences are usually invisible; below ~30 dB they are obvious.

PSNR's weakness is that it rewards blur. A slightly blurry frame has lower squared error than a sharp frame that is one pixel off — so an interpolator that hedges by averaging scores better than one that commits to a position and is nearly right. This is why PSNR alone systematically favours the ghosting you were trying to eliminate.

MetricMeasuresBlind spot
PSNRpixel-wise error, in dBrewards blur; ignores structure
SSIMlocal luminance, contrast and structure agreementstill a per-frame measure
LPIPSdistance in a learned perceptual feature spaceneeds a network; not comparable across backbones
Warping error / tOFtemporal stability across the sequencesays nothing about per-frame fidelity

The last row is the one people forget. Every metric above scores frames in isolation, but interpolation artefacts are fundamentally temporal: a sequence where each frame scores 36 dB but the artefacts jitter between frames looks far worse than one at 34 dB that is stable. Always play the result at full speed. A frame-by-frame inspection will pass footage that flickers unwatchably in motion.

What this means when you build something

Return to the 8 fps clip. The practical sequence is: run RIFE recursively — 8 → 16 → 32 fps by repeated halving, then resample to 24 — rather than asking for arbitrary t in one shot, because recursive halving keeps each interpolation between temporally close frames where motion is smallest and flow is most reliable.

The failures you will actually hit, and what causes them:

SymptomCauseFix
Ghosting on fast objects onlydisplacement exceeds what the flow estimator can findmore pyramid levels, or switch to RAFT/RIFE; or interpolate at lower resolution and upscale the flow
Torn halos around moving edgesocclusion regions blended as if validforward–backward consistency mask, or a model with a learned fusion map
Whole-frame smear at one momenta hard cut — flow between unrelated shots is meaninglessdetect scene cuts and duplicate across them instead of interpolating
Subtle stutter on slow pansnearest-neighbour sampling quantising sub-pixel motionbilinear sampling in the warp
NaN pixelsboth confidence weights zero, division by zeroepsilon in the denominator plus an explicit fallback
Good stills, bad playbackartefacts are temporal, evaluation was per-framewatch at speed; measure warping error

And know the boundary. Interpolation cannot rescue footage where the motion between frames exceeds what correspondence can resolve — a 4 fps clip of a running figure has no reliable correspondence between adjacent frames, and no interpolator will invent one. At that point you need a generative model that synthesises motion rather than one that measures it. Recognising which of the two problems you have, before you spend a day tuning flow parameters, is most of the skill.