Image and Video Generation

ControlNet, Inpainting, and Outpainting — Controlling Where Diffusion Paints


A client wants a product shot: a green bottle standing on the left third of the frame, tilted about 15 degrees to the right, with the label facing the camera. You write the prompt. You get a bottle in the centre. You add "on the left side of the frame". You get a bottle slightly left of centre, upright. You add "tilted", "angled", "leaning", "at a 15 degree angle". After forty generations you have forty compositions, none of which is the one you were asked for.

This is not a prompting skill problem. It is an information problem. Your prompt is a sequence of at most 77 CLIP tokens — call it a few hundred bits. The image is 512×512×3, and even in latent space it is 16,384 numbers. Text conditioning tells the model what to draw and roughly in what style. It has no channel through which to say where the edges go. Cross-attention gives every spatial position a soft preference over prompt tokens; it does not give any position a coordinate.

Text can specifyText cannot specify
subject, style, mood, lighting qualityexact pose or limb positions
medium: oil painting, 35mm photo, 3D renderprecise object placement and scale
rough count for small numbersexact geometry, perspective lines, camera angle
colour palette and general composition moodwhich region to change and which to leave alone

The three techniques below all attack the same gap, from different directions. ControlNet adds a second conditioning channel that carries spatial structure. Inpainting restricts where the model is allowed to paint. Image-to-image restricts how far it is allowed to move from something you already have.

One frozen model, many steering wheelsFrozen base modelCanny edges:hold the outlineDepth map: holdthe 3D layoutPose skeleton:hold the figureInpaint mask:repaint inside itOutpaint canvas:extend past the edge
Zero-initialised convolutions mean an untrained ControlNet starts as a no-op, so training can only add control.

ControlNet: a steering wheel bolted to a frozen engine

The obvious approach is to fine-tune the model on pairs of (edge map, photo). It works, and it is a bad idea. You need a large paired dataset, the run costs real money, and you have produced a model that only does edges — a second one for pose, a third for depth, and each is a full 860-million-parameter checkpoint you now have to store and swap.

ControlNet's answer: freeze the original model entirely and train a small side network that injects spatial hints into it.

The architecture, and why zero convolutions are the whole trick

Take the U-Net. Freeze every weight. Now make a trainable copy of just the encoder half — the downsampling blocks and the bottleneck — and feed it the control image (an edge map, a depth map, a pose skeleton) in addition to the noisy latent. At each resolution, the copy's output passes through a zero convolution — a 1×1 convolution whose weights and bias are both initialised to exactly zero — and the result is added into the frozen network's corresponding skip connection.

Text
noisy latent ──┬──> [FROZEN encoder] ──> [FROZEN bottleneck] ──> [FROZEN decoder] ──> noise               │           │  skip                                     ▲               │           └───────────────────────────────────────────┤ (+)               │                                                       │control image ─┴──> [zero conv] ──> [TRAINABLE copy] ──> [zero conv] ──┘

At initialisation every zero convolution outputs zero, so the sum adds zero, so the network's behaviour is bit-identical to the original model. That is the property that matters: training starts from a known-good model rather than from a perturbed one, so there is no phase where quality collapses before it recovers.

The natural objection is that zero weights cannot learn — and for a zero-initialised layer in an ordinary network, that is often true. Here it is not, and the reason is worth working through. For a zero convolution computing y = Wx + b with W = 0:

y=0,∂L∂W=∂L∂y⋅x⊤y = 0, \qquad \frac{\partial \mathcal{L}}{\partial W} = \frac{\partial \mathcal{L}}{\partial y} \cdot x^{\top}

In plain English: the gradient with respect to the weights depends on the layer's input, not on its weights. The input x is the trainable copy's activation, which is non-zero because that copy inherited the pretrained encoder's weights. And ∂L/∂y is non-zero because y is added into a live network. So the very first optimiser step moves W off zero, and from then on it is an ordinary layer. Zero initialisation buys a safe start with no cost in trainability.

ControlNet does not modify the base model at all. It computes a correction and adds it in, starting from a correction of exactly zero — which is why a half-trained ControlNet degrades gracefully instead of catastrophically.

The practical consequences: a ControlNet for SD 1.5 is roughly 360 million parameters (about 700 MB at fp16) that ride alongside the base model instead of replacing it; the base model stays swappable, so any ControlNet works with any SD 1.5 derivative; training needs on the order of 50,000 image pairs rather than millions; and you can stack several at once.

Control modalities

ModalityWhat it locksWhat it leaves freeBest for
Canny edgesevery visible outline, tightlycolour, texture, material, lightingrestyling a photo, line-art colouring
Depth map3D layout and object distanceoutlines, texture, exact silhouetteskeeping a room's geometry, changing everything in it
OpenPoseskeleton joint positions onlybody shape, clothing, face, backgroundcharacter posing, animation keyframes
Scribble / HEDrough shape, looselyalmost everything elsesketch to finished image
Segmentationregion boundaries and semantic classall appearance within each regionarchitectural and interior layouts
Normal mapsurface orientation and reliefcolour, albedo, base geometrypreserving fine sculptural detail

Choosing badly is the most common ControlNet mistake. Canny on a photograph of a person, hoping to change their clothes, will not work: canny locks the clothing outlines, so the model can only recolour what is already there. You wanted OpenPose. Conversely, depth for a line-art colouring job gives you nothing to colour inside — a flat drawing has no depth — and you wanted canny.

End to end, with canny

Python
import cv2, numpy as np, torchfrom PIL import Imagefrom diffusers import StableDiffusionControlNetPipeline, ControlNetModelsource = np.array(Image.open("reference.png").convert("RGB"))edges = cv2.Canny(source, 100, 200)                 # low / high hysteresiscontrol = Image.fromarray(np.stack([edges] * 3, axis=-1))controlnet = ControlNetModel.from_pretrained(    "lllyasviel/sd-controlnet-canny", torch_dtype=torch.float16)pipe = StableDiffusionControlNetPipeline.from_pretrained(    "stable-diffusion-v1-5/stable-diffusion-v1-5", controlnet=controlnet,    torch_dtype=torch.float16).to("cuda")image = pipe(    prompt="a green glass bottle on a marble surface, studio lighting",    image=control,    controlnet_conditioning_scale=0.8,    num_inference_steps=25,    generator=torch.Generator("cuda").manual_seed(0),).images[0]

The two canny thresholds are hysteresis bounds on gradient magnitude, measured on a 0–255 scale. Pixels above the high threshold are definitely edges; pixels between the two are edges only if connected to a definite one. Raising them from (100, 200) to (150, 250) drops fine texture and keeps only strong outlines, which gives the model far more freedom. That parameter pair matters more than most prompt words.

controlnet_conditioning_scale multiplies the injected correction before it is added:

ScaleBehaviourUse when
0.3–0.5a suggestion; the model may ignore partsloose scribbles, stylistic hints
0.6–0.9followed clearly, prompt still dominantthe usual working range
1.0strictly followedexact structure required
1.2+overrides the prompt; artefacts along edgesalmost never — usually a sign the control image is wrong

Depth control is identical in shape; only the preprocessor changes. A depth estimator such as MiDaS or Depth Anything turns the source photo into a greyscale map where brighter means closer, and the depth ControlNet consumes that. Because depth carries no outlines, the model is free to redesign every surface while keeping the room's geometry exactly — which is why depth is the standard choice for interior restyling.

Stacking several at once

Python
pipe = StableDiffusionControlNetPipeline.from_pretrained(    "stable-diffusion-v1-5/stable-diffusion-v1-5",    controlnet=[pose_net, depth_net],           # a list    torch_dtype=torch.float16).to("cuda")image = pipe(prompt="a knight in a cathedral",             image=[pose_image, depth_image],             controlnet_conditioning_scale=[0.9, 0.5]).images[0]

The corrections are summed, so scales compound. Two nets at 1.0 each behave roughly like one net at 2.0: over-constrained, and you get muddy edge artefacts and blown contrast. The working rule is to keep the total near 1.2–1.5, giving the primary control most of it. Contradictory controls — a pose skeleton standing up and a depth map of someone sitting — produce anatomical horrors, because the model is being pulled towards two incompatible images at once.

Inpainting: restricting where the model may paint

You have a good photograph with one problem: a car parked in the background. You want the car gone and the wall behind it plausible.

The naive approach, and exactly why it fails

The obvious method is called blended latent diffusion: run normal generation, but after every denoising step, overwrite the unmasked region of the latent with the original image noised to that same timestep. The masked region is generated; everything else is forced back to the original.

Python
for t in scheduler.timesteps:    latents = denoise_step(latents, t)    original_noised = scheduler.add_noise(original_latents, noise, t)    latents = latents * mask + original_noised * (1 - mask)   # force it back

It half works, and the failure is instructive. The denoiser never knows a mask exists. It sees a latent that happens to be discontinuous and treats the whole thing as one image to denoise. Inside the masked region it has no idea it is supposed to be completing a wall that continues from outside; it is generating whatever the prompt suggests, in isolation, and finding out about the surroundings only through whatever leaks across the boundary in the next step. The results are hard seams, lighting that does not match, and a patch that reads as pasted-in even when it is individually sharp.

The dedicated inpainting checkpoint

A proper inpainting model changes the U-Net's input. Where the standard SD 1.5 U-Net takes 4 input channels (the noisy latent), the inpainting variant takes 9:

ChannelsContentsWhy the model needs it
4the noisy latent being denoisedthe usual input
4VAE encoding of the image with the masked region blacked outfull-resolution context: what the surroundings actually look like
1the downsampled binary maskexplicit knowledge of which cells it is responsible for

Now the model is told where the hole is and what surrounds it, and it was trained on millions of examples of filling holes consistently. Seam quality is not close. You can check which kind of checkpoint you have loaded:

Python
print(pipe.unet.config.in_channels)   # 4 = standard, 9 = inpaintingimage = pipe(    prompt="a plain brick wall, matching afternoon light",    image=photo, mask_image=mask,            # white = repaint, black = keep    strength=1.0, num_inference_steps=30,).images[0]

Mask quality decides everything

Here is the number that explains most inpainting failures. The VAE downsamples by 8, so a 512×512 mask becomes 64×64 in latent space. One latent cell covers an 8×8 block of pixels. A mask boundary drawn with single-pixel precision is therefore rounded to the nearest 8-pixel block — your careful selection edge does not survive encoding. Worse, a hard binary edge produces an abrupt discontinuity in latent space that the model renders as a visible line.

The fix is to feather the mask by 8–32 pixels with a Gaussian blur, so the transition is gradual across several latent cells:

Python
from PIL import ImageFiltermask = mask.filter(ImageFilter.GaussianBlur(radius=12))
Mask problemVisible symptomFix
Hard binary edgesharp rectangular seam around the editGaussian blur, radius 8–32 px
Mask too tight to the objecta halo or ghost outline of the removed objectdilate by 10–20 px before blurring
Mask too largethe model reinvents surroundings you wanted kepttighten; inpaint in two smaller passes
Region smaller than ~64×64 pxblurry mush — under 8×8 latent cells to work withcrop a larger area, inpaint, composite back
Inverted polarityeverything except the target gets regeneratedwhite repaints, black keeps — verify before running

For precise object-level masks, run a segmentation model rather than drawing by hand. Segment Anything (SAM) accepts a click point and returns a pixel-accurate object mask; dilate and blur it, and you have a far better mask than any brush stroke.

Outpainting: the same mechanism, pointed outwards

Outpainting extends an image beyond its borders. Mechanically it is inpainting with a constructed mask: paste the original onto a larger canvas, mark everything outside it as "repaint", and run an inpainting model.

Python
from PIL import Imageorig = Image.open("photo.png")                  # 512 x 512canvas = Image.new("RGB", (768, 512), (127, 127, 127))canvas.paste(orig, (0, 0))                      # extend 256 px to the rightmask = Image.new("L", (768, 512), 0)            # 0 = keepmask.paste(255, (448, 0, 768, 512))             # repaint from x=448, not x=512mask = mask.filter(ImageFilter.GaussianBlur(12))

The important line is the mask starting at x = 448 rather than x = 512. That 64-pixel overlap into the original is deliberate: it gives the model a strip it is allowed to regenerate and can see, so it can blend the join instead of butting a new region against an untouched edge. Without overlap you get a vertical seam every time.

The second rule is extend in steps. Going from 512 to 2048 in one pass gives the model 1,536 pixels of unknown territory conditioned on a thin strip of context; it drifts, and the far end has nothing to do with the original. Four passes of 384 pixels each keep every new region adjacent to something real. The trade-off is compounding style drift, so re-anchor the prompt each pass and consider a final low-strength image-to-image over the whole canvas to unify it.

Image-to-image and the strength parameter

Image-to-image is the third restriction: instead of starting from pure noise, encode an existing image, noise it partway, and denoise from there. The strength parameter decides how far "partway" is, and it does two things at once that people usually only notice one of.

With num_inference_steps=50 and strength=0.6:

  • The starting latent is noised to roughly timestep 600 of 1000. On SD 1.5's schedule ā is about 0.16 there, so the signal weight √ā is about 0.40 against a noise weight of about 0.92: only ~40% of the original amplitude survives, under noise more than twice as strong.
  • Only 30 denoising steps actually run (50 × 0.6). This is the part people miss: raising strength costs proportionally more compute, and lowering it makes generation genuinely faster.
strengthSteps run (of 50)What survivesTypical use
0.15–0.258–13everything; only texture and grain changefinal polish, unifying a composite, upscale refinement
0.3–0.4515–23composition and all major shapescolour grading, subtle restyling
0.5–0.6525–33rough layout and colour massessketch to painting, meaningful style change
0.7–0.8535–43only the broadest arrangementusing the source as a loose composition guide
0.95–1.048–50essentially nothingequivalent to text-to-image; the input is wasted

Strength is not a slider from "similar" to "different" — it is the timestep you start from, and therefore also your step count. At 0.2 you are running a fifth of the compute, which is exactly why low-strength refinement passes are so cheap.

Chaining the tools

Serious work is rarely one generation. A realistic pipeline for the bottle brief from the opening:

  1. Lock the geometry. Photograph or sketch a bottle in the exact pose, run canny, generate with ControlNet at scale 0.9. Composition now matches the brief.
  2. Fix the label. The label is garbled, because the VAE cannot preserve text at that size. Mask it, dilate 15 px, blur 12 px, inpaint with a dedicated inpainting checkpoint. If it is still wrong, crop a 512×512 region around the label, inpaint at full resolution, composite back — this gives the model 8× the latent cells for the same content.
  3. Extend the frame. The client wants 16:9. Outpaint left and right, one pass of about 200 pixels per side with a 64-pixel overlap, taking the 512×512 frame to roughly 912×512.
  4. Upscale. Run a dedicated upscaler to about 2048×1152, then an image-to-image pass at strength 0.22 to reintroduce plausible detail and unify the composite.

The ordering is not arbitrary. Structure is decided first, because every later step preserves it. Extension happens before upscaling, because outpainting a 2048-wide image costs four times as much as outpainting a 1024-wide one for the same result. The low-strength pass goes last, because its whole purpose is to harmonise the seams that earlier steps introduced.

What this means when you build something

The habit to form is diagnosing which constraint is missing before reaching for a tool. Most reported failures are a mismatch between the problem and the mechanism.

SymptomReal causeFix
ControlNet output is a flat, lifeless copy of the control imageconditioning scale at or above 1.2, or edges too densedrop to 0.7; raise canny thresholds to (150, 250)
Pose control ignored entirelyOpenPose preprocessor found no person and emitted a blank skeletoninspect the control image itself — always view it before generating
Inpainted region has a visible rectangular borderhard-edged mask quantised to 8-pixel latent blocksGaussian blur radius 12; dilate first if a halo remains
Inpainted small object is blurry mushregion occupies fewer than ~8×8 latent cellscrop wider, inpaint at 512, composite back
Outpainted region drifts away in styletoo much new area per pass, too little context256–384 px per pass, 64 px overlap, re-anchor the prompt
img2img returns almost the originalstrength below ~0.2 barely moves the latentraise to 0.4, or use ControlNet if you want structure kept at higher strength
Inpainting seams even with a good maska standard 4-channel checkpoint loaded, not a 9-channel onecheck unet.config.in_channels

On performance: these pipelines are not free. A ControlNet adds roughly 30–45% to inference time because the trainable encoder copy runs every step alongside the frozen network, and two stacked ControlNets roughly double that overhead. Preprocessors — canny, depth estimation, pose detection — run once per control image, so cache their output rather than recomputing it for every seed you try. And every one of these techniques still runs classifier-free guidance, so the true step count is double what you set.

Finally, always look at the control image, the mask, and the preprocessed input before you look at the output. A startling share of "the model ignored my control" reports are a blank depth map, an inverted mask, or a pose detector that found nothing. The tools do exactly what you tell them; the bug is almost always upstream of the model.