Course Content
Deep Learning Essentials
13 sections · 61 lessons
How can a dense layer in a CNN be converted into a convolutional layer?
What you need to know
Why the two are the same computation
A dense layer with N outputs, fed a flattened map of C×H×W values, computes for each output n:
out[n] = sum over every c, h, w of W[n, c, h, w] × x[c, h, w] + b[n]A convolution with N filters, each of size C×H×W, placed at the single position where it fits, computes the same sum. The weights are the same numbers, just stored as (N, C, H, W) instead of (N, C·H·W).
Proving it in code
1import torch, torch.nn as nn23fc = nn.Linear(512 * 7 * 7, 10)4conv = nn.Conv2d(512, 10, kernel_size=7)5with torch.no_grad():6 conv.weight.copy_(fc.weight.view(10, 512, 7, 7))7 conv.bias.copy_(fc.bias)89x = torch.randn(1, 512, 7, 7)10print(torch.allclose(fc(x.flatten(1)), conv(x).flatten(1), atol=1e-5)) # True1112bigger = torch.randn(1, 512, 14, 14)13print(conv(bigger).shape) # torch.Size([1, 10, 8, 8])On the original 7×7 map both give the same 10 scores. The dense layer would crash on a 14×14 map, because it expects exactly 25,088 inputs. The convolution slides its 7×7 kernel over the bigger map and returns an 8×8 grid of 10 scores, one set per position.
Why this is useful
- Efficient sliding-window detection: running the classifier on every crop of a large image separately repeats most of the work. The converted network computes all crops in one pass, sharing the overlapping computation.
- Semantic segmentation: FCN and U-Net output a class for every pixel. They build on this idea, adding upsampling layers to get back to full resolution.
- Heatmaps and rough localisation: the output grid shows where in the image each class scores highest.
- Flexible input sizes: the same model runs on 224×224 and 448×448 images.
A related everyday case: a dense layer applied to a 1×1 map, as after global average pooling, is exactly a 1×1 convolution.
A real-life example
An agritech company trains a classifier on 224×224 crops of drone images labelled "healthy crop", "weeds" or "bare soil". Farmers then want a map of a whole 4,000×3,000 field image, not one label.
Cutting the field into overlapping 224×224 crops and classifying each one works, but with a crop every 32 pixels it means about 10,000 separate forward passes and takes minutes per image. The engineer converts the classifier's dense layers to convolutions, as above, and runs the whole field image through once. The network returns a coarse grid of class scores covering the field in a single pass, which is upsampled into a weed map for the spraying drone. The overlapping crops now share their convolution work instead of repeating it.
Follow-up questions to expect
- "What is the output resolution of the converted network?" — The input size divided by the network's total downsampling. With 5 stride-2 stages that is 32, so a 448×448 input gives a 14×14 feature map before the converted head.
- "How does U-Net get back to full resolution?" — Its decoder upsamples with transposed convolutions or interpolation and concatenates feature maps from the encoder through skip connections, so fine detail is restored.
- "Is global average pooling plus a linear layer also fully convolutional?" — Nearly. The linear layer on a pooled vector is a 1×1 convolution, and the network accepts any size, but pooling collapses the map to one prediction per image.