Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

How can a dense layer in a CNN be converted into a convolutional layer?


Reshaping a dense head into a 7×7 convolutionLinear: 25,088inputs to 10Reshape weightsto 10×512×7×7Conv2d, 10filters,kernel 77×7 input:same 10 scores14×14 input: 8×8grid of scoresNo retraining; the numbers are the same, only their layout changes.
Once the dense layer is a convolution, a bigger image yields a map of predictions in one pass.

What you need to know

Why the two are the same computation

A dense layer with N outputs, fed a flattened map of C×H×W values, computes for each output n:

Text
out[n] = sum over every c, h, w of  W[n, c, h, w] × x[c, h, w]  + b[n]

A convolution with N filters, each of size C×H×W, placed at the single position where it fits, computes the same sum. The weights are the same numbers, just stored as (N, C, H, W) instead of (N, C·H·W).

Proving it in code

Python
import torch, torch.nn as nnfc = nn.Linear(512 * 7 * 7, 10)conv = nn.Conv2d(512, 10, kernel_size=7)with torch.no_grad():    conv.weight.copy_(fc.weight.view(10, 512, 7, 7))    conv.bias.copy_(fc.bias)x = torch.randn(1, 512, 7, 7)print(torch.allclose(fc(x.flatten(1)), conv(x).flatten(1), atol=1e-5))  # Truebigger = torch.randn(1, 512, 14, 14)print(conv(bigger).shape)             # torch.Size([1, 10, 8, 8])

On the original 7×7 map both give the same 10 scores. The dense layer would crash on a 14×14 map, because it expects exactly 25,088 inputs. The convolution slides its 7×7 kernel over the bigger map and returns an 8×8 grid of 10 scores, one set per position.

Why this is useful

  • Efficient sliding-window detection: running the classifier on every crop of a large image separately repeats most of the work. The converted network computes all crops in one pass, sharing the overlapping computation.
  • Semantic segmentation: FCN and U-Net output a class for every pixel. They build on this idea, adding upsampling layers to get back to full resolution.
  • Heatmaps and rough localisation: the output grid shows where in the image each class scores highest.
  • Flexible input sizes: the same model runs on 224×224 and 448×448 images.

A related everyday case: a dense layer applied to a 1×1 map, as after global average pooling, is exactly a 1×1 convolution.

A real-life example

An agritech company trains a classifier on 224×224 crops of drone images labelled "healthy crop", "weeds" or "bare soil". Farmers then want a map of a whole 4,000×3,000 field image, not one label.

Cutting the field into overlapping 224×224 crops and classifying each one works, but with a crop every 32 pixels it means about 10,000 separate forward passes and takes minutes per image. The engineer converts the classifier's dense layers to convolutions, as above, and runs the whole field image through once. The network returns a coarse grid of class scores covering the field in a single pass, which is upsampled into a weed map for the spraying drone. The overlapping crops now share their convolution work instead of repeating it.

Follow-up questions to expect

  • "What is the output resolution of the converted network?" — The input size divided by the network's total downsampling. With 5 stride-2 stages that is 32, so a 448×448 input gives a 14×14 feature map before the converted head.
  • "How does U-Net get back to full resolution?" — Its decoder upsamples with transposed convolutions or interpolation and concatenates feature maps from the encoder through skip connections, so fine detail is restored.
  • "Is global average pooling plus a linear layer also fully convolutional?" — Nearly. The linear layer on a pooled vector is a 1×1 convolution, and the network accepts any size, but pooling collapses the map to one prediction per image.