Computer Vision Fundamentals

Convolution & Pooling Layers


Try to classify images with an ordinary fully connected network and count the cost before you write a line of code.

A modest 224×224 colour image has 224 × 224 × 3 = 150,528 values. Flatten it into a vector and connect it to a hidden layer of just 1000 neurons. Every input connects to every neuron, so that single layer needs 150,528 × 1000 ≈ 150 million weights. One layer. Before you have done anything useful.

The memory bill is bad. The real problem is worse. Suppose the network learns, in that layer, to detect a vertical edge at position (30, 40). Those weights are specific to position (30, 40). If the same edge appears at (31, 40) — one pixel across — a completely different set of weights must have independently learned to detect it. The network has to relearn "vertical edge" separately at all 50,176 positions, and it can only do that if your training data happens to contain that edge at every position.

Flattening also destroys structure. Pixel (30, 40) and pixel (31, 40) are adjacent in the image and almost certainly related. After flattening they are indices 20,200 and 20,201 in a vector, and the network has no idea they were ever neighbours. All the spatial information — the entire reason an image is a grid and not a bag of numbers — is thrown away in the first operation.

Convolution fixes both problems with a single idea: learn one small pattern detector and slide it across the whole image.

One 3 by 3 kernel on its receptive field3012715893272510131742162c0c1c2c3c4r0r1r2r3r4Nine multiplies and a sum give one output pixel; slide by the stride and repeat.
The shaded window is the whole receptive field of one output value — everything outside it is invisible to that neuron.

The convolution operation, worked by hand

A kernel (or filter) is a small grid of weights, typically 3×3. You place it over a patch of the image, multiply each kernel weight by the pixel underneath it, add up all the products, and write the result into an output grid. Then you slide the kernel one position across and repeat.

Take a 5×5 input with a bright vertical bar down the middle:

Text
Input                    Kernel (vertical edge detector)0  0  100  0  0          -1  0  10  0  100  0  0          -1  0  10  0  100  0  0          -1  0  10  0  100  0  00  0  100  0  0

Place the kernel at the top-left, covering rows 0–2 and columns 0–2:

Text
patch:   0    0   100        kernel:  -1   0   1         0    0   100                 -1   0   1         0    0   100                 -1   0   1sum = (0)(-1) + (0)(0) + (100)(1)    + (0)(-1) + (0)(0) + (100)(1)    + (0)(-1) + (0)(0) + (100)(1)    = 300

Slide one column right, so the patch covers columns 1–3:

Text
patch:   0  100    0         0  100    0     sum = 0 + 0 + 0 = 0         0  100    0

Slide again, columns 2–4:

Text
patch: 100    0    0       100    0    0     sum = -300       100    0    0

The full output is a 3×3 grid where every row reads 300, 0, -300. The kernel produced a large positive response where the image goes from dark to bright left-to-right, zero where nothing changes horizontally, and a large negative response at the bright-to-dark transition. It found the edges. Nine weights did that, and the same nine weights would find a vertical edge anywhere in a 4000-pixel-wide image.

That is the whole bargain: a small set of weights, reused at every position, detecting a local pattern wherever it occurs. Parameters collapse from millions to nine, and the detector works everywhere without being retrained per location.

Two properties fall out of this, and both have names worth knowing. Parameter sharing is the reuse of one kernel across all positions. Translation equivariance is the consequence: shift the input by three pixels and the output shifts by three pixels, with the same values. The detector does not care where the pattern is.

The formula

For a single-channel input II and a k×kk \times k kernel KK, the output at position (i,j)(i, j) is:

O(i,j)=∑m=0k−1∑n=0k−1I(i+m,  j+n)⋅K(m,n)  +  bO(i,j) = \sum_{m=0}^{k-1} \sum_{n=0}^{k-1} I(i+m,\; j+n) \cdot K(m,n) \;+\; b

The bias bb is a single learned number added to every position, letting the whole feature map shift up or down.

A pedantic note that occasionally matters: what deep learning calls convolution is, strictly, cross-correlation — true convolution flips the kernel first. Since the kernel weights are learned, a flipped kernel is just as learnable as an unflipped one, so the distinction has no practical consequence. It only matters when you compare a framework's output against a signal-processing textbook and the sign confuses you.

The four knobs that shape a convolution

Kernel size

How large a neighbourhood each output value sees. 3×3 dominates modern architectures, and the reason is a genuinely clever piece of arithmetic.

Two stacked 3×3 convolutions cover the same 5×5 region as one 5×5 convolution — the second layer sees a 3×3 window of outputs, each of which already summarised a 3×3 window of the input. But count the weights, per input and output channel:

ConfigurationReceptive fieldWeights (per channel pair)Non-linearities
One 5×5 conv5×5251
Two 3×3 convs5×5182
One 7×7 conv7×7491
Three 3×3 convs7×7273

The stack is cheaper and more expressive, because each extra layer brings another activation function and therefore another chance to represent something non-linear. This is why "just use 3×3, stacked deeper" became the default.

The 1×1 convolution deserves its own note, because it looks pointless. A 1×1 kernel sees a single pixel — how can that be useful? Because it still spans all input channels. A 1×1 convolution is a learned linear combination across the channel dimension at each spatial position. Its main jobs are cutting channel count cheaply before an expensive layer, and mixing channel information without touching spatial extent.

Stride

How far the kernel jumps between positions. Stride 1 evaluates at every pixel. Stride 2 skips every other one, halving the output dimensions and quartering the number of values.

Strided convolution is the modern alternative to pooling for downsampling. It reduces size while learning how to reduce, rather than applying a fixed rule.

Padding

Without padding, a 3×3 kernel on a 5×5 input can only be centred on the inner 3×3 region — the border pixels have no full neighbourhood. So the output shrinks, and it shrinks every layer. Stack twenty 3×3 layers and you have lost 40 pixels from each dimension. A 224×224 input becomes 184×184, and every layer has been progressively blind to its borders.

ModePadding5×5 input, 3×3 kernelEffect
Valid03×3Shrinks; no invented values
Same(k−1)/2(k-1)/25×5Preserves size; border pixels are partly artificial

"Same" padding with p=1p = 1 for a 3×3 kernel is the near-universal choice, because it lets you build deep stacks without tracking a shrinking grid. The cost is that border outputs are computed partly from padded values — usually zeros — so they carry slightly less real information. Reflecting the border content instead of padding with zeros avoids a hard artificial edge, and matters more in dense prediction tasks like segmentation than in classification.

Number of filters

One kernel detects one pattern. You want many, so a convolution layer holds a bank of filters — 64, 128, 256 — each learning a different pattern, each producing its own output grid. Those grids stack into the output's channel dimension.

Output size

The formula you will use constantly:

O=⌊I−k+2ps⌋+1O = \left\lfloor \frac{I - k + 2p}{s} \right\rfloor + 1

where II is input size, kk kernel size, pp padding, ss stride.

InputkpsWorkingOutput
224311(224 - 3 + 2)/1 + 1224
224312⌊223/2⌋ + 1 = 111 + 1112
224732⌊223/2⌋ + 1 = 111 + 1112
32501(32 - 5 + 0)/1 + 128

Learn to do this in your head. The commonest runtime error in a hand-built network is a shape mismatch where a convolutional stack meets a linear layer, and it happens because someone guessed the flattened size instead of computing it.

Multiple channels: the part that trips everyone

Real inputs have depth. An RGB image has 3 channels; a mid-network feature map might have 256. Here is the rule that people get wrong:

A filter is not 2D. Its depth always equals the input's channel count, and it produces exactly one output channel by summing across all input channels.

Work an example. Input is 32×32×3. You want 64 output channels with 3×3 kernels.

  • Each filter has shape 3×3×3 — height, width, and input depth — which is 27 weights.
  • Applying it: multiply all 27 weights by the corresponding 27 input values, sum them into one number, add a bias. Do that at every spatial position to get a 32×32×1 map.
  • You have 64 such filters, so you get 64 maps, stacked into a 32×32×64 output.
  • Total parameters: 64×(3×3×3)+64=1728+64=179264 \times (3 \times 3 \times 3) + 64 = 1728 + 64 = 1792.

The general formula for parameters in a convolution layer:

params=Cout×(k×k×Cin)+Cout\text{params} = C_{out} \times (k \times k \times C_{in}) + C_{out}

Compare 1792 parameters against the 150 million of that fully connected layer. That is the whole reason convolutional networks exist.

Python
import torch.nn as nnlayer = nn.Conv2d(in_channels=3, out_channels=64, kernel_size=3, padding=1)print(layer.weight.shape)   # torch.Size([64, 3, 3, 3])  -> C_out, C_in, k, kprint(sum(p.numel() for p in layer.parameters()))   # 1792

Convolution variants that buy efficiency

Depthwise separable convolution

Standard convolution does two jobs simultaneously: it mixes information spatially (across the 3×3 window) and across channels. Splitting those into two steps is dramatically cheaper.

  1. Depthwise: one 3×3 kernel per input channel, applied independently. No cross-channel mixing. Cost: Cin×9C_{in} \times 9.
  2. Pointwise: a 1×1 convolution mixing the channels. Cost: Cin×CoutC_{in} \times C_{out}.

For Cin=Cout=256C_{in} = C_{out} = 256 with a 3×3 kernel:

ApproachCalculationParameters
Standard 3×3256×3×3×256256 \times 3 \times 3 \times 256589,824
Depthwise256×3×3256 \times 3 \times 32,304
Pointwise256×1×1×256256 \times 1 \times 1 \times 25665,536
Separable total2,304 + 65,53667,840

An 8.7× reduction for a small accuracy cost. This is the core of MobileNet and every other architecture designed for phones.

Dilated convolution

A dilated (or atrous) convolution spreads the kernel's sampling points apart with gaps. A 3×3 kernel with dilation 2 samples pixels 2 apart, covering a 5×5 region while still using only 9 weights.

Text
dilation 1        dilation 2x x x             x . x . xx x x             . . . . .x x x             x . x . x                  . . . . .                  x . x . x

This grows the receptive field without adding parameters and without downsampling. That combination is exactly what semantic segmentation needs, since it must produce a per-pixel output at full resolution while still seeing enough context to know a pixel belongs to a bus rather than a car.

Grouped convolution

Split the input channels into gg groups, convolve each group with its own filters, concatenate. Parameters drop by a factor of gg. Originally an engineering hack to split AlexNet across two GPUs with limited memory, it turned out to be a useful design tool and reappears in ResNeXt. Depthwise convolution is simply grouped convolution taken to the extreme where g=Cing = C_{in}.

Pooling: deliberate loss of precision

Convolution is equivariant — move the input, the output moves. Often you want invariance instead: a cat one pixel to the left is still a cat, and the network should not care. Pooling provides that by summarising a neighbourhood into one number.

Max pooling takes the largest value in each window:

Text
Input 4x4              2x2 max pool, stride 2 1   3   2   4 5   6   7   8    ->    6   8 9  10  11  12         14  1613  14  15  16

Sixteen values become four. Shift the input by one pixel and most of those maxima stay the same — that is the invariance you wanted. Max pooling asks "was this feature present anywhere nearby?" and discards exactly where.

Average pooling takes the mean instead. It preserves overall intensity rather than peak response, and is gentler.

MethodKeepsGood for
MaxStrongest responseDetecting whether a feature exists; sparse features like edges
AverageOverall levelSmooth summarisation; final aggregation
Global averageOne number per channelReplacing the flatten-plus-dense head

Global average pooling is worth dwelling on. It averages each entire feature map to a single number, turning a 7×7×512 tensor into a 512-vector. Compare that with flattening: 7×7×512 = 25,088 values, which connected to a 1000-class output layer would need 25 million parameters. Global average pooling needs zero, and it accepts any input resolution, since averaging does not care about grid size. Nearly every architecture after 2014 uses it.

The modern trend is towards fewer pooling layers, replaced by strided convolutions. Pooling applies a fixed, unlearned rule; a strided convolution learns how to downsample. Pooling has not disappeared, but it is no longer automatic.

Receptive field: how much each unit actually sees

The receptive field of a unit is the region of the original input that can influence it. A first-layer 3×3 unit sees 3×3 pixels. That is far too little to recognise anything meaningful, which is why depth matters.

Stacking grows it. For layers with kernel klk_l and stride sls_l:

RFl=RFl−1+(kl−1)∏i<lsiRF_{l} = RF_{l-1} + (k_l - 1) \prod_{i<l} s_i

Concretely, with 3×3 convolutions at stride 1:

After layerReceptive fieldRoughly enough to see
13×3An edge fragment
25×5A corner
37×7A small texture motif
511×11A simple part

Growth is linear and slow — painfully slow if you need to see a whole object in a 224×224 image. Downsampling changes that. Every stride-2 layer doubles the multiplier on all subsequent growth, so the receptive field expands geometrically. This is the real reason architectures reduce spatial resolution as they go deep: not to save compute, but to let later layers see enough of the image to reason about objects rather than textures.

If a network underperforms on large objects, an insufficient receptive field is a prime suspect. Add depth, add downsampling, or add dilation.

Non-linearity, and why the whole thing collapses without it

Convolution is a linear operation. Stack two linear operations and you get a linear operation — the composition of two matrix multiplications is one matrix multiplication. A hundred convolution layers with nothing between them have exactly the representational power of one.

So an activation function goes after every convolution. ReLU is the standard:

ReLU(x)=max⁡(0,x)\text{ReLU}(x) = \max(0, x)

It is trivially cheap, and crucially its gradient is 1 for all positive inputs. Sigmoid and tanh squash their outputs into a bounded range, so their gradients shrink towards zero at the extremes; multiply many such gradients together during backpropagation through a deep network and the signal reaching the early layers vanishes. ReLU's constant gradient does not decay, which is what made networks deeper than a handful of layers trainable at all.

ReLU's own failure mode is the dying ReLU: a unit whose weights drift such that it outputs zero for every input in the dataset. Its gradient is then zero too, so it can never recover — permanently dead capacity. Leaky ReLU, which returns 0.01x0.01x for negative inputs instead of 0, keeps a small gradient flowing and prevents this.

Putting a block together

The canonical convolutional block is convolution, then normalisation, then activation, repeated, then downsampling:

Python
import torchimport torch.nn as nnclass ConvBlock(nn.Module):    def __init__(self, c_in, c_out):        super().__init__()        self.block = nn.Sequential(            # bias=False: BatchNorm has its own shift, so the conv bias is redundant            nn.Conv2d(c_in, c_out, kernel_size=3, padding=1, bias=False),            nn.BatchNorm2d(c_out),            nn.ReLU(inplace=True),            nn.Conv2d(c_out, c_out, kernel_size=3, padding=1, bias=False),            nn.BatchNorm2d(c_out),            nn.ReLU(inplace=True),            nn.MaxPool2d(kernel_size=2, stride=2),   # halves height and width        )    def forward(self, x):        return self.block(x)net = nn.Sequential(    ConvBlock(3, 64),      # 32x32 -> 16x16    ConvBlock(64, 128),    # 16x16 -> 8x8    ConvBlock(128, 256),   # 8x8   -> 4x4    nn.AdaptiveAvgPool2d(1),   # -> 256 x 1 x 1, any input size    nn.Flatten(),    nn.Linear(256, 10),)print(net(torch.randn(2, 3, 32, 32)).shape)   # torch.Size([2, 10])

The pattern of channels doubling as spatial size halves is deliberate. Each block roughly preserves the total amount of information while trading spatial detail for semantic richness: early layers hold many positions and few concepts, later layers hold few positions and many concepts.

What this means when you design a network

Compute the shapes before you write the code. Take your input size and walk it through every layer with the output-size formula, writing the numbers down. It takes two minutes and prevents the single most common class of bug, where a linear layer expects 4096 inputs and receives 8192.

When your model runs but underperforms, ask three diagnostic questions in order. Is the receptive field large enough to cover the objects you care about? If not, go deeper or add dilation. Are you downsampling too aggressively for the level of detail the task needs? Segmentation and small-object detection suffer badly from an over-reduced feature map. Is there an activation between every pair of convolutions? A missing one silently collapses two layers into one, and nothing will warn you.

And when compute is the constraint, reach for the structural savings before shrinking the network. Replacing standard convolutions with depthwise separable ones cuts parameters by roughly an order of magnitude at a small accuracy cost, and a 1×1 convolution placed before an expensive 3×3 layer to cut its input channels is nearly free. Both are far better trades than simply making the model shallower.