Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is pooling, and how does it help filter features?


2×2 max pooling on a 4×4 feature map1320461102951137Output: 6, 2, 2, 9. Sixteen values become four.
Each window keeps its strongest match and forgets where in the window it was.

What you need to know

Max and average pooling on real numbers

A 4×4 feature map, with a 2×2 window and stride 2, gives four windows:

Text
feature map            max pool 2×2       average pool 2×21  3 | 2  04  6 | 1  1            6  2               3.5  1.0-----+-----    ->0  2 | 9  5            2  9               1.0  6.01  1 | 3  7

Top-left window {1, 3, 4, 6} gives max 6 and average 3.5. Bottom-right {9, 5, 3, 7} gives max 9 and average 6. The map goes from 16 values to 4.

How that "filters" features

A feature map value says how strongly a filter matched at that position. Max pooling asks, for each region, "did this feature appear anywhere here, and how strongly?" It keeps the strongest match and throws away the weaker, noisier responses and the precise location. If the feature moves by one pixel inside the window, the output is unchanged, which is the small translation tolerance people mention.

Average pooling keeps a smoother summary, which is better when the overall amount of a feature matters more than its peak.

What happens in the backward pass

Max pooling passes the gradient only to the position that held the maximum; the other three positions in the window get zero. Average pooling splits the gradient equally among all four. Either way, there are no weights to update.

Global average pooling

Global average pooling averages each whole channel to one number. A ResNet's last feature map of 2048×7×7 becomes a vector of 2,048 values. That replaces a huge flatten-plus-dense head, reduces overfitting, and lets the network accept any input size.

Python
import torch, torch.nn as nnx = torch.randn(1, 2048, 7, 7)print(nn.MaxPool2d(2)(torch.randn(1, 64, 56, 56)).shape)   # [1, 64, 28, 28]print(nn.AdaptiveAvgPool2d(1)(x).flatten(1).shape)         # [1, 2048]

The modern view

Many newer architectures downsample with strided convolutions instead of pooling, so the network learns how to downsample. Max pooling is still common in the early layers of CNNs and remains a good default in small models.

A real-life example

In a chest X-ray model, one filter in an early layer responds to the hazy texture of a lung opacity. On a patient's image, it fires strongly (8.7) at one position and weakly (0.4 to 1.2) around it. After 2×2 max pooling, the pooled map keeps 8.7 for that region. If the same patient is imaged again a few millimetres to one side, the peak moves one pixel but stays inside the same window, so the pooled value, and everything after it, is almost unchanged.

Four pooling stages take a 512×512 map down to 32×32. The later layers, looking at 3×3 windows on that small map, are effectively seeing large regions of the lung, which is what they need to judge whether an opacity is spread across a lobe.

Follow-up questions to expect

  • "Does pooling make a CNN fully translation-invariant?" — No. It gives tolerance to small shifts within a window. Larger shifts change the output, which is why augmentation still helps.
  • "Max or average pooling?" — Max pooling inside the network, where you want to know whether a feature was present. Global average pooling at the end, to summarise each channel for the classifier.
  • "How many parameters does a pooling layer have?" — None. Window size and stride are hyperparameters.