Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is pooling, and how does it help filter features?
What you need to know
Max and average pooling on real numbers
A 4×4 feature map, with a 2×2 window and stride 2, gives four windows:
feature map max pool 2×2 average pool 2×21 3 | 2 04 6 | 1 1 6 2 3.5 1.0-----+----- ->0 2 | 9 5 2 9 1.0 6.01 1 | 3 7Top-left window {1, 3, 4, 6} gives max 6 and average 3.5. Bottom-right {9, 5, 3, 7} gives max 9 and average 6. The map goes from 16 values to 4.
How that "filters" features
A feature map value says how strongly a filter matched at that position. Max pooling asks, for each region, "did this feature appear anywhere here, and how strongly?" It keeps the strongest match and throws away the weaker, noisier responses and the precise location. If the feature moves by one pixel inside the window, the output is unchanged, which is the small translation tolerance people mention.
Average pooling keeps a smoother summary, which is better when the overall amount of a feature matters more than its peak.
What happens in the backward pass
Max pooling passes the gradient only to the position that held the maximum; the other three positions in the window get zero. Average pooling splits the gradient equally among all four. Either way, there are no weights to update.
Global average pooling
Global average pooling averages each whole channel to one number. A ResNet's last feature map of 2048×7×7 becomes a vector of 2,048 values. That replaces a huge flatten-plus-dense head, reduces overfitting, and lets the network accept any input size.
1import torch, torch.nn as nn2x = torch.randn(1, 2048, 7, 7)3print(nn.MaxPool2d(2)(torch.randn(1, 64, 56, 56)).shape) # [1, 64, 28, 28]4print(nn.AdaptiveAvgPool2d(1)(x).flatten(1).shape) # [1, 2048]The modern view
Many newer architectures downsample with strided convolutions instead of pooling, so the network learns how to downsample. Max pooling is still common in the early layers of CNNs and remains a good default in small models.
A real-life example
In a chest X-ray model, one filter in an early layer responds to the hazy texture of a lung opacity. On a patient's image, it fires strongly (8.7) at one position and weakly (0.4 to 1.2) around it. After 2×2 max pooling, the pooled map keeps 8.7 for that region. If the same patient is imaged again a few millimetres to one side, the peak moves one pixel but stays inside the same window, so the pooled value, and everything after it, is almost unchanged.
Four pooling stages take a 512×512 map down to 32×32. The later layers, looking at 3×3 windows on that small map, are effectively seeing large regions of the lung, which is what they need to judge whether an opacity is spread across a lobe.
Follow-up questions to expect
- "Does pooling make a CNN fully translation-invariant?" — No. It gives tolerance to small shifts within a window. Larger shifts change the output, which is why augmentation still helps.
- "Max or average pooling?" — Max pooling inside the network, where you want to know whether a feature was present. Global average pooling at the end, to summarise each channel for the classifier.
- "How many parameters does a pooling layer have?" — None. Window size and stride are hyperparameters.