Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the difference between valid and same padding in a CNN?
What you need to know
The output-size formula
For input size n, filter size f, padding p on each side and stride s:
output = floor((n + 2p − f) / s) + 1- Valid (
p = 0),n = 5,f = 3,s = 1:(5 − 3) / 1 + 1 = 3. - Same (
p = 1), same input:(5 + 2 − 3) / 1 + 1 = 5. - For any odd filter size with stride 1, same padding needs
p = (f − 1) / 2: 1 for 3×3, 2 for 5×5.
Why shrinking matters
Each valid 3×3 layer removes 2 pixels from each dimension. Ten such layers turn 224×224 into 204×204. That is manageable, but a 32×32 image would be down to 12×12 after ten layers, leaving little room for pooling. Same padding lets architecture designers control the size explicitly with pooling or stride, rather than losing a little every layer.
Why edges matter
With valid padding, a corner pixel is covered by only one 3×3 window position, while a central pixel is covered by nine. So the network pays less attention to anything at the border. Same padding gives border pixels more windows, at the cost of mixing in artificial zeros.
In PyTorch and Keras
1import torch.nn as nn2nn.Conv2d(16, 32, kernel_size=3, padding=0) # valid3nn.Conv2d(16, 32, kernel_size=3, padding=1) # same, for stride 14nn.Conv2d(16, 32, kernel_size=3, padding="same") # PyTorch computes it; stride 1 onlyKeras uses padding="valid" (its default) and padding="same". Other padding modes, such as padding_mode="reflect", fill the border with mirrored pixels instead of zeros, which can reduce edge artefacts in image-to-image models.
When to choose valid
- When you want the output to depend only on real pixels, for example the original U-Net used unpadded convolutions and cropped feature maps to match.
- When you are deliberately shrinking the map and do not mind losing borders.
A real-life example
A team triaging chest X-rays notices their model misses some findings near the top of the lungs, where the lung meets the edge of the cropped image. Their CNN uses valid padding throughout. After eight 3×3 layers with no padding, features computed from the outermost 8 pixels of each border had far fewer filter positions contributing, and some edge detail simply dropped out as the map shrank.
Switching to same padding keeps the full 512×512 through each block until pooling. Combined with a slightly looser crop, so the lung apex is no longer right at the edge, missed findings near the border fall noticeably on their validation set.
Follow-up questions to expect
- "Does same padding keep the size with stride 2?" — No. With stride 2 the output is about half the input. "Same" in Keras then means output = ceil(input / stride). PyTorch's
padding="same"string is only allowed with stride 1. - "Why are filter sizes usually odd?" — An odd filter has a centre pixel, so padding can be added equally on both sides. An even filter needs lopsided padding.
- "What is 'full' padding?" — Padding
f − 1on each side, so the output grows. It is rare in CNN classifiers.