Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is the difference between valid and same padding in a CNN?


A 3×3 filter on a 5×5 inputxxxxxxxxxValid: 3 positions per row, output 3×3. Same: pad 1, output 5×5.
Without padding a corner pixel is seen by one window position and a centre pixel by nine.

What you need to know

The output-size formula

For input size n, filter size f, padding p on each side and stride s:

Text
output = floor((n + 2p − f) / s) + 1
  • Valid (p = 0), n = 5, f = 3, s = 1: (5 − 3) / 1 + 1 = 3.
  • Same (p = 1), same input: (5 + 2 − 3) / 1 + 1 = 5.
  • For any odd filter size with stride 1, same padding needs p = (f − 1) / 2: 1 for 3×3, 2 for 5×5.

Why shrinking matters

Each valid 3×3 layer removes 2 pixels from each dimension. Ten such layers turn 224×224 into 204×204. That is manageable, but a 32×32 image would be down to 12×12 after ten layers, leaving little room for pooling. Same padding lets architecture designers control the size explicitly with pooling or stride, rather than losing a little every layer.

Why edges matter

With valid padding, a corner pixel is covered by only one 3×3 window position, while a central pixel is covered by nine. So the network pays less attention to anything at the border. Same padding gives border pixels more windows, at the cost of mixing in artificial zeros.

In PyTorch and Keras

Python
import torch.nn as nnnn.Conv2d(16, 32, kernel_size=3, padding=0)       # validnn.Conv2d(16, 32, kernel_size=3, padding=1)       # same, for stride 1nn.Conv2d(16, 32, kernel_size=3, padding="same")  # PyTorch computes it; stride 1 only

Keras uses padding="valid" (its default) and padding="same". Other padding modes, such as padding_mode="reflect", fill the border with mirrored pixels instead of zeros, which can reduce edge artefacts in image-to-image models.

When to choose valid

  • When you want the output to depend only on real pixels, for example the original U-Net used unpadded convolutions and cropped feature maps to match.
  • When you are deliberately shrinking the map and do not mind losing borders.

A real-life example

A team triaging chest X-rays notices their model misses some findings near the top of the lungs, where the lung meets the edge of the cropped image. Their CNN uses valid padding throughout. After eight 3×3 layers with no padding, features computed from the outermost 8 pixels of each border had far fewer filter positions contributing, and some edge detail simply dropped out as the map shrank.

Switching to same padding keeps the full 512×512 through each block until pooling. Combined with a slightly looser crop, so the lung apex is no longer right at the edge, missed findings near the border fall noticeably on their validation set.

Follow-up questions to expect

  • "Does same padding keep the size with stride 2?" — No. With stride 2 the output is about half the input. "Same" in Keras then means output = ceil(input / stride). PyTorch's padding="same" string is only allowed with stride 1.
  • "Why are filter sizes usually odd?" — An odd filter has a centre pixel, so padding can be added equally on both sides. An even filter needs lopsided padding.
  • "What is 'full' padding?" — Padding f − 1 on each side, so the output grows. It is rare in CNN classifiers.