Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are the layers of a CNN, and which of them have learnable parameters?
What you need to know
What each layer does
- Convolution: learnable filters produce feature maps. The main feature extractor.
- Batch normalisation: normalises each channel over the batch, then applies a learnable scale
γand shiftβ. Stabilises training. - Activation (ReLU): adds non-linearity.
- Pooling (max or average): shrinks the feature map, usually by 2 in each direction.
- Global average pooling or flatten: turns the final feature maps into a vector.
- Fully connected (dense): combines features into class scores.
- Dropout: randomly zeroes activations during training to reduce overfitting.
- Output: logits, turned into probabilities by softmax or sigmoid.
Count the parameters yourself
1import torch.nn as nn23model = nn.Sequential(4 nn.Conv2d(1, 16, 3, padding=1), nn.BatchNorm2d(16), nn.ReLU(), nn.MaxPool2d(2),5 nn.Conv2d(16, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),6 nn.Flatten(), nn.Dropout(0.3), nn.Linear(32 * 16 * 16, 2),7)8for name, p in model.named_parameters():9 print(name, p.numel())10print(sum(p.numel() for p in model.parameters())) # 21218For a 1×64×64 input, the printout lists only these layers:
| Layer | Output shape | Learnable parameters |
|---|---|---|
| Conv2d(1→16, 3×3) | 16×64×64 | 1×16×9 + 16 = 160 |
| BatchNorm2d(16) | 16×64×64 | 16 γ + 16 β = 32 |
| ReLU, MaxPool | 16×32×32 | 0 |
| Conv2d(16→32, 3×3) | 32×32×32 | 16×32×9 + 32 = 4,640 |
| ReLU, MaxPool, Flatten, Dropout | 8,192 | 0 |
| Linear(8192→2) | 2 | 8,192×2 + 2 = 16,386 |
Total: 21,218. Notice that the single dense layer holds 77% of the parameters. That is why modern CNNs replace flatten plus a big dense layer with global average pooling, which reduces each channel to one number: here 32 values into Linear(32, 2), just 66 parameters.
Parameters versus buffers
Batch norm also keeps a running mean and running variance per channel. They are updated during training but not by gradient descent, and they are used at inference. PyTorch stores them as buffers, so they appear in model.state_dict() but not in model.parameters(). Interviewers like this detail.
A real-life example
A team building an X-ray triage model on a small GPU finds their first design runs out of memory. It has five conv blocks, then flattens a 512×16×16 map into a dense layer of 1,024 units. That one dense layer has 512 × 16 × 16 × 1,024, about 134 million weights, far more than all the conv layers together, and it overfits their 12,000 images badly.
Replacing flatten plus dense with global average pooling and a Linear(512, 2) cuts the head from 134 million parameters to 1,026. Training fits in memory, the gap between training and validation accuracy narrows, and the model also accepts X-rays of different sizes, because global pooling works for any height and width.
Follow-up questions to expect
- "How many parameters does a conv layer have?" —
in_channels × out_channels × kernel_height × kernel_width, plusout_channelsbiases. It does not depend on the image size. - "Why is the conv bias often turned off before batch norm?" — Batch norm subtracts the mean, which cancels any constant bias, and then adds its own shift
β. Sobias=Falseon that conv saves parameters without losing anything. - "Does dropout have parameters?" — No. Its drop rate is a hyperparameter you set, not something learned.