Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What are the layers of a CNN, and which of them have learnable parameters?


Parameter count for a small X-ray classifier16×64×6416016×64×643216×32×32032×32×324,6408,1920216,386output shapeparametersConv 1 to 16BatchNormReLU, MaxPoolConv 16 to 32Pool, Flatten, DropoutLinearTotal 21,218; the one dense layer holds 77 percent.
Pooling, ReLU and dropout learn nothing, and the flatten-plus-dense head is where parameters pile up.

What you need to know

What each layer does

  • Convolution: learnable filters produce feature maps. The main feature extractor.
  • Batch normalisation: normalises each channel over the batch, then applies a learnable scale γ and shift β. Stabilises training.
  • Activation (ReLU): adds non-linearity.
  • Pooling (max or average): shrinks the feature map, usually by 2 in each direction.
  • Global average pooling or flatten: turns the final feature maps into a vector.
  • Fully connected (dense): combines features into class scores.
  • Dropout: randomly zeroes activations during training to reduce overfitting.
  • Output: logits, turned into probabilities by softmax or sigmoid.

Count the parameters yourself

Python
import torch.nn as nnmodel = nn.Sequential(    nn.Conv2d(1, 16, 3, padding=1), nn.BatchNorm2d(16), nn.ReLU(), nn.MaxPool2d(2),    nn.Conv2d(16, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),    nn.Flatten(), nn.Dropout(0.3), nn.Linear(32 * 16 * 16, 2),)for name, p in model.named_parameters():    print(name, p.numel())print(sum(p.numel() for p in model.parameters()))   # 21218

For a 1×64×64 input, the printout lists only these layers:

LayerOutput shapeLearnable parameters
Conv2d(1→16, 3×3)16×64×641×16×9 + 16 = 160
BatchNorm2d(16)16×64×6416 γ + 16 β = 32
ReLU, MaxPool16×32×320
Conv2d(16→32, 3×3)32×32×3216×32×9 + 32 = 4,640
ReLU, MaxPool, Flatten, Dropout8,1920
Linear(8192→2)28,192×2 + 2 = 16,386

Total: 21,218. Notice that the single dense layer holds 77% of the parameters. That is why modern CNNs replace flatten plus a big dense layer with global average pooling, which reduces each channel to one number: here 32 values into Linear(32, 2), just 66 parameters.

Parameters versus buffers

Batch norm also keeps a running mean and running variance per channel. They are updated during training but not by gradient descent, and they are used at inference. PyTorch stores them as buffers, so they appear in model.state_dict() but not in model.parameters(). Interviewers like this detail.

A real-life example

A team building an X-ray triage model on a small GPU finds their first design runs out of memory. It has five conv blocks, then flattens a 512×16×16 map into a dense layer of 1,024 units. That one dense layer has 512 × 16 × 16 × 1,024, about 134 million weights, far more than all the conv layers together, and it overfits their 12,000 images badly.

Replacing flatten plus dense with global average pooling and a Linear(512, 2) cuts the head from 134 million parameters to 1,026. Training fits in memory, the gap between training and validation accuracy narrows, and the model also accepts X-rays of different sizes, because global pooling works for any height and width.

Follow-up questions to expect

  • "How many parameters does a conv layer have?" — in_channels × out_channels × kernel_height × kernel_width, plus out_channels biases. It does not depend on the image size.
  • "Why is the conv bias often turned off before batch norm?" — Batch norm subtracts the mean, which cancels any constant bias, and then adds its own shift β. So bias=False on that conv saves parameters without losing anything.
  • "Does dropout have parameters?" — No. Its drop rate is a hyperparameter you set, not something learned.