Course Content
Deep Learning Essentials
13 sections · 61 lessons
How does increasing the number of layers change model capacity and risk?
What you need to know
What depth adds
Each new layer adds a weight matrix and one more level of "features built from features". That raises capacity, the range of functions the network can represent. Deeper networks can learn more abstract features and often reach better accuracy on large datasets.
The four risks
- Overfitting — more parameters can memorise a small dataset.
- Vanishing and exploding gradients — during backpropagation, the gradient is multiplied by each layer's derivative on the way back. If those factors are mostly below 1, the gradient shrinks towards zero in early layers; if above 1, it blows up.
- Degradation — in 2015, researchers found that a 56-layer plain network had higher training error than a 20-layer one. This is not overfitting; the optimiser simply cannot find good weights. Residual connections (ResNet) fixed this.
- Cost — each layer adds compute, memory for activations, and latency, because layers run one after another.
Seeing vanishing gradients
This code builds plain networks of different depths and measures the gradient reaching the first layer:
1import torch, torch.nn as nn23def first_layer_grad(depth, act, he_init=False):4 layers = []5 for _ in range(depth):6 lin = nn.Linear(64, 64)7 if he_init:8 nn.init.kaiming_normal_(lin.weight, nonlinearity="relu")9 nn.init.zeros_(lin.bias)10 layers += [lin, act()]11 net = nn.Sequential(*layers, nn.Linear(64, 1))12 net(torch.randn(32, 64)).pow(2).mean().backward()13 return net[0].weight.grad.norm().item()1415for d in [2, 10, 30]:16 print(d, first_layer_grad(d, nn.Sigmoid), first_layer_grad(d, nn.ReLU, he_init=True))In one run, with sigmoid, the first-layer gradient norm was about 0.02 at 2 layers, about 1e-9 at 10 layers, and exactly 0 at 30 layers: the first layer receives no learning signal at all. With ReLU and He initialisation it stayed between roughly 0.05 and 1 at every depth; the exact value changes from run to run, but it never collapses. Sigmoid's derivative is at most 0.25, so multiplying many of them shrinks the signal fast.
The countermeasures
| Problem | Fix |
|---|---|
| Vanishing gradients | ReLU/GELU, He initialisation, residual connections |
| Unstable activations | Batch norm or layer norm |
| Exploding gradients | Gradient clipping, lower learning rate |
| Overfitting | Dropout, weight decay, augmentation, more data |
| Degradation | Residual (skip) connections |
A real-life example
A student project classifies 10 tomato-leaf diseases from 2,000 photos. They stack 30 plain convolutional layers from scratch. Training loss barely moves for 20 epochs, then drops, but validation accuracy stays at 55% while training accuracy reaches 99%.
Two problems at once: the plain deep stack trains poorly (degradation), and once it does learn, 2,000 photos are too few for its capacity (overfitting).
The fix is not "more layers". They switch to a pretrained ResNet-18, which has residual connections and features already learned from ImageNet, fine-tune only the last block, and add flips and colour jitter. Validation accuracy rises to 91% with fewer trainable parameters.
Follow-up questions to expect
- "How do residual connections help?" — A block computes
x + F(x), so the gradient has a direct path back through the+ x. The block only needs to learn a correction, and learning "do nothing" is easy. - "Is degradation the same as overfitting?" — No. Overfitting has low training error and high validation error. Degradation has high training error: the deeper model cannot even fit the training set.
- "How do you choose the depth?" — Start from a proven architecture for the data type, then tune on validation data. I rarely design depth from scratch.