Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

How does increasing the number of layers change model capacity and risk?


Gradient norm reaching the first layer0.020.05 to 10.0000000010.05 to 10 exactly0.05 to 1SigmoidReLU + He init2 layers10 layers30 layers
Sigmoid's slope is at most 0.25, so thirty layers multiply the learning signal down to nothing before it reaches the first layer.

What you need to know

What depth adds

Each new layer adds a weight matrix and one more level of "features built from features". That raises capacity, the range of functions the network can represent. Deeper networks can learn more abstract features and often reach better accuracy on large datasets.

The four risks

  • Overfitting — more parameters can memorise a small dataset.
  • Vanishing and exploding gradients — during backpropagation, the gradient is multiplied by each layer's derivative on the way back. If those factors are mostly below 1, the gradient shrinks towards zero in early layers; if above 1, it blows up.
  • Degradation — in 2015, researchers found that a 56-layer plain network had higher training error than a 20-layer one. This is not overfitting; the optimiser simply cannot find good weights. Residual connections (ResNet) fixed this.
  • Cost — each layer adds compute, memory for activations, and latency, because layers run one after another.

Seeing vanishing gradients

This code builds plain networks of different depths and measures the gradient reaching the first layer:

Python
import torch, torch.nn as nndef first_layer_grad(depth, act, he_init=False):    layers = []    for _ in range(depth):        lin = nn.Linear(64, 64)        if he_init:            nn.init.kaiming_normal_(lin.weight, nonlinearity="relu")            nn.init.zeros_(lin.bias)        layers += [lin, act()]    net = nn.Sequential(*layers, nn.Linear(64, 1))    net(torch.randn(32, 64)).pow(2).mean().backward()    return net[0].weight.grad.norm().item()for d in [2, 10, 30]:    print(d, first_layer_grad(d, nn.Sigmoid), first_layer_grad(d, nn.ReLU, he_init=True))

In one run, with sigmoid, the first-layer gradient norm was about 0.02 at 2 layers, about 1e-9 at 10 layers, and exactly 0 at 30 layers: the first layer receives no learning signal at all. With ReLU and He initialisation it stayed between roughly 0.05 and 1 at every depth; the exact value changes from run to run, but it never collapses. Sigmoid's derivative is at most 0.25, so multiplying many of them shrinks the signal fast.

The countermeasures

ProblemFix
Vanishing gradientsReLU/GELU, He initialisation, residual connections
Unstable activationsBatch norm or layer norm
Exploding gradientsGradient clipping, lower learning rate
OverfittingDropout, weight decay, augmentation, more data
DegradationResidual (skip) connections

A real-life example

A student project classifies 10 tomato-leaf diseases from 2,000 photos. They stack 30 plain convolutional layers from scratch. Training loss barely moves for 20 epochs, then drops, but validation accuracy stays at 55% while training accuracy reaches 99%.

Two problems at once: the plain deep stack trains poorly (degradation), and once it does learn, 2,000 photos are too few for its capacity (overfitting).

The fix is not "more layers". They switch to a pretrained ResNet-18, which has residual connections and features already learned from ImageNet, fine-tune only the last block, and add flips and colour jitter. Validation accuracy rises to 91% with fewer trainable parameters.

Follow-up questions to expect

  • "How do residual connections help?" — A block computes x + F(x), so the gradient has a direct path back through the + x. The block only needs to learn a correction, and learning "do nothing" is easy.
  • "Is degradation the same as overfitting?" — No. Overfitting has low training error and high validation error. Degradation has high training error: the deeper model cannot even fit the training set.
  • "How do you choose the depth?" — Start from a proven architecture for the data type, then tune on validation data. I rarely design depth from scratch.