Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What regularizers can we use in neural networks?


What each regulariser acts onRegularisersWeights — L2 via AdamW, L1Activations — dropoutData —augmentation, MixUp, CutMixTargets — label smoothingTime — early stoppingArchitecture —smaller, stochastic depth
Grouping by what they act on shows why stacking three or four works: each removes a different way to memorise.

What you need to know

A regulariser is anything that makes a model generalise better without making it fit the training data better. Most do it by limiting how freely the model can fit noise.

On the weights: L2 and L1

  • L2 regularisation adds λ × sum(w²) to the loss. Its gradient is 2λw, so every step pulls each weight towards zero in proportion to its size. Large weights, which let the model react sharply to tiny input changes, become expensive.
  • Weight decay multiplies each weight by a factor slightly below 1 every step. For plain SGD it is equivalent to L2. For Adam it is not, which is why AdamW applies the decay separately; that is the version to use.
  • L1 regularisation adds λ × sum(|w|). Its pull is the same size for small and large weights, which pushes many weights to exactly zero. Useful for sparse or prunable models, rarer in practice.

It is common to skip weight decay for biases and normalisation parameters, since shrinking them does not reduce overfitting.

On the activations: dropout

Randomly zero a fraction of activations during training, so the network cannot rely on specific neurons working together. Rates of 0.1 in transformers and 0.2 to 0.5 in dense heads are typical. Variants include spatial dropout, which drops whole feature maps in CNNs.

On the data: augmentation and mixing

Augmentation shows label-preserving variations of each example. MixUp blends two images and their labels (for example 70% cat, 30% dog), and CutMix pastes a patch of one image into another. Both stop the model from becoming certain about any single training example.

On the targets: label smoothing

Instead of a hard target of 1 for the correct class and 0 elsewhere, label smoothing with ε = 0.1 over 3 classes uses [0.033, 0.033, 0.933]. The model is never rewarded for pushing its confidence to 100%, which reduces over-confidence and often improves calibration.

On training time: early stopping

Stop when validation loss stops improving, and keep the best checkpoint. Simple, nearly free, and used almost everywhere.

On the architecture

Fewer parameters, CNN weight sharing, bottleneck layers, and stochastic depth, which randomly skips whole residual blocks during training, all limit capacity.

A typical modern setup in PyTorch:

Python
criterion = nn.CrossEntropyLoss(label_smoothing=0.1)optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.05)# plus: augmentation in the DataLoader, nn.Dropout in the head, early stopping

Not a regulariser: gradient clipping

Gradient clipping is often listed here, but it only caps the size of updates to keep training stable. It does not reduce overfitting.

A real-life example

An online grocery app classifies product photos uploaded by 400 partner stores into 150 categories, with about 90 images per category. A pretrained EfficientNet with no regularisation reaches 99% training accuracy and 76% validation accuracy, and often says "100% sure" about wrong answers, such as confusing two brands of atta that look alike.

The team adds four regularisers, one at a time, measuring each: random crops and colour jitter (validation 76% to 82%), AdamW weight decay 0.05 (to 83%), label smoothing 0.1 (84%, and confidence on wrong answers drops sharply), and early stopping at epoch 14 instead of running all 40. Dropout in the head adds nothing measurable, so they leave it out.

Follow-up questions to expect

  • "What is the difference between L2 regularisation and weight decay?" — For SGD they are the same. For Adam, L2 is added to the gradient and then divided by Adam's adaptive scale, so large-gradient weights get almost no decay; weight decay in AdamW is applied directly to the weights.
  • "Why does L1 give sparse weights and L2 does not?" — L1's pull towards zero has constant size, so small weights are pushed all the way to zero. L2's pull shrinks as the weight shrinks, so weights get small but rarely reach exactly zero.
  • "Does label smoothing have downsides?" — It can hurt when you need well-separated feature representations, for example in some knowledge-distillation setups, and it changes what the output probabilities mean.