Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

How does L1/L2 regularization help in deep learning models?


What you need to know

The penalties

Text
L2:  loss = data_loss + lambda * sum(w * w)L1:  loss = data_loss + lambda * sum(|w|)

lambda (often called weight_decay or alpha) controls how strong the penalty is. Too small and it does nothing; too large and the model underfits.

What happens to the weights

The gradient tells us what each penalty does at every step:

  • L2 adds 2 * lambda * w to each weight's gradient. A big weight gets a big push towards zero; a small weight gets a small push. Weights shrink in proportion to their size but rarely reach exactly zero. Hence the name weight decay: each step multiplies weights by a factor slightly below 1.
  • L1 adds lambda * sign(w). Every non-zero weight gets the same fixed push towards zero, whatever its size. Small weights reach zero and stay there. The result is sparsity.

Why smaller weights generalise better

A network with large weights can make its output change sharply when an input changes a little. That lets it fit every bump in the training data, including noise. Small weights force a smoother function, which is more likely to be right between training points.

How it is done in PyTorch

For L2, use the optimiser's weight_decay:

Python
import torchdecay, no_decay = [], []for name, p in model.named_parameters():    (no_decay if p.ndim == 1 else decay).append(p)   # biases and norm params are 1-Dopt = torch.optim.AdamW([    {"params": decay, "weight_decay": 0.01},    {"params": no_decay, "weight_decay": 0.0},], lr=1e-3)

Biases and normalisation parameters are usually excluded, because they are few and shrinking them does not reduce overfitting.

Why AdamW? With plain SGD, adding L2 to the loss and applying weight decay are the same thing. With Adam they are not: Adam divides each gradient by a running size estimate, so an L2 term in the loss gets rescaled and weights with large gradients are barely regularised. AdamW applies the decay directly to the weights, separate from the gradient step. It is the default choice for training transformers.

For L1 there is no optimiser flag; add it to the loss yourself:

Python
l1 = sum(p.abs().sum() for p in model.parameters() if p.ndim > 1)   # weight matrices onlyloss = criterion(model(x), y) + 1e-5 * l1

L2 (weight decay)

  • Penalty: sum of squared weights
  • Shrinks all weights smoothly
  • Weights rarely become exactly zero
  • The standard choice in deep learning

L1

  • Penalty: sum of absolute weights
  • Same push on every weight
  • Many weights become exactly zero
  • Used for sparsity and feature selection

A real-life example

A voice-assistant team trains a keyword-spotting model to detect "stop" and "go" from 1,500 recordings by 40 speakers. Without regularisation, it reaches 99% on training speakers and 81% on new speakers. Looking at the first layer, a few weights are 30 to 50 times larger than the rest: the model has latched onto the exact microphone hum of a few speakers.

With AdamW weight_decay=0.05, the largest weights shrink by more than half, and new-speaker accuracy rises to 89%. With weight_decay=0.5 the model underfits: 84% on training and 80% on new speakers. They keep 0.05.

Later, they need to run the model on a cheap microcontroller. They fine-tune with an L1 penalty on the first layer, and 60% of its weights go to zero. Those connections are pruned, making the model smaller and faster with a 1-point accuracy drop.

Follow-up questions to expect

  • "Why does L1 give zeros and L2 does not?" — L1's push is constant, so small weights get pushed all the way to zero. L2's push shrinks as the weight shrinks, so it approaches zero but rarely reaches it.
  • "What is Elastic Net?" — L1 plus L2 together: some sparsity with L2's stability.
  • "Is weight decay the same as L2?" — For SGD, yes. For adaptive optimisers like Adam, no; that is why AdamW decouples the decay from the gradient update.