Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How does L2 regularization help prevent overfitting?


What you need to know

Text
L2 loss = prediction error + λ × (w₁² + w₂² + ... + wₙ²)

Squaring has a useful effect: a weight of 10 costs 100, but two weights of 5 cost only 50. So L2 prefers to spread influence across several features rather than put it all on one. And because the penalty on a small weight is tiny (0.1² = 0.01), L2 has little reason to push a weight all the way to zero — it just keeps shrinking it.

Why smaller weights mean less overfitting

A model with huge weights is very sensitive: a small change in one input causes a large change in the output. That sensitivity is exactly what lets it fit noise. Small weights make the model smoother — similar inputs get similar predictions. In bias–variance terms (see that section), L2 accepts a little extra bias in exchange for a large drop in variance.

The correlated-features problem

This is where L2 shines. Suppose you have both "straight-line distance" and "road distance" — almost the same number. An unregularised model cannot tell which one to credit, so it can give one a huge positive weight and the other a huge negative weight that nearly cancel. Those weights change wildly with each new sample:

Python
import numpy as npfrom sklearn.linear_model import LinearRegression, Ridgerng = np.random.default_rng(0)n = 50km = rng.uniform(1, 8, n)road_km = km + rng.normal(0, 0.01, n)          # road distance: almost the same columnX = np.column_stack([km, road_km])minutes = 15 + 3 * km + rng.normal(0, 3, n)for name, m in [("LinearRegression", LinearRegression()), ("Ridge alpha=1", Ridge(alpha=1.0))]:    coefs = [m.fit(X[s], minutes[s]).coef_.round(1) for s in (slice(0, 25), slice(25, 50))]    print(f"{name:16s} first half {coefs[0]}   second half {coefs[1]}")
Text
LinearRegression first half [-21.8  24.7]   second half [-50.8  53.5]Ridge alpha=1    first half [1.4 1.5]   second half [1.3 1.5]

The true effect is 3 minutes per km. The unregularised model trained on two halves of the same data gives -21.8 and +24.7, then -50.8 and +53.5 — the pair still adds up to roughly 3, but each weight is meaningless and unstable. Ridge shares the effect evenly (about 1.5 each, slightly shrunk) and gives nearly the same answer on both halves. That stability is lower variance.

Where you meet L2 in practice

  • Ridge(alpha=...) for linear regression.
  • LogisticRegression uses L2 by default, with strength set through C (smaller C = stronger penalty).
  • Weight decay in neural networks. In PyTorch, AdamW(weight_decay=0.01) is the common choice. With plain Adam, adding an L2 penalty to the loss is not the same as true weight decay, because Adam rescales each gradient; AdamW was introduced to apply the decay directly to the weights.

Scale first

The penalty treats every weight equally, but a weight's size depends on its feature's units. Standardise features (the Feature Scaling section explains how) inside a Pipeline, or the penalty will hit features unevenly.

A real-life example

A used-car marketplace predicts resale prices from 60 features, many of them overlapping: engine size and horsepower, age and odometer reading, city and state. The unregularised model gives "engine size" a weight of +₹4 lakh per litre and "horsepower" a weight of -₹1,500 per hp — nonsense that sales staff notice immediately. With Ridge, both weights become modest and positive, the model's error on next month's listings drops, and the weights finally match a dealer's intuition.

Follow-up questions to expect

  • "Why does L2 not produce exact zeros?" — Near zero, the squared penalty's slope is almost flat, so there is almost no push to go the last bit to zero. L1's penalty has the same slope everywhere, which is why it can reach zero.
  • "Is L2 the same as weight decay?" — For plain gradient descent, yes, mathematically. For adaptive optimisers like Adam they differ, which is why AdamW applies decay directly to the weights.
  • "How do you choose lambda?" — Try values on a log scale (0.001 to 1,000) with cross-validation, for example with RidgeCV, and pick the best validation score.