Machine Learning Essentials

Regularization (Ridge, Lasso, Elastic Net)


A team fits an ordinary least-squares model to predict house prices. They have 400 sales and, after one-hot encoding neighbourhoods and adding a few interaction terms, 180 features. The training R² comes back as 0.998. On the held-out test set, R² is −3.4 — meaning its squared error is more than four times that of simply guessing the average price for every house.

They print the coefficients to see what went wrong:

Text
bedrooms            +  4,182,000rooms_total         -  4,177,300bathrooms           +  9,310,500bathrooms_per_room  -  9,308,900garage_area         +    812,400driveway_area       -    811,600...

Nothing here says "a bedroom adds £4.18 million". What the model has found is a set of features that nearly duplicate each other, and it has assigned each pair an enormous positive and an enormous negative coefficient that almost exactly cancel. On the training data the cancellation is precise and the fit is perfect. On a new house where bedrooms and rooms_total stand in a slightly different relationship — as they inevitably will — the cancellation fails by a fraction of a percent, and a fraction of a percent of £4 million is a £30,000 error appearing out of nowhere.

This is what unconstrained least squares does when features are correlated or numerous. It has one instruction — minimise training error — and no reason to prefer a modest, stable solution to a wild one that fits marginally better.

Why lasso zeros coefficients and ridge does notRidge, squared penalty• Shrinks every coefficient toward zero• Never reaches exactly zero• Splits weight across correlated columns• Keeps all 180 features, all smallLasso, absolute penalty• Constraint region has sharp corners• Optimum lands on acorner, so terms vanish• Picks one of acorrelated group arbitrarily• Returns a short, readable feature list
Both charge rent on size; only the absolute-value penalty has corners, and corners are where coefficients become exactly zero.

Large coefficients are the disease

The connection between coefficient size and instability is worth making precise, because it is the entire justification for what follows.

A model's prediction changes by βj\beta_j for each unit change in feature jj. If βj=4,182,000\beta_j = 4{,}182{,}000, then a measurement error of 0.01 in that feature — a rounding difference, a slightly different data entry convention — moves the prediction by £41,820. The model is not robust; it is balanced on a knife edge that only the training data happens to hold steady.

Equally, if you retrain on a slightly different sample, those cancelling pairs will land on completely different values. High-variance behaviour, in the sense of "a model that changes drastically with the training sample", shows up concretely as coefficients that change drastically.

Regularisation is the idea that a model should have to pay for complexity. If a large coefficient does not buy a large improvement in fit, the model should not be allowed to have it.

Charging rent on coefficients

The mechanism is a single added term. Instead of minimising squared error alone, minimise squared error plus a penalty on the size of the coefficients:

J(β)=∑i=1n(yi−y^i)2⏟fit the data+α⋅P(β)⏟stay smallJ(\boldsymbol{\beta}) = \underbrace{\sum_{i=1}^{n}\left(y_i - \hat{y}_i\right)^2}_{\text{fit the data}} + \underbrace{\alpha \cdot P(\boldsymbol{\beta})}_{\text{stay small}}

The two terms pull in opposite directions. The fit term wants coefficients that reduce error. The penalty term wants coefficients near zero. α\alpha sets the exchange rate: how much reduction in training error a unit of coefficient magnitude must buy in order to be worth keeping.

Different choices of PP give different behaviour, and the difference is more interesting than it first appears.

Ridge regression: penalise squared magnitude

J=∑i(yi−y^i)2+α∑j=1pβj2J = \sum_{i}\left(y_i - \hat{y}_i\right)^2 + \alpha \sum_{j=1}^{p} \beta_j^2

Ridge (also called L2 regularisation, or Tikhonov regularisation) has an exact solution, and looking at it explains the name:

βridge=(X⊤X+αI)−1X⊤y\boldsymbol{\beta}_{\text{ridge}} = (X^\top X + \alpha I)^{-1} X^\top \mathbf{y}

Compare with ordinary least squares, which is the same thing with α=0\alpha = 0. The problem with OLS on correlated data is that X⊤XX^\top X is close to singular — nearly impossible to invert, which is exactly what produces the exploding coefficients. Adding αI\alpha I places a small positive value along the diagonal, a "ridge" running down the matrix, which makes it invertible again and numerically stable. Ridge regression was invented to fix a numerical problem, and the statistical benefits came afterwards.

What ridge does to coefficients

Ridge shrinks every coefficient towards zero, but never to exactly zero. As α\alpha increases, all coefficients smoothly contract:

αbedroomsrooms_totalgarage_arealot_sizeTest RMSE
0 (OLS)4,182,000−4,177,300812,40018.2184,000
0.141,300−40,9009,12018.161,200
16,840−6,4101,98017.934,800
102,110−1,64087016.428,300
100640−28031011.231,900
10,00012−480.978,400

The pattern is the trade-off in miniature. Too little penalty and the model is unstable. Too much and every coefficient is crushed towards zero, the model approaches "predict the mean for everything", and error rises again. The minimum is somewhere in the middle and is found empirically.

Notice also what ridge does with the correlated pair. Rather than picking one, it spreads the effect across both: the enormous cancelling pair becomes a modest pair pointing in the same general direction as the true joint effect. That is the right behaviour when both features genuinely carry the same signal.

Lasso: penalise absolute magnitude

J=∑i(yi−y^i)2+α∑j=1p∣βj∣J = \sum_{i}\left(y_i - \hat{y}_i\right)^2 + \alpha \sum_{j=1}^{p} |\beta_j|

A single change — squares become absolute values — and the behaviour changes qualitatively. Lasso does not merely shrink coefficients; it sets some of them to exactly zero, removing those features from the model entirely. It performs feature selection as a side effect of fitting.

Why absolute values produce exact zeros

This surprises people, so it is worth understanding rather than memorising.

Think about the pressure the penalty applies to a coefficient as it approaches zero.

For ridge, the penalty is αβ2\alpha\beta^2 and its derivative is 2αβ2\alpha\beta. As β\beta shrinks towards zero, the pressure to shrink further also shrinks towards zero. At β=0.001\beta = 0.001 the penalty barely notices. So there is always some tiny value where the pressure from the penalty is weaker than the pull from the data, and the coefficient settles there rather than at zero.

For lasso, the penalty is α∣β∣\alpha|\beta| and its derivative is α\alpha for any positive β\beta, however small. The pressure is constant all the way down. At β=0.001\beta = 0.001 the penalty pushes just as hard as at β=100\beta = 100. So if a feature's contribution to reducing error is weaker than α\alpha, nothing stops it being driven to precisely zero and held there.

The geometric version of the same fact: with two coefficients, the penalty defines a region the solution must stay inside. For ridge that region is a circle; for lasso it is a diamond with corners on the axes. The solution is where the expanding ellipses of the loss first touch that region — and a diamond is far more likely to be touched at a corner, where one coordinate is exactly zero, than a circle is to be touched at exactly the point where it crosses an axis.

Text
      RIDGE (circle)                LASSO (diamond)   b2                            b2    |    ,--.                     |     /\    |   /    \    .-- loss        |    /  \      .-- loss    |  |  o---|--'   contours     |   /  o \ ---'   contours    |   \    /                    |   \    /    |    `--'                     |    \  /    +-------------- b1            +-----\/------- b1                                        ^   touches off-axis:              touches at a corner:   both b1, b2 nonzero            b2 = 0 exactly

Ridge versus lasso, in practice

Ridge (L2)Lasso (L1)
Penaltyα∑βj2\alpha\sum\beta_j^2α∑∣βj∣\alpha\sum|\beta_j|
Coefficients driven to zero?No — only close to itYes, exactly
Performs feature selection?NoYes
Closed-form solution?YesNo — needs iterative optimisation
With correlated featuresShares the weight between themPicks one arbitrarily, zeroes the rest
When p > nWorks fineSelects at most n features
Best whenMany features each contributing a littleFew features contributing a lot
InterpretabilityAll features retainedSparse, readable model

The "picks one arbitrarily" row is the practical weakness of lasso and deserves emphasis. Given three features correlated at 0.97 — say three different measures of house size — lasso will typically keep one and zero the other two. Which one it keeps can flip if you resample the data or reorder the columns. If you then present that model as evidence that "living area matters and plot footprint does not", you are reporting an artefact of the optimiser, not a fact about houses.

Elastic Net: both penalties

J=∑i(yi−y^i)2+α(ρ∑j∣βj∣+1−ρ2∑jβj2)J = \sum_i (y_i - \hat{y}_i)^2 + \alpha\left(\rho\sum_j|\beta_j| + \frac{1-\rho}{2}\sum_j\beta_j^2\right)

Here α\alpha still controls overall strength and ρ\rho (scikit-learn calls it l1_ratio) controls the mix: ρ=1\rho = 1 is pure lasso, ρ=0\rho = 0 is pure ridge, and anything in between blends them.

The reason to bother is the correlated-group problem. Elastic net has a grouping effect: correlated features tend to enter or leave the model together, with similar coefficients, rather than one being picked at random. You still get sparsity — irrelevant features go to zero — but among a correlated cluster of genuinely useful features, you keep the cluster.

It also lifts lasso's ceiling. Pure lasso can select at most nn features when there are more features than samples, which is a hard limit that has nothing to do with the problem. Elastic net has no such restriction.

SituationReach forBecause
Many features, most probably relevantRidgeYou want shrinkage, not deletion
Many features, most probably irrelevantLassoAutomatic selection is the point
Features come in correlated groupsElastic NetKeeps or drops groups together
p > n (more features than rows)Ridge or Elastic NetLasso caps out at n features
You must explain the model to a regulatorLassoA 12-feature model is explainable; a 400-feature one is not
You have no ideaElastic Net, tune bothIt contains ridge and lasso as special cases

The step everyone skips: scaling

The penalty ∑βj2\sum \beta_j^2 treats every coefficient identically. But coefficients are in the units of their features, and features are not in comparable units.

Suppose area_m2 ranges over 40–300 and area_hectares is the same information divided by 10,000. To have the same effect on predictions, the coefficient on the hectares version must be 10,000 times larger — so the penalty punishes it 100 million times as hard. The identical feature, measured differently, is nearly eliminated by regularisation for no reason at all.

Therefore: always standardise features before regularising. Not usually. Always.

Python
from sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerfrom sklearn.linear_model import Ridgemodel = make_pipeline(StandardScaler(), Ridge(alpha=10))model.fit(X_train, y_train)

Putting the scaler in a pipeline also means it is fitted on the training fold only during cross-validation, which is the other half of doing this correctly.

One related detail that scikit-learn handles for you: the intercept is never penalised. Shrinking β0\beta_0 towards zero would drag every prediction towards zero rather than towards the mean, which is not shrinkage but sabotage.

Choosing α

There is no formula. Cross-validate over a logarithmic grid — the useful range spans many orders of magnitude, so linear grids waste nearly all their points.

Python
import numpy as npfrom sklearn.linear_model import RidgeCV, LassoCV, ElasticNetCValphas = np.logspace(-4, 4, 100)   # 0.0001 to 10,000ridge = make_pipeline(StandardScaler(),                      RidgeCV(alphas=alphas, cv=5)).fit(X_train, y_train)print("ridge alpha:", ridge[-1].alpha_)lasso = make_pipeline(StandardScaler(),                      LassoCV(alphas=alphas, cv=5, max_iter=50_000)                      ).fit(X_train, y_train)print("lasso alpha:", lasso[-1].alpha_)print("features kept:", (lasso[-1].coef_ != 0).sum(), "of", X_train.shape[1])enet = make_pipeline(StandardScaler(),                     ElasticNetCV(l1_ratio=[0.1, 0.5, 0.7, 0.9, 0.95, 1.0],                                  alphas=alphas, cv=5, max_iter=50_000)                     ).fit(X_train, y_train)print("enet alpha:", enet[-1].alpha_, "l1_ratio:", enet[-1].l1_ratio_)

The dedicated *CV classes are considerably faster than a generic grid search because they exploit warm starts: the solution at one α\alpha is an excellent starting point for the next, so fitting a hundred values costs far less than a hundred separate fits.

The one-standard-error rule

Cross-validation scores are noisy estimates. If α=3\alpha = 3 gives CV RMSE 28,300 ± 1,900 and α=30\alpha = 30 gives 29,100 ± 1,800, those are indistinguishable — but the second is a simpler, more heavily regularised model. The convention is to take the largest α\alpha whose score is within one standard error of the best. You give up a statistically meaningless amount of fit for a genuinely more stable model.

Reading a coefficient path

Plotting how coefficients evolve as α\alpha increases is the most informative single diagnostic here.

Python
import matplotlib.pyplot as pltfrom sklearn.linear_model import lasso_pathX_std = StandardScaler().fit_transform(X_train)alphas_path, coefs, _ = lasso_path(X_std, y_train, alphas=np.logspace(-2, 3, 100))plt.semilogx(alphas_path, coefs.T)plt.gca().invert_xaxis()          # strong penalty on the leftplt.xlabel("alpha"); plt.ylabel("coefficient")

Read it from the right (weak penalty, all features active) towards the left (strong penalty). Features whose coefficients survive longest as the penalty tightens are the ones carrying the most signal. Features that drop to zero almost immediately were contributing nothing. And a pair of lines that cross wildly and swap places is the visual signature of correlated features fighting over the same variance.

Where the analogy misleads

Two clarifications that prevent common errors.

Lasso's feature selection is not a statistical test. A coefficient of zero means "this feature did not earn its keep at this penalty level given the other features present". It does not mean the feature is unrelated to the target. Change α\alpha slightly, or drop a correlated neighbour, and it may reappear with a substantial coefficient.

Regularisation is not only for linear models. The same idea appears everywhere under different names, and recognising it saves a lot of relearning:

ModelThe regularisationParameter
Ridge / LassoPenalty on coefficient sizealpha — larger is stronger
Logistic regression / SVMSame penalty, inverted parameterC — smaller is stronger
Decision treeLimit on depth or leaf sizemax_depth, min_samples_leaf
Random forestAveraging plus feature subsamplingmax_features
Gradient boostingSmall steps, subsampled rows, shallow treeslearning_rate, subsample
Neural networkWeight decay, dropout, early stoppingweight_decay, dropout rate

The inversion in the second row catches people constantly. In LogisticRegression and SVC, C is the inverse of regularisation strength: C=0.01 is heavy regularisation and C=100 is almost none. Setting C=100 because you wanted "more regularisation" does the exact opposite of what you intended.

What this means when you build something

Make regularised regression your default rather than your fallback. Plain LinearRegression is the special case α=0\alpha = 0, and there is rarely a good reason to insist on that particular value when cross-validation can choose a better one — including choosing something very close to zero if that is genuinely best.

The practical recipe is short. Put a StandardScaler and a RidgeCV in a pipeline with a log-spaced alpha grid; that is a strong, stable baseline in three lines. If you need a model someone has to read and defend — a credit decision, a clinical score — switch to LassoCV and accept some accuracy in exchange for a model with twelve features instead of four hundred. If your features arrive in correlated families, use ElasticNetCV and tune the mix.

Then check the coefficients. If any of them still look like £4 million per bedroom, the penalty is not doing its job, and the most likely reason is that someone forgot the scaler.