Course Content
Machine Learning Essentials
6 sections · 16 lessons
Regularization (Ridge, Lasso, Elastic Net)
A team fits an ordinary least-squares model to predict house prices. They have 400 sales and, after one-hot encoding neighbourhoods and adding a few interaction terms, 180 features. The training R² comes back as 0.998. On the held-out test set, R² is −3.4 — meaning its squared error is more than four times that of simply guessing the average price for every house.
They print the coefficients to see what went wrong:
bedrooms + 4,182,000rooms_total - 4,177,300bathrooms + 9,310,500bathrooms_per_room - 9,308,900garage_area + 812,400driveway_area - 811,600...Nothing here says "a bedroom adds £4.18 million". What the model has found is a set of features that nearly duplicate each other, and it has assigned each pair an enormous positive and an enormous negative coefficient that almost exactly cancel. On the training data the cancellation is precise and the fit is perfect. On a new house where bedrooms and rooms_total stand in a slightly different relationship — as they inevitably will — the cancellation fails by a fraction of a percent, and a fraction of a percent of £4 million is a £30,000 error appearing out of nowhere.
This is what unconstrained least squares does when features are correlated or numerous. It has one instruction — minimise training error — and no reason to prefer a modest, stable solution to a wild one that fits marginally better.
Large coefficients are the disease
The connection between coefficient size and instability is worth making precise, because it is the entire justification for what follows.
A model's prediction changes by βj for each unit change in feature j. If βj=4,182,000, then a measurement error of 0.01 in that feature — a rounding difference, a slightly different data entry convention — moves the prediction by £41,820. The model is not robust; it is balanced on a knife edge that only the training data happens to hold steady.
Equally, if you retrain on a slightly different sample, those cancelling pairs will land on completely different values. High-variance behaviour, in the sense of "a model that changes drastically with the training sample", shows up concretely as coefficients that change drastically.
Regularisation is the idea that a model should have to pay for complexity. If a large coefficient does not buy a large improvement in fit, the model should not be allowed to have it.
Charging rent on coefficients
The mechanism is a single added term. Instead of minimising squared error alone, minimise squared error plus a penalty on the size of the coefficients:
The two terms pull in opposite directions. The fit term wants coefficients that reduce error. The penalty term wants coefficients near zero. α sets the exchange rate: how much reduction in training error a unit of coefficient magnitude must buy in order to be worth keeping.
Different choices of P give different behaviour, and the difference is more interesting than it first appears.
Ridge regression: penalise squared magnitude
Ridge (also called L2 regularisation, or Tikhonov regularisation) has an exact solution, and looking at it explains the name:
Compare with ordinary least squares, which is the same thing with α=0. The problem with OLS on correlated data is that X⊤X is close to singular — nearly impossible to invert, which is exactly what produces the exploding coefficients. Adding αI places a small positive value along the diagonal, a "ridge" running down the matrix, which makes it invertible again and numerically stable. Ridge regression was invented to fix a numerical problem, and the statistical benefits came afterwards.
What ridge does to coefficients
Ridge shrinks every coefficient towards zero, but never to exactly zero. As α increases, all coefficients smoothly contract:
| α | bedrooms | rooms_total | garage_area | lot_size | Test RMSE |
|---|---|---|---|---|---|
| 0 (OLS) | 4,182,000 | −4,177,300 | 812,400 | 18.2 | 184,000 |
| 0.1 | 41,300 | −40,900 | 9,120 | 18.1 | 61,200 |
| 1 | 6,840 | −6,410 | 1,980 | 17.9 | 34,800 |
| 10 | 2,110 | −1,640 | 870 | 16.4 | 28,300 |
| 100 | 640 | −280 | 310 | 11.2 | 31,900 |
| 10,000 | 12 | −4 | 8 | 0.9 | 78,400 |
The pattern is the trade-off in miniature. Too little penalty and the model is unstable. Too much and every coefficient is crushed towards zero, the model approaches "predict the mean for everything", and error rises again. The minimum is somewhere in the middle and is found empirically.
Notice also what ridge does with the correlated pair. Rather than picking one, it spreads the effect across both: the enormous cancelling pair becomes a modest pair pointing in the same general direction as the true joint effect. That is the right behaviour when both features genuinely carry the same signal.
Lasso: penalise absolute magnitude
A single change — squares become absolute values — and the behaviour changes qualitatively. Lasso does not merely shrink coefficients; it sets some of them to exactly zero, removing those features from the model entirely. It performs feature selection as a side effect of fitting.
Why absolute values produce exact zeros
This surprises people, so it is worth understanding rather than memorising.
Think about the pressure the penalty applies to a coefficient as it approaches zero.
For ridge, the penalty is αβ2 and its derivative is 2αβ. As β shrinks towards zero, the pressure to shrink further also shrinks towards zero. At β=0.001 the penalty barely notices. So there is always some tiny value where the pressure from the penalty is weaker than the pull from the data, and the coefficient settles there rather than at zero.
For lasso, the penalty is α∣β∣ and its derivative is α for any positive β, however small. The pressure is constant all the way down. At β=0.001 the penalty pushes just as hard as at β=100. So if a feature's contribution to reducing error is weaker than α, nothing stops it being driven to precisely zero and held there.
The geometric version of the same fact: with two coefficients, the penalty defines a region the solution must stay inside. For ridge that region is a circle; for lasso it is a diamond with corners on the axes. The solution is where the expanding ellipses of the loss first touch that region — and a diamond is far more likely to be touched at a corner, where one coordinate is exactly zero, than a circle is to be touched at exactly the point where it crosses an axis.
RIDGE (circle) LASSO (diamond) b2 b2 | ,--. | /\ | / \ .-- loss | / \ .-- loss | | o---|--' contours | / o \ ---' contours | \ / | \ / | `--' | \ / +-------------- b1 +-----\/------- b1 ^ touches off-axis: touches at a corner: both b1, b2 nonzero b2 = 0 exactlyRidge versus lasso, in practice
| Ridge (L2) | Lasso (L1) | |
|---|---|---|
| Penalty | α∑βj2 | α∑∣βj∣ |
| Coefficients driven to zero? | No — only close to it | Yes, exactly |
| Performs feature selection? | No | Yes |
| Closed-form solution? | Yes | No — needs iterative optimisation |
| With correlated features | Shares the weight between them | Picks one arbitrarily, zeroes the rest |
| When p > n | Works fine | Selects at most n features |
| Best when | Many features each contributing a little | Few features contributing a lot |
| Interpretability | All features retained | Sparse, readable model |
The "picks one arbitrarily" row is the practical weakness of lasso and deserves emphasis. Given three features correlated at 0.97 — say three different measures of house size — lasso will typically keep one and zero the other two. Which one it keeps can flip if you resample the data or reorder the columns. If you then present that model as evidence that "living area matters and plot footprint does not", you are reporting an artefact of the optimiser, not a fact about houses.
Elastic Net: both penalties
Here α still controls overall strength and ρ (scikit-learn calls it l1_ratio) controls the mix: ρ=1 is pure lasso, ρ=0 is pure ridge, and anything in between blends them.
The reason to bother is the correlated-group problem. Elastic net has a grouping effect: correlated features tend to enter or leave the model together, with similar coefficients, rather than one being picked at random. You still get sparsity — irrelevant features go to zero — but among a correlated cluster of genuinely useful features, you keep the cluster.
It also lifts lasso's ceiling. Pure lasso can select at most n features when there are more features than samples, which is a hard limit that has nothing to do with the problem. Elastic net has no such restriction.
| Situation | Reach for | Because |
|---|---|---|
| Many features, most probably relevant | Ridge | You want shrinkage, not deletion |
| Many features, most probably irrelevant | Lasso | Automatic selection is the point |
| Features come in correlated groups | Elastic Net | Keeps or drops groups together |
| p > n (more features than rows) | Ridge or Elastic Net | Lasso caps out at n features |
| You must explain the model to a regulator | Lasso | A 12-feature model is explainable; a 400-feature one is not |
| You have no idea | Elastic Net, tune both | It contains ridge and lasso as special cases |
The step everyone skips: scaling
The penalty ∑βj2 treats every coefficient identically. But coefficients are in the units of their features, and features are not in comparable units.
Suppose area_m2 ranges over 40–300 and area_hectares is the same information divided by 10,000. To have the same effect on predictions, the coefficient on the hectares version must be 10,000 times larger — so the penalty punishes it 100 million times as hard. The identical feature, measured differently, is nearly eliminated by regularisation for no reason at all.
Therefore: always standardise features before regularising. Not usually. Always.
1from sklearn.pipeline import make_pipeline2from sklearn.preprocessing import StandardScaler3from sklearn.linear_model import Ridge45model = make_pipeline(StandardScaler(), Ridge(alpha=10))6model.fit(X_train, y_train)Putting the scaler in a pipeline also means it is fitted on the training fold only during cross-validation, which is the other half of doing this correctly.
One related detail that scikit-learn handles for you: the intercept is never penalised. Shrinking β0 towards zero would drag every prediction towards zero rather than towards the mean, which is not shrinkage but sabotage.
Choosing α
There is no formula. Cross-validate over a logarithmic grid — the useful range spans many orders of magnitude, so linear grids waste nearly all their points.
1import numpy as np2from sklearn.linear_model import RidgeCV, LassoCV, ElasticNetCV34alphas = np.logspace(-4, 4, 100) # 0.0001 to 10,00056ridge = make_pipeline(StandardScaler(),7 RidgeCV(alphas=alphas, cv=5)).fit(X_train, y_train)8print("ridge alpha:", ridge[-1].alpha_)910lasso = make_pipeline(StandardScaler(),11 LassoCV(alphas=alphas, cv=5, max_iter=50_000)12 ).fit(X_train, y_train)13print("lasso alpha:", lasso[-1].alpha_)14print("features kept:", (lasso[-1].coef_ != 0).sum(), "of", X_train.shape[1])1516enet = make_pipeline(StandardScaler(),17 ElasticNetCV(l1_ratio=[0.1, 0.5, 0.7, 0.9, 0.95, 1.0],18 alphas=alphas, cv=5, max_iter=50_000)19 ).fit(X_train, y_train)20print("enet alpha:", enet[-1].alpha_, "l1_ratio:", enet[-1].l1_ratio_)The dedicated *CV classes are considerably faster than a generic grid search because they exploit warm starts: the solution at one α is an excellent starting point for the next, so fitting a hundred values costs far less than a hundred separate fits.
The one-standard-error rule
Cross-validation scores are noisy estimates. If α=3 gives CV RMSE 28,300 ± 1,900 and α=30 gives 29,100 ± 1,800, those are indistinguishable — but the second is a simpler, more heavily regularised model. The convention is to take the largest α whose score is within one standard error of the best. You give up a statistically meaningless amount of fit for a genuinely more stable model.
Reading a coefficient path
Plotting how coefficients evolve as α increases is the most informative single diagnostic here.
1import matplotlib.pyplot as plt2from sklearn.linear_model import lasso_path34X_std = StandardScaler().fit_transform(X_train)5alphas_path, coefs, _ = lasso_path(X_std, y_train, alphas=np.logspace(-2, 3, 100))67plt.semilogx(alphas_path, coefs.T)8plt.gca().invert_xaxis() # strong penalty on the left9plt.xlabel("alpha"); plt.ylabel("coefficient")Read it from the right (weak penalty, all features active) towards the left (strong penalty). Features whose coefficients survive longest as the penalty tightens are the ones carrying the most signal. Features that drop to zero almost immediately were contributing nothing. And a pair of lines that cross wildly and swap places is the visual signature of correlated features fighting over the same variance.
Where the analogy misleads
Two clarifications that prevent common errors.
Lasso's feature selection is not a statistical test. A coefficient of zero means "this feature did not earn its keep at this penalty level given the other features present". It does not mean the feature is unrelated to the target. Change α slightly, or drop a correlated neighbour, and it may reappear with a substantial coefficient.
Regularisation is not only for linear models. The same idea appears everywhere under different names, and recognising it saves a lot of relearning:
| Model | The regularisation | Parameter |
|---|---|---|
| Ridge / Lasso | Penalty on coefficient size | alpha — larger is stronger |
| Logistic regression / SVM | Same penalty, inverted parameter | C — smaller is stronger |
| Decision tree | Limit on depth or leaf size | max_depth, min_samples_leaf |
| Random forest | Averaging plus feature subsampling | max_features |
| Gradient boosting | Small steps, subsampled rows, shallow trees | learning_rate, subsample |
| Neural network | Weight decay, dropout, early stopping | weight_decay, dropout rate |
The inversion in the second row catches people constantly. In LogisticRegression and SVC, C is the inverse of regularisation strength: C=0.01 is heavy regularisation and C=100 is almost none. Setting C=100 because you wanted "more regularisation" does the exact opposite of what you intended.
What this means when you build something
Make regularised regression your default rather than your fallback. Plain LinearRegression is the special case α=0, and there is rarely a good reason to insist on that particular value when cross-validation can choose a better one — including choosing something very close to zero if that is genuinely best.
The practical recipe is short. Put a StandardScaler and a RidgeCV in a pipeline with a log-spaced alpha grid; that is a strong, stable baseline in three lines. If you need a model someone has to read and defend — a credit decision, a clinical score — switch to LassoCV and accept some accuracy in exchange for a model with twelve features instead of four hundred. If your features arrive in correlated families, use ElasticNetCV and tune the mix.
Then check the coefficients. If any of them still look like £4 million per bedroom, the penalty is not doing its job, and the most likely reason is that someone forgot the scaler.