Course Content
Mathematics for Machine Learning
5 sections · 13 lessons
Regularization — Keeping Models Honest
Take twenty points that lie roughly along a gentle curve, with a little measurement noise on each. Fit a straight line and you get a sensible but imperfect fit — training error around 0.9, test error around 1.0. Now fit a degree-15 polynomial to the same twenty points. The training error drops to something like 10−6. The curve passes through essentially every point exactly.
Then evaluate it on new data from the same source, and the test error is in the hundreds. Plot the fitted curve and you see why: between the training points it swings wildly up and down, shooting off to enormous values in the gaps, hooking violently at the edges. Look at the coefficients and they are absurd — numbers like +4.2×106 next to −3.9×106, enormous terms cancelling each other almost exactly to produce a curve that happens to thread through your twenty points.
The model did exactly what you asked. You said "minimise squared error on this data" and it found the parameters that do that. It memorised the noise, because with sixteen free parameters and twenty points there is more than enough freedom to do so, and nothing in the objective said memorising was bad.
The fix is to change the question. Instead of "which parameters fit this data best", ask "which parameters fit this data well while staying small". Adding that second clause to the objective is what regularisation means, and it turns out to be one of the highest-leverage ideas in applied machine learning.
Where the error actually comes from
Before fixing overfitting it helps to know precisely what you are trading against what. For squared error, the expected test error at a point decomposes into exactly three pieces:
Imagine training your model many times on many different datasets drawn from the same source, and collecting all the predictions it makes at one particular input x.
Bias is how far the average of those predictions sits from the truth. High bias means the model is systematically wrong in the same direction every time — it is too simple to represent the real pattern. A straight line fitted to a curve has high bias no matter how much data you give it.
Variance is how much those predictions scatter around their own average. High variance means the model is unstable: change the training data slightly and the prediction changes a lot. The degree-15 polynomial has enormous variance — refit it on a different twenty points and you get a completely different curve.
Noise σ2 is the irreducible part. It is the randomness in y that no function of x could ever predict. This term does not depend on your model at all, and no algorithm, architecture or quantity of data will reduce it. It is the floor.
You cannot beat the noise floor. Everything you can control lives in bias and variance, and pushing one down usually pushes the other up.
| Underfitting | Overfitting | |
|---|---|---|
| Training error | High | Very low |
| Test error | High | High |
| Gap between them | Small | Large |
| Dominant term | Bias | Variance |
| Symptom | Model misses obvious structure | Coefficients enormous; predictions unstable |
| Fix | More capacity, better features, train longer | Regularise, get more data, reduce capacity |
The diagnostic is the gap. A model with 2% training error and 3% test error is not overfitting no matter how complex it is — it is simply good. A model with 0.1% training error and 25% test error is overfitting badly, regardless of size. Always look at both numbers.
Ridge regression: penalise large weights
Add a term to the objective that grows when the weights grow:
The first term is the usual squared error — it wants the predictions to match. The second is the penalty — it wants the weights to be small. The parameter λ≥0 sets the exchange rate between them. At λ=0 you recover ordinary least squares. As λ→∞ the penalty dominates and every weight is driven to zero.
The reason this works is that wild overfitting requires large weights. To make a curve oscillate violently between data points you need huge coefficients that nearly cancel. Making largeness expensive removes the model's ability to do that, without removing any features.
The closed form, and where the extra term comes from
Expand the objective and take the gradient with respect to w. The error term contributes −2XTy+2XTXw, exactly as it would without a penalty. The new term λwTw differentiates to 2λw. Setting the total to zero:
The identity matrix appears because ∇(λwTw)=2λw=2λIw, and factoring the w out on the left requires writing that scalar as a matrix. So the entire effect of the penalty is to add λ to every diagonal entry before inverting.
Why that always has a solution
Ordinary least squares fails when XTX is singular — when features are perfectly collinear, or when you have more features than observations (d>n), which makes the matrix singular automatically. The inverse does not exist and there is no unique answer.
Ridge never has this problem. Look at what λI does to the eigenvalues. If XTX has eigenvalues σ1,…,σd, then XTX+λI has eigenvalues σi+λ. Every eigenvalue is shifted up by λ. Since XTX is positive semi-definite, all its eigenvalues are at least zero, so all the shifted eigenvalues are at least λ>0. None can be zero, so the matrix is invertible, always.
Adding λ to the diagonal lifts every eigenvalue away from zero. That single change turns an unsolvable problem into a well-conditioned one.
Worked example on a nearly-singular problem
Two features that are almost identical, standardised so that
The eigenvalues of that matrix are 1+0.99=1.99 and 1−0.99=0.01, so the condition number is 199 — very unbalanced.
At λ=0, the determinant is 1−0.9801=0.0199 and
Two nearly identical features, and the model assigns one a large positive weight and the other a large negative weight. That is nonsense as an interpretation — it is the model exploiting the tiny difference between the features to fit noise.
At λ=0.1, the matrix becomes [1.10.990.991.1] with determinant 1.21−0.9801=0.2299, giving
| λ | w1 | w2 | ∥w∥22 | Eigenvalues | Condition number |
|---|---|---|---|---|---|
| 0 | 1.497 | -0.503 | 2.495 | 1.99, 0.01 | 199 |
| 0.1 | 0.565 | 0.383 | 0.465 | 2.09, 0.11 | 19 |
| 1.0 | 0.341 | 0.321 | 0.219 | 2.99, 1.01 | 3 |
Three things happen at once as λ grows. The weights shrink towards zero. The absurd +1.5 / −0.5 split becomes a sensible split of similar credit between two similar features. And the conditioning improves dramatically, which means the answer stops being sensitive to tiny perturbations in the data.
Biased, and better for it
Ridge is a biased estimator. Its expected value is not the true weight vector — the shrinkage systematically pulls estimates towards zero. Ordinary least squares, by contrast, is unbiased.
Yet ridge often has lower total error. The reason is exactly the decomposition above: total error is bias squared plus variance plus noise, and ridge trades a small increase in the first for a large decrease in the second. When features are correlated or data is scarce, the variance term dominates, so the trade is overwhelmingly worth making. This is the practical reason "unbiased" is not the same as "good".
Two things that must be right
Do not penalise the intercept. The intercept represents the baseline level of the target. Shrinking it towards zero forces predictions towards zero, which is meaningless unless zero is a special value for your target. Every library handles this for you, but if you implement ridge yourself, exclude the intercept column from the penalty.
Standardise the features first. This is the mistake that costs people the most, and it silently produces a worse model rather than an error. The penalty ∑jwj2 treats all weights identically, but a weight's natural size depends entirely on its feature's units. Suppose one feature is a distance in metres and another the same distance in kilometres. The kilometre version needs a weight 1000 times larger to express the same relationship — and therefore incurs a penalty a million times larger. Ridge will crush it to nothing purely because of the unit choice.
Subtract the mean and divide by the standard deviation of each feature before fitting, so that all weights are measured on a comparable scale and the penalty means the same thing for each.
Lasso: the penalty that produces exact zeros
Swap the squared penalty for a sum of absolute values:
This looks like a minor change. It is not. Ridge shrinks every weight towards zero but essentially never reaches it. Lasso drives some weights to exactly zero, which removes those features from the model entirely. It performs feature selection as a side effect of fitting.
There is no closed-form solution, because ∣w∣ is not differentiable at zero. Solvers use coordinate descent or proximal methods instead. That non-differentiability is not an inconvenience to be worked around — it is the source of the whole behaviour.
Why L1 zeros things out, geometrically
Both penalties can be rewritten as a constraint: minimise the squared error subject to the weights staying inside a region whose size depends on λ.
For L2 the region is a circle (a sphere in higher dimensions). For L1 it is a diamond — a square rotated 45 degrees, with sharp corners sitting exactly on the coordinate axes.
The squared-error contours are ellipses centred on the unconstrained least-squares solution. The constrained optimum is where the smallest ellipse first touches the region. A circle is smooth everywhere, so the touching point is generically somewhere in the interior of a quadrant, with both coordinates non-zero. A diamond has corners, and an expanding ellipse is disproportionately likely to strike a corner — and a corner is a point where one coordinate is exactly zero.
Why L1 zeros things out, analytically
The geometry is suggestive; the algebra is conclusive. In the simplified case where features are uncorrelated and scaled so that XTX=I, both problems have explicit solutions in terms of the unpenalised estimate zj. (The lasso formula below assumes the squared error is halved, 21∥y−Xw∥2+λ∥w∥1, the convention most libraries use; with the unhalved objective above, the threshold is λ/2.)
The lasso formula is called the soft-thresholding operator. It subtracts λ from the magnitude and clamps at zero: anything smaller than λ is not shrunk, it is deleted.
Take three unpenalised estimates z=[3.0,0.4,−1.8] and apply both with λ=0.5:
| zj | Ridge: zj/1.5 | Lasso: sign(zj)max(∣zj∣−0.5,0) |
|---|---|---|
| 3.0 | 2.000 | 2.5 |
| 0.4 | 0.267 | 0.0 |
| -1.8 | -1.200 | -1.3 |
The small coefficient 0.4 is reduced to 0.267 by ridge — smaller, but still there, still contributing to every prediction. Lasso sets it to exactly 0. The feature is gone.
Notice also that ridge shrinks proportionally, so it hits large coefficients hardest in absolute terms (3.0→2.0 loses 1.0), while lasso shrinks everything by the same fixed amount (3.0→2.5 loses 0.5). Different philosophies: ridge distributes shrinkage by size, lasso applies a flat toll and kills anything that cannot pay it.
The caveat with correlated features
Given ten near-identical features, lasso tends to keep one at random and zero the other nine. The chosen one is arbitrary — refit on slightly different data and it may pick a different member of the group. If you are using the selected feature set to draw conclusions about which variables matter, this is a serious problem. Ridge instead spreads weight roughly evenly across the group, which is more stable but selects nothing.
Elastic net: use both
Here λ controls the overall strength and α∈[0,1] the mix: α=1 is pure lasso, α=0 is pure ridge.
The L1 part still produces sparsity. The L2 part supplies the grouping effect: correlated features are encouraged to receive similar weights, so instead of arbitrarily picking one from a group of ten, elastic net tends to keep or drop the group together. That makes the selected set far more reproducible.
| Ridge (L2) | Lasso (L1) | Elastic Net | |
|---|---|---|---|
| Closed form? | Yes | No | No |
| Produces exact zeros? | No | Yes | Yes |
| Works when d>n? | Yes | Yes, but selects at most n features | Yes, no such cap |
| Correlated features | Shares weight; stable | Picks one arbitrarily; unstable | Keeps groups together |
| Hyperparameters | λ | λ | λ and α |
| Reach for it when | All features plausibly matter; multicollinearity | You believe most features are irrelevant | Many correlated features, some irrelevant |
Choosing λ
λ cannot be fitted from the training data, because the training data always prefers λ=0 — that is what minimises training error by definition. It has to be chosen by measuring performance on data the model did not see.
k-fold cross-validation, step by step
- Split the training data into k equal folds (5 or 10 are standard).
- Pick a grid of candidate values, spaced logarithmically: 10−4,10−3,…,102. Linear spacing is wrong here, because the effect of λ is multiplicative.
- For each candidate λ and each fold i: train on the other k−1 folds and evaluate on fold i.
- Average the k validation scores to get one number per λ.
- Pick the λ with the best average score, then refit on all the training data with it.
Plotting validation error against logλ gives a U shape. The left side is overfitting (too little penalty), the right side is underfitting (too much), and the bottom is your answer. If the curve is still falling at the edge of your grid, extend the grid — the minimum is outside it.
A refinement worth knowing is the one standard error rule: rather than taking the exact minimum, take the largest λ whose average error is within one standard error of the minimum. The minimum itself is a noisy estimate, and among statistically indistinguishable options the simpler, more heavily regularised model is the safer bet.
Tracing the coefficients across the whole λ grid gives a regularisation path. For lasso it is particularly informative: you can see exactly at which λ each feature drops out, and the order in which features die is a readable ranking of their usefulness.
The leakage trap
Standardisation must happen inside each fold, using only that fold's training portion. Standardising the whole dataset before splitting leaks information from the validation fold into the training procedure, because the mean and standard deviation you subtracted were computed partly from data you are about to evaluate on.
The consequence is a cross-validation score that is optimistically biased. The model looks better than it is, you pick a λ that is too small, and performance in production is worse than your numbers promised. Pipelines exist precisely to make this impossible to get wrong.
Any transformation fitted on data — scaling, imputation, encoding — must be fitted inside the fold, never before the split.
The Bayesian reading, in two sentences
Both penalties are exactly what you get from maximum a posteriori estimation with a prior on the weights. A Gaussian prior centred at zero contributes −logp(w)∝∥w∥22, which is ridge. A Laplace prior, which is much more sharply peaked at zero, contributes ∝∥w∥1, which is lasso — and its sharp peak is the probabilistic counterpart of the diamond's corners.
This also gives λ a meaning: it is the ratio of the noise variance to the prior variance. A large λ says you strongly believe the weights are small before seeing any data.
Regularisation beyond the penalty term
Anything that restricts a model's ability to fit noise is regularisation, whether or not it appears in the objective function. Early stopping halts training when validation error starts rising, which limits how far the weights can travel from their small initial values — it is approximately equivalent to an L2 penalty. Dropout randomly zeros activations during training so no single unit can be relied upon. Data augmentation expands the effective dataset with transformations that should not change the label. Batch normalisation adds noise through batch statistics. In practice you use several at once.
Mistakes that quietly cost accuracy
| Mistake | What actually happens |
|---|---|
| Not standardising before penalising | The penalty is applied inconsistently across features; features in large units are wrongly crushed |
| Penalising the intercept | Predictions are dragged towards zero for no reason |
| Scaling before the train/test split | Leakage; validation scores are optimistic and λ is chosen too small |
| Searching λ on a linear grid | Wastes nearly all candidates in a narrow range; use a logarithmic grid |
| "Lasso zeroed it, so it is irrelevant" | With correlated features, lasso drops informative variables arbitrarily |
| Regularising an underfitting model | Makes it worse. Check the train/test gap first; a small gap means the problem is bias, not variance |
| Selecting λ on the test set | The test set is now a validation set and no longer estimates generalisation |
Doing it properly in code
1import numpy as np2from sklearn.linear_model import Ridge, Lasso, ElasticNet, RidgeCV, LassoCV3from sklearn.pipeline import make_pipeline4from sklearn.preprocessing import StandardScaler5from sklearn.model_selection import train_test_split, cross_val_score67rng = np.random.default_rng(0)8n, d = 100, 409X = rng.normal(size=(n, d))10w_true = np.zeros(d)11w_true[:5] = [3.0, -2.0, 1.5, 0.4, -1.8] # only 5 of 40 features matter12y = X @ w_true + 0.5 * rng.normal(size=n)1314X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=0)1516# The scaler is INSIDE the pipeline, so cross_val_score refits it on each17# training fold only. Scaling X before this point would leak.18alphas = np.logspace(-4, 2, 40)1920ridge = make_pipeline(StandardScaler(), RidgeCV(alphas=alphas, cv=5))21lasso = make_pipeline(StandardScaler(), LassoCV(alphas=alphas, cv=5,22 max_iter=10000, random_state=0))23enet = make_pipeline(StandardScaler(), ElasticNet(alpha=0.05, l1_ratio=0.5,24 max_iter=10000))2526for name, model in [("ridge", ridge), ("lasso", lasso), ("enet ", enet)]:27 model.fit(X_tr, y_tr)28 coefs = model[-1].coef_29 nonzero = int(np.sum(np.abs(coefs) > 1e-8))30 print(f"{name} test R2 = {model.score(X_te, y_te):.3f} "31 f"non-zero coefficients: {nonzero}/{d}")3233# Manual sweep to see the U-shaped validation curve.34for a in [1e-3, 1e-2, 1e-1, 1.0, 10.0, 100.0]:35 pipe = make_pipeline(StandardScaler(), Ridge(alpha=a))36 score = cross_val_score(pipe, X_tr, y_tr, cv=5, scoring="r2").mean()37 print(f"alpha={a:8.3f} cv R2 = {score:.4f}")Run it and the pattern is clear. Ridge keeps all 40 coefficients but shrinks the 35 irrelevant ones to small values. Lasso zeros most of the irrelevant ones: on this seed 17 of 40 survive, including all 5 real features, because cross-validation picks the penalty that predicts best, not the one that is sparsest. And the manual sweep traces the U shape (upside down, since higher R2 is better): it peaks around α=1, loses a little on the left by overfitting and a lot on the right by underfitting.
Applying this to a model you are building
Start by measuring the gap between training and validation error, because that number tells you which problem you have. A large gap is a variance problem and regularisation is the right tool. A small gap with both errors high is a bias problem, and adding a penalty will only make things worse — you need more capacity or better features instead.
Once you know it is variance, default to ridge. It is stable, has a closed form, needs one hyperparameter, and rarely does harm. Reach for lasso when you have far more features than you believe are useful and you want the model to tell you which ones. Reach for elastic net when those features come in correlated groups — which, in most real datasets, they do.
Always put the scaler and the model in a pipeline and cross-validate the pipeline as a unit. This is not a stylistic preference. It is the only way to guarantee that the standardisation statistics are computed without touching validation data, and leakage of that kind is invisible: it produces better-looking numbers and a worse model, which is the most dangerous combination there is.
Finally, remember what λ actually encodes. It is a statement of how much you trust your data versus how much you believe the underlying relationship is simple. With ten million clean rows you can afford a small λ and let the data speak. With two hundred noisy rows and forty features, a large λ is not a hedge — it is the honest expression of how little those two hundred rows can really tell you.