Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is regularization in Machine Learning?
What you need to know
A model learns by minimising a loss — a number that measures how wrong its predictions are on the training data. Left alone, a flexible model will push that loss as low as possible, even if that means bending itself around random noise. That is overfitting (covered in the Overfitting section).
Regularisation changes what the model is trying to minimise:
total loss = prediction error + λ × penalty(weights)The first part rewards fitting the data. The second part charges a price for large weights. The model now has to balance the two: it will only use a large weight if that weight reduces the error by more than it costs.
λ (lambda) is the regularisation strength. In scikit-learn it is called alpha for Ridge and Lasso. Careful: in LogisticRegression and SVMs the setting is C, which is the inverse of strength — a smaller C means stronger regularisation.
Seeing it work
Here a very flexible model — a degree-12 polynomial — is fitted to just 15 delivery orders where the true relationship is a straight line plus noise:
1import numpy as np2from sklearn.pipeline import make_pipeline3from sklearn.preprocessing import PolynomialFeatures, StandardScaler4from sklearn.linear_model import LinearRegression, Ridge5from sklearn.metrics import mean_absolute_error67rng = np.random.default_rng(1)8x = rng.uniform(0, 10, 15).reshape(-1, 1) # 15 orders: distance in km9y = 15 + 3 * x.ravel() + rng.normal(0, 3, 15) # delivery minutes: a straight line + noise10x_new = rng.uniform(0, 10, 1000).reshape(-1, 1)11y_new = 15 + 3 * x_new.ravel() + rng.normal(0, 3, 1000)1213for name, reg in [("no penalty", LinearRegression()), ("ridge alpha=1", Ridge(alpha=1.0))]:14 model = make_pipeline(PolynomialFeatures(degree=12), StandardScaler(), reg).fit(x, y)15 print(f"{name:13s} train MAE={mean_absolute_error(y, model.predict(x)):.1f} "16 f"new-data MAE={mean_absolute_error(y_new, model.predict(x_new)):.1f}")no penalty train MAE=1.0 new-data MAE=9.9ridge alpha=1 train MAE=1.9 new-data MAE=2.6Without a penalty, the curve wiggles through the 15 points (train error 1.0 minute) and is wildly wrong between them (9.9 minutes on new orders). With a small L2 penalty, the same degree-12 model is forced to keep its weights small, so it cannot wiggle much. Training error gets a little worse; new-data error drops to 2.6 — close to the noise we added. The StandardScaler step matters: penalties compare weights, so features must be on the same scale first.
The regularisation family
| Technique | Where | How it limits complexity |
|---|---|---|
| L2 (Ridge, weight decay) | Linear models, neural nets | Penalises squared weights; shrinks all of them |
| L1 (Lasso) | Linear models | Penalises absolute weights; sets some to exactly zero |
| Elastic Net | Linear models | Mix of L1 and L2 |
| Dropout | Neural nets | Randomly switches off neurons during training |
| Early stopping | Iteratively trained models | Stops before the model starts memorising |
max_depth, min_samples_leaf | Trees | Stops splits on tiny groups of rows |
| Data augmentation | Images, text | More varied examples make memorising harder |
They all say the same thing: "do not become more complicated than the data justifies".
A real-life example
A quick-commerce company predicts how many units of each product a dark store will sell tomorrow. With 300 features (weather, promotions, day of week, lags of past sales) and only 120 days of history for a new store, an unregularised linear model fits the past perfectly and gives absurd forecasts — negative units for some items. The team adds an L2 penalty and tunes alpha with time-based cross-validation. Training error rises slightly, but forecast error on the next two weeks drops by about a third, and the negative forecasts disappear.
Follow-up questions to expect
- "Is regularisation only for linear models?" — No. Neural networks use weight decay and dropout, trees use depth and leaf-size limits, and gradient-boosting libraries such as XGBoost also have L1 and L2 terms on leaf weights.
- "Why not just use a simpler model?" — Sometimes that is the right answer. Regularisation lets you keep a flexible model but control how much of its flexibility it uses, tuned by one number.
- "Is the intercept regularised?" — Usually not. Penalising the intercept would pull predictions towards zero for no reason; scikit-learn's Ridge and Lasso leave it out of the penalty.