Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is regularization in Machine Learning?


What you need to know

A model learns by minimising a loss — a number that measures how wrong its predictions are on the training data. Left alone, a flexible model will push that loss as low as possible, even if that means bending itself around random noise. That is overfitting (covered in the Overfitting section).

Regularisation changes what the model is trying to minimise:

Text
total loss = prediction error + λ × penalty(weights)

The first part rewards fitting the data. The second part charges a price for large weights. The model now has to balance the two: it will only use a large weight if that weight reduces the error by more than it costs.

λ (lambda) is the regularisation strength. In scikit-learn it is called alpha for Ridge and Lasso. Careful: in LogisticRegression and SVMs the setting is C, which is the inverse of strength — a smaller C means stronger regularisation.

Seeing it work

Here a very flexible model — a degree-12 polynomial — is fitted to just 15 delivery orders where the true relationship is a straight line plus noise:

Python
import numpy as npfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import PolynomialFeatures, StandardScalerfrom sklearn.linear_model import LinearRegression, Ridgefrom sklearn.metrics import mean_absolute_errorrng = np.random.default_rng(1)x = rng.uniform(0, 10, 15).reshape(-1, 1)            # 15 orders: distance in kmy = 15 + 3 * x.ravel() + rng.normal(0, 3, 15)        # delivery minutes: a straight line + noisex_new = rng.uniform(0, 10, 1000).reshape(-1, 1)y_new = 15 + 3 * x_new.ravel() + rng.normal(0, 3, 1000)for name, reg in [("no penalty", LinearRegression()), ("ridge alpha=1", Ridge(alpha=1.0))]:    model = make_pipeline(PolynomialFeatures(degree=12), StandardScaler(), reg).fit(x, y)    print(f"{name:13s} train MAE={mean_absolute_error(y, model.predict(x)):.1f}  "          f"new-data MAE={mean_absolute_error(y_new, model.predict(x_new)):.1f}")
Text
no penalty    train MAE=1.0  new-data MAE=9.9ridge alpha=1 train MAE=1.9  new-data MAE=2.6

Without a penalty, the curve wiggles through the 15 points (train error 1.0 minute) and is wildly wrong between them (9.9 minutes on new orders). With a small L2 penalty, the same degree-12 model is forced to keep its weights small, so it cannot wiggle much. Training error gets a little worse; new-data error drops to 2.6 — close to the noise we added. The StandardScaler step matters: penalties compare weights, so features must be on the same scale first.

The regularisation family

TechniqueWhereHow it limits complexity
L2 (Ridge, weight decay)Linear models, neural netsPenalises squared weights; shrinks all of them
L1 (Lasso)Linear modelsPenalises absolute weights; sets some to exactly zero
Elastic NetLinear modelsMix of L1 and L2
DropoutNeural netsRandomly switches off neurons during training
Early stoppingIteratively trained modelsStops before the model starts memorising
max_depth, min_samples_leafTreesStops splits on tiny groups of rows
Data augmentationImages, textMore varied examples make memorising harder

They all say the same thing: "do not become more complicated than the data justifies".

A real-life example

A quick-commerce company predicts how many units of each product a dark store will sell tomorrow. With 300 features (weather, promotions, day of week, lags of past sales) and only 120 days of history for a new store, an unregularised linear model fits the past perfectly and gives absurd forecasts — negative units for some items. The team adds an L2 penalty and tunes alpha with time-based cross-validation. Training error rises slightly, but forecast error on the next two weeks drops by about a third, and the negative forecasts disappear.

Follow-up questions to expect

  • "Is regularisation only for linear models?" — No. Neural networks use weight decay and dropout, trees use depth and leaf-size limits, and gradient-boosting libraries such as XGBoost also have L1 and L2 terms on leaf weights.
  • "Why not just use a simpler model?" — Sometimes that is the right answer. Regularisation lets you keep a flexible model but control how much of its flexibility it uses, tuned by one number.
  • "Is the intercept regularised?" — Usually not. Penalising the intercept would pull predictions towards zero for no reason; scikit-learn's Ridge and Lasso leave it out of the penalty.