Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

When might regularization hurt model performance?


How an unscaled feature gets deleted by LassoEMI ratio runsfrom 0 to 0.6So it needsa largeweight, near -50L1 charges byweight size,not by valueLasso setsEMI ratio toexactly zeroStandardisefirst: bothfeatures surviveIncome in rupees needs a tiny weight, so it is barely penalised.
A penalty on weight size punishes features measured in small units, so scaling must come before regularisation.

What you need to know

Regularisation is a cure for high variance (overfitting). Given to a model with a different problem, it does harm. Here are the situations to name.

1. The model is underfitting

If training and validation scores are both poor and close together, the model is too simple already (high bias). Adding a penalty makes it simpler still. The right move is the opposite: less regularisation, more features, or a more flexible model.

2. λ is too large

A strong penalty pushes useful weights towards zero. In the validation curve from the previous lesson, alpha = 1,000 dropped validation R² to about zero — the model was predicting roughly the average for everyone.

3. Features are not scaled

L1 and L2 penalise the size of each weight, but a weight's size depends on its feature's units. A feature with large units (income in rupees) needs only a tiny weight, so it is barely penalised. A feature with small units (a ratio between 0 and 0.6) needs a large weight, so it is penalised heavily. Look what Lasso does to a loan-scoring model:

Python
import numpy as npfrom sklearn.linear_model import Lassofrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerrng = np.random.default_rng(0)n = 500income = rng.uniform(20_000, 200_000, n)        # rupees per month: big numbersemi_ratio = rng.uniform(0.0, 0.6, n)            # share of income already on EMIs: small numbersscore = 0.0001 * income - 50 * emi_ratio + rng.normal(0, 3, n)   # both matterX = np.column_stack([income, emi_ratio])raw = Lasso(alpha=2.0).fit(X, score)scaled = make_pipeline(StandardScaler(), Lasso(alpha=2.0)).fit(X, score)print(f"unscaled: income={raw.coef_[0]:.4f}  emi_ratio={raw.coef_[1]:.2f}")print(f"scaled:   income={scaled[-1].coef_[0]:.2f}  emi_ratio={scaled[-1].coef_[1]:.2f}")
Text
unscaled: income=0.0001  emi_ratio=0.00scaled:   income=3.20  emi_ratio=-6.50

Without scaling, Lasso throws away the EMI ratio completely — one of the most important signals for loan risk — simply because it is measured in small numbers. After standardising, both features survive, and EMI ratio correctly has the larger effect. Always put a scaler before a regularised linear model, inside a Pipeline.

4. There is plenty of data

With 10 million rows and 50 features, a linear model has little room to overfit. A penalty mostly adds bias. Cross-validation will usually choose a very small λ — let it.

5. You need the coefficients to mean something

Regularised coefficients are biased towards zero on purpose. If a report needs "each extra year of age adds ₹X to the premium" in true units, heavy regularisation will understate X. L1 may even drop a feature that an actuary knows matters.

6. The wrong tool for the model

For tree ensembles, the main complexity controls are depth, leaf size, learning rate and subsampling. L1/L2 terms on leaf weights exist in libraries like XGBoost and LightGBM, but tuning them first rarely helps as much as tuning depth and learning rate.

A real-life example

A food-delivery company's ETA model is a linear regression on 12 hand-picked features, trained on 30 million deliveries. A new engineer adds strong L2 regularisation "to be safe". Training MAE rises from 4.1 to 5.3 minutes and validation MAE rises by the same amount — both got worse together, the classic underfitting signature. With 30 million rows and 12 features, there was no variance to fix. They remove the penalty and instead add features for rain and restaurant load, which is what actually brings the error down.

Follow-up questions to expect

  • "How do you detect that regularisation is hurting?" — Training and validation scores are both low and close, and they improve when you lower lambda.
  • "Do you regularise one-hot encoded columns?" — Yes, they are usually included; since they are already 0 or 1, many teams leave them unscaled and scale only the numeric columns with a ColumnTransformer.
  • "Should you regularise when you only care about explanation, not prediction?" — Be careful. Regularised coefficients are shrunk, so they understate effects. For explanation, a well-specified unregularised model on enough data is often preferred.