Course Content
Machine Learning Foundations
14 sections · 70 lessons
When might regularization hurt model performance?
What you need to know
Regularisation is a cure for high variance (overfitting). Given to a model with a different problem, it does harm. Here are the situations to name.
1. The model is underfitting
If training and validation scores are both poor and close together, the model is too simple already (high bias). Adding a penalty makes it simpler still. The right move is the opposite: less regularisation, more features, or a more flexible model.
2. λ is too large
A strong penalty pushes useful weights towards zero. In the validation curve from the previous lesson, alpha = 1,000 dropped validation R² to about zero — the model was predicting roughly the average for everyone.
3. Features are not scaled
L1 and L2 penalise the size of each weight, but a weight's size depends on its feature's units. A feature with large units (income in rupees) needs only a tiny weight, so it is barely penalised. A feature with small units (a ratio between 0 and 0.6) needs a large weight, so it is penalised heavily. Look what Lasso does to a loan-scoring model:
1import numpy as np2from sklearn.linear_model import Lasso3from sklearn.pipeline import make_pipeline4from sklearn.preprocessing import StandardScaler56rng = np.random.default_rng(0)7n = 5008income = rng.uniform(20_000, 200_000, n) # rupees per month: big numbers9emi_ratio = rng.uniform(0.0, 0.6, n) # share of income already on EMIs: small numbers10score = 0.0001 * income - 50 * emi_ratio + rng.normal(0, 3, n) # both matter11X = np.column_stack([income, emi_ratio])1213raw = Lasso(alpha=2.0).fit(X, score)14scaled = make_pipeline(StandardScaler(), Lasso(alpha=2.0)).fit(X, score)15print(f"unscaled: income={raw.coef_[0]:.4f} emi_ratio={raw.coef_[1]:.2f}")16print(f"scaled: income={scaled[-1].coef_[0]:.2f} emi_ratio={scaled[-1].coef_[1]:.2f}")unscaled: income=0.0001 emi_ratio=0.00scaled: income=3.20 emi_ratio=-6.50Without scaling, Lasso throws away the EMI ratio completely — one of the most important signals for loan risk — simply because it is measured in small numbers. After standardising, both features survive, and EMI ratio correctly has the larger effect. Always put a scaler before a regularised linear model, inside a Pipeline.
4. There is plenty of data
With 10 million rows and 50 features, a linear model has little room to overfit. A penalty mostly adds bias. Cross-validation will usually choose a very small λ — let it.
5. You need the coefficients to mean something
Regularised coefficients are biased towards zero on purpose. If a report needs "each extra year of age adds ₹X to the premium" in true units, heavy regularisation will understate X. L1 may even drop a feature that an actuary knows matters.
6. The wrong tool for the model
For tree ensembles, the main complexity controls are depth, leaf size, learning rate and subsampling. L1/L2 terms on leaf weights exist in libraries like XGBoost and LightGBM, but tuning them first rarely helps as much as tuning depth and learning rate.
A real-life example
A food-delivery company's ETA model is a linear regression on 12 hand-picked features, trained on 30 million deliveries. A new engineer adds strong L2 regularisation "to be safe". Training MAE rises from 4.1 to 5.3 minutes and validation MAE rises by the same amount — both got worse together, the classic underfitting signature. With 30 million rows and 12 features, there was no variance to fix. They remove the penalty and instead add features for rain and restaurant load, which is what actually brings the error down.
Follow-up questions to expect
- "How do you detect that regularisation is hurting?" — Training and validation scores are both low and close, and they improve when you lower lambda.
- "Do you regularise one-hot encoded columns?" — Yes, they are usually included; since they are already 0 or 1, many teams leave them unscaled and scale only the numeric columns with a
ColumnTransformer. - "Should you regularise when you only care about explanation, not prediction?" — Be careful. Regularised coefficients are shrunk, so they understate effects. For explanation, a well-specified unregularised model on enough data is often preferred.