Course Content
Machine Learning Foundations
14 sections · 70 lessons
How does regularization affect model complexity?
What you need to know
Model complexity is how flexible a model is — how many different shapes of relationship it can fit. There are two kinds:
- Structural complexity — the number of features, parameters, layers, or the polynomial degree.
- Effective complexity — how much of that flexibility the model actually uses.
Regularisation changes only the second. A degree-12 polynomial with a strong penalty behaves more like a gentle curve than a wild one, even though it still has 13 coefficients.
Turning the dial
| λ | Weights | Model behaviour | Diagnosis |
|---|---|---|---|
| 0 | As large as needed | Fits every point, including noise | High variance, overfitting |
| Moderate | Smaller, spread out | Smooth, follows the trend | Balanced |
| Very large | Close to zero | Predicts roughly the average | High bias, underfitting |
A validation curve
validation_curve trains the model at several strengths and reports training and validation scores. Here there are 60 rows and 40 features, a setting that invites overfitting:
1from sklearn.datasets import make_regression2from sklearn.linear_model import Ridge3from sklearn.model_selection import validation_curve45X, y = make_regression(n_samples=60, n_features=40, n_informative=10,6 noise=80, random_state=0)7alphas = [0.001, 1, 10, 100, 1000]8train, val = validation_curve(Ridge(), X, y, param_name="alpha",9 param_range=alphas, cv=5, scoring="r2")10for a, t, v in zip(alphas, train.mean(axis=1), val.mean(axis=1)):11 print(f"alpha={a:<7} train R2={t:5.2f} validation R2={v:5.2f}")alpha=0.001 train R2= 0.99 validation R2= 0.53alpha=1 train R2= 0.99 validation R2= 0.81alpha=10 train R2= 0.93 validation R2= 0.65alpha=100 train R2= 0.58 validation R2= 0.28alpha=1000 train R2= 0.12 validation R2=-0.01R² is the share of the target's variation the model explains: 1 is perfect, 0 is no better than predicting the average. Read the two columns together:
- Training R² falls steadily as alpha rises — the penalty removes flexibility.
- Validation R² first rises (0.53 → 0.81), then falls (→ 0.65 → 0.28 → about 0).
The peak at alpha = 1 is the sweet spot. To its left the model overfits (big train–validation gap); to its right it underfits (both scores low). That rise-then-fall shape is the bias–variance trade-off drawn with real numbers. In practice you would test many more values on a log scale, or use RidgeCV / LassoCV, which do this search for you.
Other dials that do the same job
- Tree depth and
min_samples_leafin decision trees. - Dropout rate and weight decay in neural networks.
- Number of boosting rounds with early stopping.
Each lowers effective complexity, and each is tuned the same way: sweep the value, read the validation curve.
A real-life example
A movie-streaming service fits a linear model to predict how many minutes a user will watch this week from 500 features about their past viewing. With no penalty, training R² is 0.92 and next week's R² is 0.40. The team sweeps alpha over 12 values from 0.01 to 10,000. Validation R² peaks at 0.61 around alpha = 30. Beyond 1,000 the model predicts nearly the same number for everyone, and both scores collapse. They ship alpha = 30 and set a reminder to re-sweep after each quarterly retrain, since the best value changes as the data grows.
Follow-up questions to expect
- "Does more data change the best lambda?" — Usually the best lambda falls as data grows, because more data already reduces variance.
- "How is lambda related to C in logistic regression?" — C is the inverse of the regularisation strength, roughly 1/λ. Smaller C means a stronger penalty and a simpler model.
- "Can regularisation reduce training time?" — Sometimes: it can make optimisation better-behaved, and L1's sparse models are faster to predict with. But its main purpose is generalisation, not speed.