Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How does regularization affect model complexity?


Validation R squared as alpha grows from 0.001 to 1,0000.530.810.650.28-0.0101234alpha0.001: overfitsalpha 1: bestalpha 1000:underfitsTraining R squared falls steadily: 0.99, 0.99, 0.93, 0.58, 0.12.
Validation rises then falls as the penalty grows — the peak is where bias and variance balance.

What you need to know

Model complexity is how flexible a model is — how many different shapes of relationship it can fit. There are two kinds:

  • Structural complexity — the number of features, parameters, layers, or the polynomial degree.
  • Effective complexity — how much of that flexibility the model actually uses.

Regularisation changes only the second. A degree-12 polynomial with a strong penalty behaves more like a gentle curve than a wild one, even though it still has 13 coefficients.

Turning the dial

λWeightsModel behaviourDiagnosis
0As large as neededFits every point, including noiseHigh variance, overfitting
ModerateSmaller, spread outSmooth, follows the trendBalanced
Very largeClose to zeroPredicts roughly the averageHigh bias, underfitting

A validation curve

validation_curve trains the model at several strengths and reports training and validation scores. Here there are 60 rows and 40 features, a setting that invites overfitting:

Python
from sklearn.datasets import make_regressionfrom sklearn.linear_model import Ridgefrom sklearn.model_selection import validation_curveX, y = make_regression(n_samples=60, n_features=40, n_informative=10,                       noise=80, random_state=0)alphas = [0.001, 1, 10, 100, 1000]train, val = validation_curve(Ridge(), X, y, param_name="alpha",                              param_range=alphas, cv=5, scoring="r2")for a, t, v in zip(alphas, train.mean(axis=1), val.mean(axis=1)):    print(f"alpha={a:<7} train R2={t:5.2f}  validation R2={v:5.2f}")
Text
alpha=0.001   train R2= 0.99  validation R2= 0.53alpha=1       train R2= 0.99  validation R2= 0.81alpha=10      train R2= 0.93  validation R2= 0.65alpha=100     train R2= 0.58  validation R2= 0.28alpha=1000    train R2= 0.12  validation R2=-0.01

R² is the share of the target's variation the model explains: 1 is perfect, 0 is no better than predicting the average. Read the two columns together:

  • Training R² falls steadily as alpha rises — the penalty removes flexibility.
  • Validation R² first rises (0.53 → 0.81), then falls (→ 0.65 → 0.28 → about 0).

The peak at alpha = 1 is the sweet spot. To its left the model overfits (big train–validation gap); to its right it underfits (both scores low). That rise-then-fall shape is the bias–variance trade-off drawn with real numbers. In practice you would test many more values on a log scale, or use RidgeCV / LassoCV, which do this search for you.

Other dials that do the same job

  • Tree depth and min_samples_leaf in decision trees.
  • Dropout rate and weight decay in neural networks.
  • Number of boosting rounds with early stopping.

Each lowers effective complexity, and each is tuned the same way: sweep the value, read the validation curve.

A real-life example

A movie-streaming service fits a linear model to predict how many minutes a user will watch this week from 500 features about their past viewing. With no penalty, training R² is 0.92 and next week's R² is 0.40. The team sweeps alpha over 12 values from 0.01 to 10,000. Validation R² peaks at 0.61 around alpha = 30. Beyond 1,000 the model predicts nearly the same number for everyone, and both scores collapse. They ship alpha = 30 and set a reminder to re-sweep after each quarterly retrain, since the best value changes as the data grows.

Follow-up questions to expect

  • "Does more data change the best lambda?" — Usually the best lambda falls as data grows, because more data already reduces variance.
  • "How is lambda related to C in logistic regression?" — C is the inverse of the regularisation strength, roughly 1/λ. Smaller C means a stronger penalty and a simpler model.
  • "Can regularisation reduce training time?" — Sometimes: it can make optimisation better-behaved, and L1's sparse models are faster to predict with. But its main purpose is generalisation, not speed.