Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

Why are complex models more prone to overfitting?


What you need to know

Capacity can memorise pure noise

The clearest way to see this: give a model labels that are completely random. There is nothing to learn, so an honest model should stay near 50%.

Python
import numpy as npfrom sklearn.tree import DecisionTreeClassifierrng = np.random.default_rng(0)X = rng.normal(size=(1000, 10))       # 1,000 customers, 10 random featuresy = rng.integers(0, 2, 1000)          # random "churn" labels: there is NO patternX_new, y_new = rng.normal(size=(1000, 10)), rng.integers(0, 2, 1000)for depth in [2, 5, 10, None]:    tree = DecisionTreeClassifier(max_depth=depth, random_state=0).fit(X, y)    print(f"max_depth={str(depth):>4}: {tree.get_n_leaves():>3} leaves | "          f"train acc {tree.score(X, y):.2f} | new data acc {tree.score(X_new, y_new):.2f}")
Text
max_depth=   2:   4 leaves | train acc 0.56 | new data acc 0.50max_depth=   5:  24 leaves | train acc 0.60 | new data acc 0.48max_depth=  10:  65 leaves | train acc 0.73 | new data acc 0.50max_depth=None: 225 leaves | train acc 1.00 | new data acc 0.49

As capacity grows, training accuracy climbs from 0.56 to a perfect 1.00, on labels that are pure noise. On new data every version stays at a coin flip. The unlimited tree built 225 leaves, each a tiny box around a few customers, and simply memorised who had which label. That is overfitting in its purest form: the training score measures memory, not learning.

Why data size matters

Capacity is only dangerous relative to data. A tree with 225 leaves fitted to 1,000 rows has about 4 rows per leaf, easy to memorise. The same tree shape on 10 million rows has thousands of rows per leaf, and noise averages out. This is why large neural networks work on huge datasets: the data is big enough to keep their capacity busy with real patterns.

Examples of high-capacity models

  • Decision trees with no depth limit.
  • k-nearest neighbours with k = 1 (it memorises every point).
  • High-degree polynomials.
  • Neural networks with millions of parameters trained on a few thousand examples.
  • Gradient boosting with many deep trees and no early stopping.

The counterweights

More data, regularisation (penalties on large weights, limits on depth), dropout in neural networks, early stopping, ensembling such as random forests, and always an honest validation set.

A real-life example

A start-up trains a deep neural network with 5 million parameters to predict loan default from 4,000 applications. Training accuracy is 99%; validation is 74%. A plain logistic regression with 12 carefully chosen features gets 81% on validation. The network had enough capacity to memorise all 4,000 applicants. The team ships the logistic regression, and revisits the network once they have 400,000 applications.

Follow-up questions to expect

  • "Why do huge LLMs not simply memorise everything?" — They are trained on trillions of tokens, often for about one pass, so the data is vast relative to capacity. They do still memorise some repeated text, which is a known issue.
  • "Is a complex model ever the right choice?" — Yes, when there is enough data and the pattern is genuinely complex, such as images or language, and with regularisation and validation in place.
  • "How do random forests reduce the overfitting of deep trees?" — Each tree overfits differently because it sees a different sample of rows and features; averaging many of them cancels much of that noise.