Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is overfitting in Machine Learning?


What you need to know

Signal and noise

Every dataset has two parts:

  • Signal — the real pattern that will repeat on new data. Longer distance means longer delivery time.
  • Noise — random variation that will not repeat. This one order was slow because the lift was broken.

A good model learns the signal and ignores the noise. An overfit model learns both, because from inside the training data it cannot tell them apart.

Watch it happen

Here we predict delivery time from distance, using only 15 past orders, with three models of increasing flexibility.

Python
import numpy as npfrom sklearn.linear_model import LinearRegressionfrom sklearn.metrics import mean_absolute_errorfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import MinMaxScaler, PolynomialFeaturesrng = np.random.default_rng(0)def make_orders(n):    km = rng.uniform(0.5, 10, n)                          # distance to customer    minutes = 12 + 4 * km - 0.15 * km**2 + rng.normal(0, 3, n)    return km.reshape(-1, 1), minutesX_tr, y_tr = make_orders(15)        # only 15 past orders to learn fromX_te, y_te = make_orders(500)       # 500 new ordersfor degree in [1, 2, 12]:    model = make_pipeline(MinMaxScaler(), PolynomialFeatures(degree), LinearRegression())    model.fit(X_tr, y_tr)    tr = mean_absolute_error(y_tr, model.predict(X_tr))    te = mean_absolute_error(y_te, model.predict(X_te))    print(f"degree {degree:>2}: train MAE {tr:5.2f} min | test MAE {te:6.2f} min")
Text
degree  1: train MAE  1.78 min | test MAE   2.47 mindegree  2: train MAE  1.62 min | test MAE   2.25 mindegree 12: train MAE  1.30 min | test MAE   6.44 min

A degree-12 polynomial is a very bendy curve. It has the lowest training error, 1.30 minutes, because it bends towards each of the 15 training points. But on 500 new orders it is off by 6.44 minutes, nearly three times worse than the simple curve. It learned the noise in those 15 orders. The degree-2 curve matches the true shape and generalises best.

Common causes

  • A model too flexible for the data: deep trees, high-degree polynomials, big networks on small data.
  • Too many features compared with rows, so some features match the labels by chance.
  • Training for too long (too many epochs or boosting rounds).
  • Too little regularisation.
  • Leakage or duplicate rows, which can look like overfitting and should be checked first.

A real-life example

Production scenario. A retail chain builds a store-sales model with 300 features for 45 stores. Training R² is 0.99; next quarter's R² is 0.40. Many features, like "number of sunny days in week 12", matched past sales by coincidence. The team cuts to 25 features chosen with store managers, adds regularisation, and validates on a later quarter. Training R² drops to 0.82, but next-quarter R² rises to 0.76. A lower training score was the sign of a better model.

Follow-up questions to expect

  • "How do you detect overfitting?" — Compare training and validation scores; a large gap is the signal. For iterative models, plot both over training time and look for validation loss rising while training loss keeps falling.
  • "Can a model overfit with lots of data?" — Yes, if it is flexible enough, but more data makes it much harder, because noise averages out and the real pattern dominates.
  • "Is 100% training accuracy always overfitting?" — Not always; on an easy, clean problem it can be fine. It is overfitting only if validation accuracy is clearly lower.