Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What happens if you tune hyperparameters on the test set?


Selecting the luckiest model on coin-flip labelsTrue accuracy ofany model: 50%200 experimentson one test setKeep the best:62% reportedFreshdata: 51.5%The 12-point gap is pure selection luck.
Picking the best of many scores on the same rows fits your choices to those rows, so the reported number is always too optimistic.

What you need to know

Picking the best of many is itself a kind of learning

Every score has some luck in it. If you try 200 settings and keep the one with the highest test score, you are partly selecting the one that got the luckiest rows. That luck will not repeat on new data.

A demonstration with pure noise

Here the labels are coin flips, so no model can truly beat 50%. We "tune" by trying 200 feature subsets and keeping whichever scores best on a 100-row test set.

Python
import numpy as npfrom sklearn.linear_model import LogisticRegressionrng = np.random.default_rng(0)# Pure noise: 30 random features, coin-flip labels. True accuracy of ANY model is 50%.X = rng.normal(size=(2100, 30))y = rng.integers(0, 2, 2100)X_tr, y_tr = X[:1000], y[:1000]X_te, y_te = X[1000:1100], y[1000:1100]      # the "test set" we keep peeking atX_new, y_new = X[1100:], y[1100:]            # fresh data nobody has touchedbest_score, best_cols = 0, Nonefor trial in range(200):                     # 200 "experiments"    cols = rng.choice(30, size=5, replace=False)    model = LogisticRegression().fit(X_tr[:, cols], y_tr)    score = model.score(X_te[:, cols], y_te)    if score > best_score:        best_score, best_cols = score, colsfinal = LogisticRegression().fit(X_tr[:, best_cols], y_tr)print("best score on the peeked-at test set:", best_score)print("same model on fresh data            :", final.score(X_new[:, best_cols], y_new))
Text
best score on the peeked-at test set: 0.62same model on fresh data            : 0.515

The "best" model reports 62% on the test set it was chosen with. On 1,000 fresh rows it scores 51.5%, which is the coin-flip truth. The 12-point gap is pure selection luck. Real projects have real signal, so the gap is usually smaller, but it is always in the same direction: too optimistic.

Why the damage is worse than it looks

  • The inflated number goes into a report, a slide or a launch decision.
  • The model chosen is not the truly best one, just the luckiest on those rows.
  • You can no longer tell whether the model is good, because you have no untouched data left.

How to stay honest

  • Tune on a validation set or with cross-validation.
  • Look at the test set once, at the end.
  • If you must iterate after seeing the test score, set aside new data (for example, next month's) for a fresh final check.
  • For honest tuning with small data, use nested cross-validation: an inner loop tunes, an outer loop estimates performance.

A real-life example

A Kaggle-style public leaderboard shows the same effect. Teams submit many times and tune to the public leaderboard's portion of the test data. When the private leaderboard (the untouched portion) is revealed, many teams drop dozens of places. This "leaderboard shake-up" is exactly test-set overfitting.

In a company, it looks like this: a churn team runs 60 experiments over three months, each time checking "the test set". The final report says AUC 0.88. Live, the model gives 0.81. Nobody changed the code; the 0.88 was simply the best of 60 lucky draws.

Follow-up questions to expect

  • "How many times can I look at the test set?" — Ideally once. Each extra look that changes your model makes the estimate less trustworthy.
  • "What if I already tuned on the test set?" — Say so, treat the number as optimistic, and get fresh held-out data for a new final evaluation.
  • "What is nested cross-validation?" — Cross-validation inside cross-validation: the inner loop picks hyperparameters on each outer training fold, and the outer folds measure how well the whole tuning procedure generalises.