Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What are the risks of excessive hyperparameter tuning?


Tuning a tree on coin-flip labels0.5670.5010.487012best CV scoreof 200 triessame model,fresh datanested CV estimateThe labels are random; the only honest answer is about 0.5.
Pick the best of 200 noisy scores and luck looks like skill — nested CV or an untouched test set removes it.

What you need to know

Every time you check a validation score and make a choice, the validation data leaks a little information into your model. Do it a few times and the effect is tiny. Do it hundreds of times and the "best" configuration is partly the one that got lucky on those particular validation rows.

Tuning can find signal in pure noise

Here the labels are coin flips — there is nothing to learn. A random search tries 200 decision-tree configurations:

Python
import numpy as npfrom scipy.stats import randintfrom sklearn.tree import DecisionTreeClassifierfrom sklearn.model_selection import RandomizedSearchCV, cross_val_scorerng = np.random.default_rng(0)X = rng.normal(size=(300, 10))y = rng.integers(0, 2, 300)                 # labels are coin flips: nothing to learnX_test = rng.normal(size=(5000, 10))y_test = rng.integers(0, 2, 5000)space = {"max_depth": randint(1, 20), "min_samples_leaf": randint(1, 50),         "max_features": [0.3, 0.6, 1.0]}search = RandomizedSearchCV(DecisionTreeClassifier(random_state=0), space,                            n_iter=200, cv=5, random_state=0).fit(X, y)print(f"best CV accuracy after 200 tries: {search.best_score_:.3f}")print(f"same model on fresh data:         {search.score(X_test, y_test):.3f}")nested = cross_val_score(search, X, y, cv=5)   # the search runs inside each outer foldprint(f"nested CV estimate:               {nested.mean():.3f}")
Text
best CV accuracy after 200 tries: 0.567same model on fresh data:         0.501nested CV estimate:               0.487

The search reports 56.7% — apparently better than chance — on data with no signal at all. On fresh data the same model scores 50.1%, a coin flip. The 6.7-point "improvement" was pure selection luck.

Nested cross-validation fixes the estimate. An outer loop splits the data; inside each outer training part, the whole search runs with its own inner CV; the chosen model is then scored on the outer fold, which the search never saw. Its estimate here, 48.7%, is honest. In scikit-learn, nesting is simply cross_val_score(search, X, y).

The other risks

  • Chasing noise — in the grid search lesson, the top settings differed by 0.01 while the fold-to-fold standard deviation was 0.02. Choosing between them is a coin toss.
  • Diminishing returns — the first 20 trials usually find most of the gain. Trial 500 might add 0.1%.
  • Opportunity cost — a week of tuning while a leakage bug or a missing feature sits unfixed.
  • Compute cost — for large models, tuning can cost more than training the final model many times over.
  • Fragility — a very specific configuration, such as a learning rate tuned to four decimal places, can be tuned to quirks of this month's data and degrade faster under drift.

Sensible practice

  1. Fix the metric and the budget before you start — for example 50 trials.
  2. Use cross-validation, and look at the standard deviation, not just the mean.
  3. Apply the one-standard-error rule — among settings whose score is within one standard deviation of the best, pick the simplest (shallowest tree, strongest regularisation).
  4. Keep a test set that tuning never touches, and use it once. Use nested CV when data is small.
  5. Log every trial, so the number of configurations tried is visible to reviewers.

A real-life example

A Kaggle-style internal competition at a bank ranks teams on a public validation set of 5,000 loans. The winning team submitted over 300 versions and tops the board by 0.4 points of ROC-AUC. On the private test set of 50,000 later loans, they fall to fifth place, and a team that submitted only six carefully cross-validated versions wins. The bank's model-risk team now requires every production model to report how many configurations were tried and to show a test score from data no one tuned on.

Follow-up questions to expect

  • "What is nested cross-validation?" — An outer CV loop estimates performance, and an inner CV loop inside each outer training fold chooses hyperparameters. It gives an unbiased estimate of the whole tuning procedure.
  • "If nested CV picks different hyperparameters in each fold, which do you deploy?" — Nested CV estimates performance of the procedure. For the final model, run the search once on all the training data and deploy its choice.
  • "How do you know when to stop tuning?" — When the best score stops improving by more than the fold-to-fold standard deviation over a set number of trials, or when the planned budget is spent.