Course Content
Machine Learning Foundations
14 sections · 70 lessons
What are the risks of excessive hyperparameter tuning?
What you need to know
Every time you check a validation score and make a choice, the validation data leaks a little information into your model. Do it a few times and the effect is tiny. Do it hundreds of times and the "best" configuration is partly the one that got lucky on those particular validation rows.
Tuning can find signal in pure noise
Here the labels are coin flips — there is nothing to learn. A random search tries 200 decision-tree configurations:
1import numpy as np2from scipy.stats import randint3from sklearn.tree import DecisionTreeClassifier4from sklearn.model_selection import RandomizedSearchCV, cross_val_score56rng = np.random.default_rng(0)7X = rng.normal(size=(300, 10))8y = rng.integers(0, 2, 300) # labels are coin flips: nothing to learn9X_test = rng.normal(size=(5000, 10))10y_test = rng.integers(0, 2, 5000)1112space = {"max_depth": randint(1, 20), "min_samples_leaf": randint(1, 50),13 "max_features": [0.3, 0.6, 1.0]}14search = RandomizedSearchCV(DecisionTreeClassifier(random_state=0), space,15 n_iter=200, cv=5, random_state=0).fit(X, y)16print(f"best CV accuracy after 200 tries: {search.best_score_:.3f}")17print(f"same model on fresh data: {search.score(X_test, y_test):.3f}")1819nested = cross_val_score(search, X, y, cv=5) # the search runs inside each outer fold20print(f"nested CV estimate: {nested.mean():.3f}")best CV accuracy after 200 tries: 0.567same model on fresh data: 0.501nested CV estimate: 0.487The search reports 56.7% — apparently better than chance — on data with no signal at all. On fresh data the same model scores 50.1%, a coin flip. The 6.7-point "improvement" was pure selection luck.
Nested cross-validation fixes the estimate. An outer loop splits the data; inside each outer training part, the whole search runs with its own inner CV; the chosen model is then scored on the outer fold, which the search never saw. Its estimate here, 48.7%, is honest. In scikit-learn, nesting is simply cross_val_score(search, X, y).
The other risks
- Chasing noise — in the grid search lesson, the top settings differed by 0.01 while the fold-to-fold standard deviation was 0.02. Choosing between them is a coin toss.
- Diminishing returns — the first 20 trials usually find most of the gain. Trial 500 might add 0.1%.
- Opportunity cost — a week of tuning while a leakage bug or a missing feature sits unfixed.
- Compute cost — for large models, tuning can cost more than training the final model many times over.
- Fragility — a very specific configuration, such as a learning rate tuned to four decimal places, can be tuned to quirks of this month's data and degrade faster under drift.
Sensible practice
- Fix the metric and the budget before you start — for example 50 trials.
- Use cross-validation, and look at the standard deviation, not just the mean.
- Apply the one-standard-error rule — among settings whose score is within one standard deviation of the best, pick the simplest (shallowest tree, strongest regularisation).
- Keep a test set that tuning never touches, and use it once. Use nested CV when data is small.
- Log every trial, so the number of configurations tried is visible to reviewers.
A real-life example
A Kaggle-style internal competition at a bank ranks teams on a public validation set of 5,000 loans. The winning team submitted over 300 versions and tops the board by 0.4 points of ROC-AUC. On the private test set of 50,000 later loans, they fall to fifth place, and a team that submitted only six carefully cross-validated versions wins. The bank's model-risk team now requires every production model to report how many configurations were tried and to show a test score from data no one tuned on.
Follow-up questions to expect
- "What is nested cross-validation?" — An outer CV loop estimates performance, and an inner CV loop inside each outer training fold chooses hyperparameters. It gives an unbiased estimate of the whole tuning procedure.
- "If nested CV picks different hyperparameters in each fold, which do you deploy?" — Nested CV estimates performance of the procedure. For the final model, run the search once on all the training data and deploy its choice.
- "How do you know when to stop tuning?" — When the best score stops improving by more than the fold-to-fold standard deviation over a set number of trials, or when the planned budget is spent.