Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

Why is hyperparameter tuning necessary?


What you need to know

Most hyperparameters control complexity — how flexible the model is allowed to be. Too little and it underfits; too much and it overfits (see the Bias–Variance section). The best setting depends on the data, so no default can be right for every problem.

One hyperparameter, very different results

Python
from sklearn.datasets import make_classificationfrom sklearn.model_selection import cross_val_scorefrom sklearn.tree import DecisionTreeClassifierX, y = make_classification(n_samples=1000, n_features=20, n_informative=6,                           flip_y=0.1, random_state=0)for depth in [1, 3, 5, 8, None]:    scores = cross_val_score(DecisionTreeClassifier(max_depth=depth, random_state=0), X, y, cv=5)    print(f"max_depth={str(depth):4s}  CV accuracy={scores.mean():.3f}")
Text
max_depth=1     CV accuracy=0.743max_depth=3     CV accuracy=0.811max_depth=5     CV accuracy=0.818max_depth=8     CV accuracy=0.832max_depth=None  CV accuracy=0.799

The scikit-learn default for a decision tree is max_depth=None (grow until the leaves are pure). On this data that overfits: 0.799. A depth of 1 underfits: 0.743. Depth 8 is best at 0.832. Same algorithm, same data, a nine-point spread — just from one setting. cross_val_score averages five train/validation splits, so the comparison does not depend on one lucky split.

What tuning can and cannot do

  • It can move a model to the right complexity for its data, and fix defaults that are badly wrong for your problem (a learning rate that makes a neural net diverge, a k that is too small).
  • It cannot create information the features do not contain. If the features cannot separate fraud from genuine payments, no setting will.
  • Typical gains for a sensible model with reasonable defaults are modest. Gradient-boosting libraries in particular have good defaults; tuning often adds a point or two of ROC-AUC. Fixing a data bug or adding a strong feature can add ten.

Where tuning fits in a project

  1. Clean data and a baseline — a simple model with default settings.
  2. Features — the biggest gains usually come here.
  3. Choose the model family — compare a few with defaults.
  4. Tune the chosen model with cross-validation.
  5. Final check — one score on the untouched test set.

Tuning before step 2 wastes compute on a model whose ceiling is set by weak features.

A real-life example

An e-commerce company's product-recommendation model uses gradient boosting with default settings and scores precision@10 of 0.21. A junior engineer spends two weeks on a large search and reaches 0.22. A senior colleague looks at the data instead, adds "categories this user browsed in the last 7 days" as a feature, and — with default hyperparameters again — gets 0.27. A short tuning pass on top brings it to 0.28. The lesson the team writes down: tune last, and budget time in proportion to the likely gain.

Follow-up questions to expect

  • "Which hyperparameters would you tune first for gradient boosting?" — The learning rate together with the number of trees (often using early stopping), then tree size (max_leaf_nodes or max_depth) and min_samples_leaf, then regularisation and subsampling.
  • "How do you know tuning is done?" — When the best few settings differ by less than the fold-to-fold standard deviation, further search is chasing noise.
  • "Do deep learning models need more tuning?" — Usually yes, especially the learning rate, which can decide whether training converges at all.