Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is the purpose of a validation set?


Three sets, three jobsTrain 60% — fits tree splits and weightsValidation 20% — picks depth 2 of sevenRetrain on train plus validationTest 20% — looked at once: 0.853
The test score is honest only because it played no part in choosing the depth.

What you need to know

Three sets, three jobs

  1. Training set — the model learns its parameters here, such as tree splits or regression weights.
  2. Validation set — you compare options here: tree depth, learning rate, which features, which algorithm, which threshold.
  3. Test set — used once, after every decision is made, to report how the final model will perform.

A typical split is 60/20/20 or 70/15/15.

Choosing a tree depth with a validation set

Python
import numpy as npfrom sklearn.tree import DecisionTreeClassifierfrom sklearn.model_selection import train_test_splitrng = np.random.default_rng(3)n = 3000income, emi_ratio, score = rng.normal(60, 20, n), rng.uniform(0, 0.8, n), rng.normal(700, 60, n)risk = -5.5 + 7 * emi_ratio - 0.03 * (score - 700) - 0.03 * (income - 60)default = rng.random(n) < 1 / (1 + np.exp(-risk))X = np.column_stack([income, emi_ratio, score])# 60% train, 20% validation, 20% testX_tmp, X_te, y_tmp, y_te = train_test_split(X, default, test_size=0.2, stratify=default, random_state=0)X_tr, X_va, y_tr, y_va = train_test_split(X_tmp, y_tmp, test_size=0.25, stratify=y_tmp, random_state=0)best_depth, best_val = None, 0for depth in [2, 3, 4, 5, 6, 8, None]:    model = DecisionTreeClassifier(max_depth=depth, random_state=0).fit(X_tr, y_tr)    val = model.score(X_va, y_va)    print(f"max_depth={depth}: validation accuracy {val:.3f}")    if val > best_val:        best_depth, best_val = depth, valfinal = DecisionTreeClassifier(max_depth=best_depth, random_state=0).fit(X_tmp, y_tmp)print("chosen depth:", best_depth, "| test accuracy, used once:", round(final.score(X_te, y_te), 3))
Text
max_depth=2: validation accuracy 0.848max_depth=3: validation accuracy 0.845max_depth=4: validation accuracy 0.842max_depth=5: validation accuracy 0.832max_depth=6: validation accuracy 0.842max_depth=8: validation accuracy 0.830max_depth=None: validation accuracy 0.812chosen depth: 2 | test accuracy, used once: 0.853

We tried seven depths and let the validation set pick. The unlimited tree is worst, as the previous lesson predicted. After choosing depth 2, we retrain on train plus validation (more data helps) and check the test set exactly once. The test score is a number we can report honestly because it played no part in the choice.

The validation score is slightly optimistic

We picked the best of seven validation scores. Some of that "best" is luck: one depth happened to suit those 600 rows. With 7 options the effect is small; with 500 options it can be large. That is why the final number comes from the test set, not the validation set.

Cross-validation as a replacement

With little data, a single validation set is small and noisy. K-fold cross-validation splits the training data into k parts, validates on each part in turn, and averages. scikit-learn's GridSearchCV does this for you; the hyperparameter tuning section covers it.

A real-life example

A food-delivery company trains a neural network to predict ETAs. After each training pass over the data (an epoch), it checks error on the validation set. Validation error falls for 14 epochs, then starts rising while training error keeps falling. Early stopping saves the model from epoch 14. The team also uses the validation set to compare a gradient-boosting model against the network. Only when both decisions are made do they run the test set: the last two weeks of orders, never touched before.

Follow-up questions to expect

  • "Can I retrain on train plus validation after tuning?" — Yes, and it is common. Fix the hyperparameters, retrain on the combined data, then evaluate once on the test set.
  • "What is early stopping?" — Stopping training when validation error stops improving, and keeping the best checkpoint. It is a validation-set decision.
  • "Is the validation set needed if I use cross-validation?" — Cross-validation plays the validation role. You still keep a separate test set for the final estimate.