Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How can you identify overfitting from training and validation curves?


Validation loss as boosting rounds grow0.4360.3470.3440.3570.4060.476012345100 trees,best was 90600 treesRounds: 10, 50, 100, 200, 400, 600. Training loss keeps falling, to 0.006.
Training loss never stops falling, so only the validation curve shows where learning turned into memorising.

What you need to know

Two kinds of curve

  • Training curve (loss against time). For models that learn step by step: neural networks (epochs) and gradient boosting (number of trees).
  • Learning curve (score against training-set size). For any model: train on 10%, 20%, ... 100% of the data and plot both scores.

Reading the shapes

ShapeWhat you seeMeaningAction
Good fitBoth losses fall and flatten togetherThe model learned the patternKeep it
OverfittingTraining loss keeps falling, validation loss turns upwardLearning noise after a pointStop at the lowest validation loss; regularise
UnderfittingBoth losses stay high and flatModel too weak or features missingAdd capacity or features
Still learningBoth still falling at the endStopped too earlyTrain longer

A curve you can produce yourself

Gradient boosting adds one tree per round, so each round works like an epoch. staged_predict_proba gives predictions after every round.

Python
from sklearn.datasets import make_classificationfrom sklearn.ensemble import GradientBoostingClassifierfrom sklearn.metrics import log_lossfrom sklearn.model_selection import train_test_split# Churn-like data: 2,000 customers, 20 features, some label noiseX, y = make_classification(n_samples=2000, n_features=20, n_informative=5,                           flip_y=0.1, random_state=0)X_tr, X_va, y_tr, y_va = train_test_split(X, y, test_size=0.3, random_state=0)gb = GradientBoostingClassifier(n_estimators=600, learning_rate=0.1, max_depth=4,                                random_state=0).fit(X_tr, y_tr)# Each boosting round adds one tree, like one more epoch of trainingtrain_curve = [log_loss(y_tr, p) for p in gb.staged_predict_proba(X_tr)]val_curve = [log_loss(y_va, p) for p in gb.staged_predict_proba(X_va)]for rounds in [10, 50, 100, 200, 400, 600]:    print(f"{rounds:>3} trees: train loss {train_curve[rounds-1]:.3f} | validation loss {val_curve[rounds-1]:.3f}")best = min(range(len(val_curve)), key=val_curve.__getitem__) + 1print("lowest validation loss at", best, "trees")
Text
 10 trees: train loss 0.385 | validation loss 0.436 50 trees: train loss 0.195 | validation loss 0.347100 trees: train loss 0.128 | validation loss 0.344200 trees: train loss 0.064 | validation loss 0.357400 trees: train loss 0.019 | validation loss 0.406600 trees: train loss 0.006 | validation loss 0.476lowest validation loss at 90 trees

Read it as a story. Up to about 90 trees both losses fall: the model is learning real patterns. After that, training loss keeps dropping to almost zero, but validation loss climbs from 0.344 to 0.476. Those extra 510 trees are memorising the labels we deliberately made noisy (flip_y=0.1 scrambles about 10% of them). Early stopping would keep the 90-tree model. In scikit-learn you can do this automatically with n_iter_no_change and validation_fraction; Keras has an EarlyStopping callback, and plain PyTorch loops usually code the same check by hand.

Learning curves for any model

scikit-learn's learning_curve trains on growing slices of the data. If the validation score keeps rising as data grows and the gap is shrinking, more data will help (overfitting). If both curves have flattened at a poor score, more data will not help (underfitting).

A real-life example

A team fine-tunes an image model to spot damaged parcels from warehouse photos. After 30 epochs, training accuracy is 99.6% and they are pleased. The validation curve tells another story: validation loss was lowest at epoch 8 and rose afterwards. The epoch-30 model misses 11% more damaged parcels than the epoch-8 checkpoint. They switch on early stopping with a patience of 3 epochs, meaning training stops once validation loss has not improved for 3 epochs in a row, and add image augmentation.

Follow-up questions to expect

  • "Why watch loss rather than accuracy?" — Loss changes smoothly and shows overconfidence early; accuracy can stay flat while the model is already becoming overconfident on wrong answers.
  • "What is early-stopping patience?" — The number of epochs to wait for validation loss to improve before stopping, so a single noisy epoch does not stop training too soon.
  • "What if validation loss is lower than training loss?" — It can happen when dropout or augmentation is active only during training, or when the validation set is easier. Check that the splits are fair before celebrating.