Course Content
Machine Learning Foundations
14 sections · 70 lessons
How can you identify overfitting from training and validation curves?
What you need to know
Two kinds of curve
- Training curve (loss against time). For models that learn step by step: neural networks (epochs) and gradient boosting (number of trees).
- Learning curve (score against training-set size). For any model: train on 10%, 20%, ... 100% of the data and plot both scores.
Reading the shapes
| Shape | What you see | Meaning | Action |
|---|---|---|---|
| Good fit | Both losses fall and flatten together | The model learned the pattern | Keep it |
| Overfitting | Training loss keeps falling, validation loss turns upward | Learning noise after a point | Stop at the lowest validation loss; regularise |
| Underfitting | Both losses stay high and flat | Model too weak or features missing | Add capacity or features |
| Still learning | Both still falling at the end | Stopped too early | Train longer |
A curve you can produce yourself
Gradient boosting adds one tree per round, so each round works like an epoch. staged_predict_proba gives predictions after every round.
1from sklearn.datasets import make_classification2from sklearn.ensemble import GradientBoostingClassifier3from sklearn.metrics import log_loss4from sklearn.model_selection import train_test_split56# Churn-like data: 2,000 customers, 20 features, some label noise7X, y = make_classification(n_samples=2000, n_features=20, n_informative=5,8 flip_y=0.1, random_state=0)9X_tr, X_va, y_tr, y_va = train_test_split(X, y, test_size=0.3, random_state=0)1011gb = GradientBoostingClassifier(n_estimators=600, learning_rate=0.1, max_depth=4,12 random_state=0).fit(X_tr, y_tr)1314# Each boosting round adds one tree, like one more epoch of training15train_curve = [log_loss(y_tr, p) for p in gb.staged_predict_proba(X_tr)]16val_curve = [log_loss(y_va, p) for p in gb.staged_predict_proba(X_va)]17for rounds in [10, 50, 100, 200, 400, 600]:18 print(f"{rounds:>3} trees: train loss {train_curve[rounds-1]:.3f} | validation loss {val_curve[rounds-1]:.3f}")19best = min(range(len(val_curve)), key=val_curve.__getitem__) + 120print("lowest validation loss at", best, "trees") 10 trees: train loss 0.385 | validation loss 0.436 50 trees: train loss 0.195 | validation loss 0.347100 trees: train loss 0.128 | validation loss 0.344200 trees: train loss 0.064 | validation loss 0.357400 trees: train loss 0.019 | validation loss 0.406600 trees: train loss 0.006 | validation loss 0.476lowest validation loss at 90 treesRead it as a story. Up to about 90 trees both losses fall: the model is learning real patterns. After that, training loss keeps dropping to almost zero, but validation loss climbs from 0.344 to 0.476. Those extra 510 trees are memorising the labels we deliberately made noisy (flip_y=0.1 scrambles about 10% of them). Early stopping would keep the 90-tree model. In scikit-learn you can do this automatically with n_iter_no_change and validation_fraction; Keras has an EarlyStopping callback, and plain PyTorch loops usually code the same check by hand.
Learning curves for any model
scikit-learn's learning_curve trains on growing slices of the data. If the validation score keeps rising as data grows and the gap is shrinking, more data will help (overfitting). If both curves have flattened at a poor score, more data will not help (underfitting).
A real-life example
A team fine-tunes an image model to spot damaged parcels from warehouse photos. After 30 epochs, training accuracy is 99.6% and they are pleased. The validation curve tells another story: validation loss was lowest at epoch 8 and rose afterwards. The epoch-30 model misses 11% more damaged parcels than the epoch-8 checkpoint. They switch on early stopping with a patience of 3 epochs, meaning training stops once validation loss has not improved for 3 epochs in a row, and add image augmentation.
Follow-up questions to expect
- "Why watch loss rather than accuracy?" — Loss changes smoothly and shows overconfidence early; accuracy can stay flat while the model is already becoming overconfident on wrong answers.
- "What is early-stopping patience?" — The number of epochs to wait for validation loss to improve before stopping, so a single noisy epoch does not stop training too soon.
- "What if validation loss is lower than training loss?" — It can happen when dropout or augmentation is active only during training, or when the validation set is easier. Check that the splits are fair before celebrating.