Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is the difference between training performance and generalization performance?
What you need to know
Training performance
- Measured on rows the model learned from
- Can reach 100% by memorising
- Useful for spotting underfitting
- Never a promise about new data
Generalisation performance
- Measured on rows the model never saw
- Estimated with validation, cross-validation or a test set
- The number that predicts production
- Only valid if the unseen data looks like production
Generalisation means the model has learned patterns that hold beyond its training examples. You cannot measure it directly — the future has not happened yet — so you estimate it on held-out data.
Reading the two numbers together
| Training | Validation | Diagnosis | Next step |
|---|---|---|---|
| Low | Low, close to training | Underfitting (high bias) | More expressive model, better features, less regularisation |
| High | Much lower | Overfitting (high variance) or leakage | More data, regularisation, simpler model; check for leakage |
| High | High, close | Healthy | Check segments, then test set |
| Low | Higher than training | Unusual — check for a bug | Data mix-up, heavy dropout or augmentation only in training |
The Overfitting and Bias–Variance sections cover these diagnoses in depth. The point here is measurement.
The gap is a clue, not the score
A common misunderstanding is "the smaller the gap, the better the model". It is the validation score you choose on. Random forests, for example, almost always score close to 100% on training data:
1from sklearn.datasets import make_classification2from sklearn.model_selection import cross_validate3from sklearn.ensemble import RandomForestClassifier45X, y = make_classification(n_samples=1500, n_features=20, n_informative=5,6 flip_y=0.1, random_state=0)7for leaf in [1, 20]:8 rf = RandomForestClassifier(min_samples_leaf=leaf, random_state=0)9 s = cross_validate(rf, X, y, cv=5, return_train_score=True)10 print(f"min_samples_leaf={leaf:2d} train={s['train_score'].mean():.2f} "11 f"validation={s['test_score'].mean():.2f} ± {s['test_score'].std():.2f}")min_samples_leaf= 1 train=1.00 validation=0.86 ± 0.01min_samples_leaf=20 train=0.87 validation=0.83 ± 0.02The first forest has a 14-point gap and the second only 4 points — yet the first generalises better (0.86 against 0.83). Restricting the leaves closed the gap by lowering the training score, not by raising the validation score. cross_validate with return_train_score=True gives you both numbers from 5 folds, and the "±" (standard deviation across folds) tells you how much of a difference is just noise.
Why even validation is a little optimistic
Every time you look at the validation score and change something, you fit to it slightly. After 200 experiments, the best validation score is partly luck. That is why a final, untouched test set exists. And even a clean test score assumes production data looks like the test data; if users or seasons change (drift), real performance can fall, so you monitor after launch.
A real-life example
A ride-hailing company trains a surge-pricing demand model on January to June data. Training MAE is 4 rides per zone per hour; validation MAE on a random 20% of the same months is 5. The team is happy — until they test on July, which the model never saw, and MAE is 11. The random validation split had rows from the same days as training, so it measured memory of those days, not generalisation to new ones. They switch to a time-based split (train on January to May, validate on June), which gives an honest MAE of 9, and add weather and holiday features to close the gap.
Follow-up questions to expect
- "Can validation be better than training?" — Occasionally, by chance on small data, or because regularisation like dropout is active only during training. A large, persistent reverse gap usually signals a bug or leakage.
- "How do you get a more reliable generalisation estimate on small data?" — Use k-fold cross-validation and report the mean and standard deviation across folds.
- "What is the generalisation gap?" — Training score minus validation score. It measures how much the model relies on memorising its training rows.