Course Content
Machine Learning Foundations
14 sections · 70 lessons
What steps do you take when training accuracy is high but test accuracy is low?
What you need to know
"Train 98%, test 72%" has four possible causes. They need different fixes, so identify which one you have.
| Cause | How to recognise it | Fix |
|---|---|---|
| Overfitting | Gap shrinks with regularisation or more data | Reduce variance |
| Split or feature problem | Test has new groups; ID-like features in the model | Group-aware split, drop ID columns |
| Distribution shift | Train and test rows are easy to tell apart | Train on data like the test period; reweight |
| Unlucky split | Gap varies a lot across CV folds | Use k-fold CV, larger test set |
Step 1: check the split
- Time — if the test set is a later period than training, the gap may be drift, not overfitting.
- New groups — if training repeats the same customers many times and the test set holds new customers, the model may have learned who each customer is. Training looks great; new customers look bad. Validate with
GroupKFoldso you see this during development, not at the end. - ID-like features — a user ID, order number or raw timestamp passed in as a number lets a tree memorise individual rows. Drop them.
- Duplicates — near-identical rows on both sides make the test number unreliable. Deduplicate before splitting.
Step 2: check whether the test set looks different
Adversarial validation is a quick, practical check. Label every training row 0 and every test row 1, then train a classifier to tell them apart. If it cannot (AUC near 0.5), train and test look alike. If it can, they differ, and the "overfitting" may really be a shift.
1import numpy as np, pandas as pd2from sklearn.ensemble import HistGradientBoostingClassifier3from sklearn.model_selection import cross_val_score45rng = np.random.default_rng(0)6def orders(n, avg_basket): # two features per order7 return pd.DataFrame({"basket_value": rng.normal(avg_basket, 300, n),8 "items": rng.poisson(4, n)})910train = orders(5000, avg_basket=900) # January to June11test_same = orders(2000, avg_basket=900) # a random slice of the same months12test_sale = orders(2000, avg_basket=1400) # festive-sale week1314for name, test in [("same period", test_same), ("sale week", test_sale)]:15 X = pd.concat([train, test])16 is_test = np.r_[np.zeros(len(train)), np.ones(len(test))] # label = "which set is this row from?"17 auc = cross_val_score(HistGradientBoostingClassifier(), X, is_test, cv=5, scoring="roc_auc").mean()18 print(f"{name:11s} train-vs-test AUC = {auc:.2f}")same period train-vs-test AUC = 0.49sale week train-vs-test AUC = 0.87A test set from the same months is indistinguishable from training (0.49). The festive-sale week is easy to spot (0.87), because basket values are much higher. A model that does badly on sale week may not be overfitting at all; it has simply never seen sale-week behaviour. The fix is to include past sale periods in training, not to add regularisation. The classifier's feature importances also tell you which features shifted.
Step 3: if it really is overfitting, reduce variance
- More training data — the most reliable fix.
- Regularisation — L1/L2, dropout,
min_samples_leaf,max_depth(see the Regularization section). - Simpler model or fewer features — especially when features outnumber the signal.
- Early stopping for boosting and neural networks.
- Ensembling — random forests and bagging average away variance.
Step 4: confirm with cross-validation
One train/test split can be unlucky. Run 5-fold CV and look at the spread; if the gap appears in every fold, it is real.
A real-life example
A bank's loan-default model scores 0.95 ROC-AUC on training data and 0.74 on the test set. The junior data scientist adds regularisation; the test score barely moves. The team runs adversarial validation and gets 0.93: the test set is easy to tell apart from training. The top distinguishing feature is loan_product. The test period included a new "instant app loan" product that did not exist in training. The model was not overfitting; it was being tested on a new population. They set up a separate evaluation for the new product and collect a few months of its outcomes before trusting any score on it.
Follow-up questions to expect
- "What if the test score is higher than the training score?" — Usually a small, noisy test set, heavy regularisation such as dropout active only during training, or leakage into the test set. Investigate before celebrating.
- "Would you retrain on train plus test once you're happy?" — Often yes, for the final production model, but only after the test score has been recorded; you then lose your untouched estimate, so keep monitoring.
- "How big a gap is acceptable?" — There is no fixed number. Compare validation scores across models; a gap is fine if validation performance is the best you can get and meets the business target.