Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What steps do you take when training accuracy is high but test accuracy is low?


Adversarial validation: can a model tell train from test?Label trainingrows 0, test rows 1Train aclassifieron that labelSamemonths: AUC 0.49Festive-saleweek: AUC 0.87Top featureshows what shiftedShift, not overfitting, needs new training data rather than more regularisation.
If a classifier can separate the two sets, the gap comes from a different test population, not from memorising.

What you need to know

"Train 98%, test 72%" has four possible causes. They need different fixes, so identify which one you have.

CauseHow to recognise itFix
OverfittingGap shrinks with regularisation or more dataReduce variance
Split or feature problemTest has new groups; ID-like features in the modelGroup-aware split, drop ID columns
Distribution shiftTrain and test rows are easy to tell apartTrain on data like the test period; reweight
Unlucky splitGap varies a lot across CV foldsUse k-fold CV, larger test set

Step 1: check the split

  • Time — if the test set is a later period than training, the gap may be drift, not overfitting.
  • New groups — if training repeats the same customers many times and the test set holds new customers, the model may have learned who each customer is. Training looks great; new customers look bad. Validate with GroupKFold so you see this during development, not at the end.
  • ID-like features — a user ID, order number or raw timestamp passed in as a number lets a tree memorise individual rows. Drop them.
  • Duplicates — near-identical rows on both sides make the test number unreliable. Deduplicate before splitting.

Step 2: check whether the test set looks different

Adversarial validation is a quick, practical check. Label every training row 0 and every test row 1, then train a classifier to tell them apart. If it cannot (AUC near 0.5), train and test look alike. If it can, they differ, and the "overfitting" may really be a shift.

Python
import numpy as np, pandas as pdfrom sklearn.ensemble import HistGradientBoostingClassifierfrom sklearn.model_selection import cross_val_scorerng = np.random.default_rng(0)def orders(n, avg_basket):                       # two features per order    return pd.DataFrame({"basket_value": rng.normal(avg_basket, 300, n),                         "items": rng.poisson(4, n)})train = orders(5000, avg_basket=900)             # January to Junetest_same = orders(2000, avg_basket=900)         # a random slice of the same monthstest_sale = orders(2000, avg_basket=1400)        # festive-sale weekfor name, test in [("same period", test_same), ("sale week", test_sale)]:    X = pd.concat([train, test])    is_test = np.r_[np.zeros(len(train)), np.ones(len(test))]   # label = "which set is this row from?"    auc = cross_val_score(HistGradientBoostingClassifier(), X, is_test, cv=5, scoring="roc_auc").mean()    print(f"{name:11s} train-vs-test AUC = {auc:.2f}")
Text
same period train-vs-test AUC = 0.49sale week   train-vs-test AUC = 0.87

A test set from the same months is indistinguishable from training (0.49). The festive-sale week is easy to spot (0.87), because basket values are much higher. A model that does badly on sale week may not be overfitting at all; it has simply never seen sale-week behaviour. The fix is to include past sale periods in training, not to add regularisation. The classifier's feature importances also tell you which features shifted.

Step 3: if it really is overfitting, reduce variance

  • More training data — the most reliable fix.
  • Regularisation — L1/L2, dropout, min_samples_leaf, max_depth (see the Regularization section).
  • Simpler model or fewer features — especially when features outnumber the signal.
  • Early stopping for boosting and neural networks.
  • Ensembling — random forests and bagging average away variance.

Step 4: confirm with cross-validation

One train/test split can be unlucky. Run 5-fold CV and look at the spread; if the gap appears in every fold, it is real.

A real-life example

A bank's loan-default model scores 0.95 ROC-AUC on training data and 0.74 on the test set. The junior data scientist adds regularisation; the test score barely moves. The team runs adversarial validation and gets 0.93: the test set is easy to tell apart from training. The top distinguishing feature is loan_product. The test period included a new "instant app loan" product that did not exist in training. The model was not overfitting; it was being tested on a new population. They set up a separate evaluation for the new product and collect a few months of its outcomes before trusting any score on it.

Follow-up questions to expect

  • "What if the test score is higher than the training score?" — Usually a small, noisy test set, heavy regularisation such as dropout active only during training, or leakage into the test set. Investigate before celebrating.
  • "Would you retrain on train plus test once you're happy?" — Often yes, for the final production model, but only after the test score has been recorded; you then lose your untouched estimate, so keep monitoring.
  • "How big a gap is acceptable?" — There is no fixed number. Compare validation scores across models; a gap is fine if validation performance is the best you can get and meets the business target.