Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

Why do we split data into training and testing sets?


What you need to know

The goal is generalisation, not memory

A model is only useful on future cases: tomorrow's loan applications, next month's customers. You cannot measure the future, so you simulate it. You hide part of your data, train on the rest, and check the hidden part. The hidden part is the test set.

See memorisation happen

A decision tree with no limits can memorise every training row. Here it predicts loan default on simulated data.

Python
import numpy as npfrom sklearn.tree import DecisionTreeClassifierfrom sklearn.model_selection import train_test_splitrng = np.random.default_rng(3)n = 3000income = rng.normal(60, 20, n)            # monthly income, thousand rupeesemi_ratio = rng.uniform(0, 0.8, n)        # share of income already going to EMIsscore = rng.normal(700, 60, n)            # credit scorerisk = -5.5 + 7 * emi_ratio - 0.03 * (score - 700) - 0.03 * (income - 60)default = rng.random(n) < 1 / (1 + np.exp(-risk))X = np.column_stack([income, emi_ratio, score])X_tr, X_te, y_tr, y_te = train_test_split(X, default, test_size=0.2,                                          stratify=default, random_state=0)tree = DecisionTreeClassifier(random_state=0).fit(X_tr, y_tr)   # no depth limitprint("accuracy on training data:", round(tree.score(X_tr, y_tr), 3))print("accuracy on test data    :", round(tree.score(X_te, y_te), 3))print("always predict 'repays'  :", round(1 - y_te.mean(), 3))
Text
accuracy on training data: 1.0accuracy on test data    : 0.832always predict 'repays'  : 0.807

Training accuracy is a perfect 1.0, which is a lie. On unseen applicants the tree scores 0.832, only a little better than the lazy rule "everyone repays" at 0.807. Without the split you would ship this model believing it was perfect.

How to split

  • Size. 80/20 or 70/30 is common. With millions of rows, 1% as test can be plenty; what matters is enough test rows to give a stable number.
  • Stratify for classification: stratify=y keeps the same share of defaulters in both parts. Without it, a small test set might get too few positives.
  • Fix random_state so the split is repeatable and teammates can compare results.
  • Respect time and groups. For time-based data, train on the past and test on the most recent period. If one customer has many rows, keep all of them on one side.

Why "use it once" matters

Every time you look at the test score and change the model because of it, a bit of the test set leaks into your decisions. After enough rounds you have tuned the model to those specific rows. That is why we also need a validation set, covered in the next lesson.

A real-life example

Production scenario. A team building a spam filter reports 99.2% accuracy. A reviewer asks how it was measured, and the answer is "on the training emails". Re-measured on a held-out 20%, it is 91%. Worse, when tested on emails from the last month only, it is 86%, because spammers had changed their wording. The team now keeps the most recent month as the test set, since that best matches what the model will see next.

Follow-up questions to expect

  • "What ratio would you use?" — 80/20 is a sensible default. With very large data a smaller test share is fine; with small data use cross-validation instead of one split.
  • "What does stratify do?" — It keeps class proportions equal in train and test, so a rare class is not accidentally missing from one of them.
  • "When is a random split wrong?" — For time series (you would train on the future) and for grouped data (the same user appears on both sides). Use time-based or group-based splits.