Course Content
Machine Learning Foundations
14 sections · 70 lessons
Why do we split data into training and testing sets?
What you need to know
The goal is generalisation, not memory
A model is only useful on future cases: tomorrow's loan applications, next month's customers. You cannot measure the future, so you simulate it. You hide part of your data, train on the rest, and check the hidden part. The hidden part is the test set.
See memorisation happen
A decision tree with no limits can memorise every training row. Here it predicts loan default on simulated data.
1import numpy as np2from sklearn.tree import DecisionTreeClassifier3from sklearn.model_selection import train_test_split45rng = np.random.default_rng(3)6n = 30007income = rng.normal(60, 20, n) # monthly income, thousand rupees8emi_ratio = rng.uniform(0, 0.8, n) # share of income already going to EMIs9score = rng.normal(700, 60, n) # credit score10risk = -5.5 + 7 * emi_ratio - 0.03 * (score - 700) - 0.03 * (income - 60)11default = rng.random(n) < 1 / (1 + np.exp(-risk))1213X = np.column_stack([income, emi_ratio, score])14X_tr, X_te, y_tr, y_te = train_test_split(X, default, test_size=0.2,15 stratify=default, random_state=0)1617tree = DecisionTreeClassifier(random_state=0).fit(X_tr, y_tr) # no depth limit18print("accuracy on training data:", round(tree.score(X_tr, y_tr), 3))19print("accuracy on test data :", round(tree.score(X_te, y_te), 3))20print("always predict 'repays' :", round(1 - y_te.mean(), 3))accuracy on training data: 1.0accuracy on test data : 0.832always predict 'repays' : 0.807Training accuracy is a perfect 1.0, which is a lie. On unseen applicants the tree scores 0.832, only a little better than the lazy rule "everyone repays" at 0.807. Without the split you would ship this model believing it was perfect.
How to split
- Size. 80/20 or 70/30 is common. With millions of rows, 1% as test can be plenty; what matters is enough test rows to give a stable number.
- Stratify for classification:
stratify=ykeeps the same share of defaulters in both parts. Without it, a small test set might get too few positives. - Fix
random_stateso the split is repeatable and teammates can compare results. - Respect time and groups. For time-based data, train on the past and test on the most recent period. If one customer has many rows, keep all of them on one side.
Why "use it once" matters
Every time you look at the test score and change the model because of it, a bit of the test set leaks into your decisions. After enough rounds you have tuned the model to those specific rows. That is why we also need a validation set, covered in the next lesson.
A real-life example
Production scenario. A team building a spam filter reports 99.2% accuracy. A reviewer asks how it was measured, and the answer is "on the training emails". Re-measured on a held-out 20%, it is 91%. Worse, when tested on emails from the last month only, it is 86%, because spammers had changed their wording. The team now keeps the most recent month as the test set, since that best matches what the model will see next.
Follow-up questions to expect
- "What ratio would you use?" — 80/20 is a sensible default. With very large data a smaller test share is fine; with small data use cross-validation instead of one split.
- "What does
stratifydo?" — It keeps class proportions equal in train and test, so a rare class is not accidentally missing from one of them. - "When is a random split wrong?" — For time series (you would train on the future) and for grouped data (the same user appears on both sides). Use time-based or group-based splits.