Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is data leakage, and how does improper splitting cause it?
What you need to know
Leakage through the split
- Preprocessing before splitting. A
StandardScalerfitted on all rows learns the mean of the test rows too. An imputer fills missing values using test statistics. A feature selector picks columns because they correlate with the test labels. - Duplicates across splits. The same email forwarded twice, the same flat listed by two brokers. The model has effectively seen the test row.
- Random split on time-ordered data. Predicting March sales with a model that saw April.
- Groups split randomly. Ten scans from one patient, eight in train and two in test. The model learns to recognise the patient, not the disease.
Leakage through features (not the split, but related)
A feature computed from the outcome itself: "number of collection calls" to predict loan default, or "refund issued" to predict a complaint. These exist only after the event.
How big can leakage be? A demonstration
Here there is nothing to learn: 5,000 random columns and random labels. The only difference between the two numbers is when feature selection runs.
1import numpy as np2from sklearn.feature_selection import SelectKBest, f_classif3from sklearn.linear_model import LogisticRegression4from sklearn.model_selection import cross_val_score5from sklearn.pipeline import make_pipeline67rng = np.random.default_rng(0)8X = rng.normal(size=(200, 5000)) # 200 rows, 5,000 random columns9y = rng.integers(0, 2, 200) # random labels: nothing to learn1011# WRONG: pick the 20 "best" columns using ALL rows, then cross-validate12X_picked = SelectKBest(f_classif, k=20).fit_transform(X, y)13leaky = cross_val_score(LogisticRegression(), X_picked, y, cv=5).mean()1415# RIGHT: selection happens inside each training fold only16pipe = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression())17honest = cross_val_score(pipe, X, y, cv=5).mean()1819print("leaky accuracy :", round(leaky, 3))20print("honest accuracy:", round(honest, 3))leaky accuracy : 0.795honest accuracy: 0.51The leaky version reports 79.5% on data that has no pattern at all. It chose the 20 columns that, by chance, lined up with the labels of every row, including the rows later used for validation. The pipeline version re-runs selection inside each fold, sees only training rows, and reports the truth: 51%, a coin flip.
The rule this teaches: anything that is fitted, including scalers, imputers, encoders and selectors, is part of the model, and must be fitted on training data only. A scikit-learn Pipeline enforces this automatically.
A real-life example
A hospital team builds a model to detect pneumonia from chest X-rays and gets 97% accuracy. A reviewer notices the dataset has several images per patient, split randomly. Re-splitting by patient drops accuracy to 84%. A second leak appears: images from the intensive-care unit used a portable machine, and nearly all ICU patients had pneumonia, so the model was partly learning "which machine took this photo". The fix was a group split by patient and removing the scanner-specific markings.
Follow-up questions to expect
- "How do you detect leakage?" — Be suspicious of scores that are too good, check feature importance for a single dominant feature, ask "would I know this value at prediction time?", and look for duplicates and shared IDs across splits.
- "How does a Pipeline prevent leakage?" — During
fitit fits every step on the training data only, and during cross-validation it refits all steps inside each fold. - "How do you split time-series data?" — Chronologically: train on earlier periods, validate and test on later ones. scikit-learn's
TimeSeriesSplitdoes this for cross-validation.