Course Content
Machine Learning Foundations
14 sections · 70 lessons
How do you split data when the dataset is very small?
What you need to know
Why one split fails on small data
With 120 rows, a 20% test set is 24 rows. One wrong prediction moves accuracy by about 4 points. So the score depends heavily on which 24 rows you happened to pick.
See the difference
1import numpy as np2from sklearn.datasets import load_breast_cancer3from sklearn.linear_model import LogisticRegression4from sklearn.model_selection import StratifiedKFold, cross_val_score, train_test_split5from sklearn.pipeline import make_pipeline6from sklearn.preprocessing import StandardScaler78X, y = load_breast_cancer(return_X_y=True)9X, y = X[:120], y[:120] # pretend we only have 120 patients10model = make_pipeline(StandardScaler(), LogisticRegression())1112# One 80/20 split: the score depends on which 24 rows landed in the test set13for seed in [0, 1, 2]:14 X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=seed)15 print(f"single split, seed {seed}:", round(model.fit(X_tr, y_tr).score(X_te, y_te), 3))1617# Stratified 5-fold: every row is tested exactly once18cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)19scores = cross_val_score(model, X, y, cv=cv)20print("fold scores:", scores.round(3))21print(f"mean {scores.mean():.3f} +/- {scores.std():.3f}")single split, seed 0: 0.958single split, seed 1: 1.0single split, seed 2: 0.917fold scores: [0.875 0.917 0.958 0.958 0.917]mean 0.925 +/- 0.031Three single splits give 0.917, 0.958 and a perfect 1.0, depending only on the random seed. Which one would you report? Cross-validation answers with 0.925 plus or minus 0.031, which is both more stable and more honest about the uncertainty. The scaler sits inside the pipeline, so it is refitted in every fold without leakage.
Choosing the right kind of cross-validation
| Situation | Use | Why |
|---|---|---|
| Classification, especially imbalanced | StratifiedKFold | Keeps class ratios in every fold |
| Several rows per customer or patient | GroupKFold | Keeps one person's rows in one fold |
| Time-ordered data | TimeSeriesSplit | Always validates on later data |
| Very small data (a few dozen rows) | Leave-one-out, or repeated k-fold | Uses nearly all data for training each time |
| Also tuning hyperparameters | Nested cross-validation | Keeps the estimate honest |
Other habits for small data
- Prefer simpler, regularised models. A deep tree or big network will memorise 120 rows.
- Report the spread across folds, not just the mean.
- Keep a small untouched test set if you can afford it, or collect a few new cases later to confirm.
A real-life example
A diagnostics start-up has 180 labelled samples from a new blood test, 40 of them positive. A 20% test set would contain only 8 positives, so a single missed case changes recall by 12.5 points. The team uses stratified 5-fold cross-validation repeated 10 times with different shuffles, and reports recall as 0.81 with a range of 0.70 to 0.90 across runs. That range goes into the report to doctors, because a single number would hide how uncertain the estimate is with so little data.
Follow-up questions to expect
- "Why 5 or 10 folds?" — They balance cost and stability: each model trains on 80–90% of the data, and you get enough scores to see the spread. More folds cost more training time.
- "What is the downside of leave-one-out?" — It trains n models, which is slow, and its estimate can have high variance. 5- or 10-fold is usually preferred unless data is tiny.
- "After cross-validation, which model do you deploy?" — None of the k fold models. You retrain one final model on all the data with the chosen settings.