Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How do you split data when the dataset is very small?


Stratified 5-fold on 120 patientsvaltraintraintraintraintrainvaltraintraintraintraintrainvaltraintraintraintraintrainvaltraintraintraintraintrainvalPart APart BPart CPart DPart EFold 1: 0.875Fold 2: 0.917Fold 3: 0.958Fold 4: 0.958Fold 5: 0.917Mean 0.925, plus or minus 0.031.
Every patient is tested exactly once, and the spread across folds is reported alongside the mean.

What you need to know

Why one split fails on small data

With 120 rows, a 20% test set is 24 rows. One wrong prediction moves accuracy by about 4 points. So the score depends heavily on which 24 rows you happened to pick.

See the difference

Python
import numpy as npfrom sklearn.datasets import load_breast_cancerfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import StratifiedKFold, cross_val_score, train_test_splitfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerX, y = load_breast_cancer(return_X_y=True)X, y = X[:120], y[:120]                  # pretend we only have 120 patientsmodel = make_pipeline(StandardScaler(), LogisticRegression())# One 80/20 split: the score depends on which 24 rows landed in the test setfor seed in [0, 1, 2]:    X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=seed)    print(f"single split, seed {seed}:", round(model.fit(X_tr, y_tr).score(X_te, y_te), 3))# Stratified 5-fold: every row is tested exactly oncecv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)scores = cross_val_score(model, X, y, cv=cv)print("fold scores:", scores.round(3))print(f"mean {scores.mean():.3f} +/- {scores.std():.3f}")
Text
single split, seed 0: 0.958single split, seed 1: 1.0single split, seed 2: 0.917fold scores: [0.875 0.917 0.958 0.958 0.917]mean 0.925 +/- 0.031

Three single splits give 0.917, 0.958 and a perfect 1.0, depending only on the random seed. Which one would you report? Cross-validation answers with 0.925 plus or minus 0.031, which is both more stable and more honest about the uncertainty. The scaler sits inside the pipeline, so it is refitted in every fold without leakage.

Choosing the right kind of cross-validation

SituationUseWhy
Classification, especially imbalancedStratifiedKFoldKeeps class ratios in every fold
Several rows per customer or patientGroupKFoldKeeps one person's rows in one fold
Time-ordered dataTimeSeriesSplitAlways validates on later data
Very small data (a few dozen rows)Leave-one-out, or repeated k-foldUses nearly all data for training each time
Also tuning hyperparametersNested cross-validationKeeps the estimate honest

Other habits for small data

  • Prefer simpler, regularised models. A deep tree or big network will memorise 120 rows.
  • Report the spread across folds, not just the mean.
  • Keep a small untouched test set if you can afford it, or collect a few new cases later to confirm.

A real-life example

A diagnostics start-up has 180 labelled samples from a new blood test, 40 of them positive. A 20% test set would contain only 8 positives, so a single missed case changes recall by 12.5 points. The team uses stratified 5-fold cross-validation repeated 10 times with different shuffles, and reports recall as 0.81 with a range of 0.70 to 0.90 across runs. That range goes into the report to doctors, because a single number would hide how uncertain the estimate is with so little data.

Follow-up questions to expect

  • "Why 5 or 10 folds?" — They balance cost and stability: each model trains on 80–90% of the data, and you get enough scores to see the spread. More folds cost more training time.
  • "What is the downside of leave-one-out?" — It trains n models, which is slow, and its estimate can have high variance. 5- or 10-fold is usually preferred unless data is tiny.
  • "After cross-validation, which model do you deploy?" — None of the k fold models. You retrain one final model on all the data with the chosen settings.