Course Content
Machine Learning Foundations
14 sections · 70 lessons
How does the bias–variance tradeoff affect model selection?
What you need to know
The U-shaped curve
As you turn up complexity:
- Bias falls — the model can fit more shapes.
- Variance rises — it also fits more noise.
- Validation error first falls (bias shrinking), bottoms out, then rises (variance growing).
Model selection means finding the bottom of that U. A validation curve plots it directly.
A validation curve you can run
In k-nearest neighbours, a small k (copy the nearest customer) is very flexible; a large k (average hundreds of customers) is very smooth.
1from sklearn.datasets import make_classification2from sklearn.model_selection import validation_curve3from sklearn.neighbors import KNeighborsClassifier4from sklearn.pipeline import make_pipeline5from sklearn.preprocessing import StandardScaler67# Telecom churn-like data: 1,500 customers, 10 features, some noisy labels8X, y = make_classification(n_samples=1500, n_features=10, n_informative=4,9 flip_y=0.15, random_state=1)10model = make_pipeline(StandardScaler(), KNeighborsClassifier())11ks = [1, 3, 10, 30, 100, 300, 700]1213train, val = validation_curve(model, X, y, param_name="kneighborsclassifier__n_neighbors",14 param_range=ks, cv=5)15for k, tr, va in zip(ks, train.mean(axis=1), val.mean(axis=1)):16 print(f"k={k:>3}: train {tr:.3f} | validation {va:.3f} | gap {tr - va:.3f}")k= 1: train 1.000 | validation 0.741 | gap 0.259k= 3: train 0.877 | validation 0.787 | gap 0.090k= 10: train 0.847 | validation 0.812 | gap 0.035k= 30: train 0.834 | validation 0.822 | gap 0.012k=100: train 0.816 | validation 0.809 | gap 0.007k=300: train 0.799 | validation 0.801 | gap -0.002k=700: train 0.758 | validation 0.752 | gap 0.006Read it from both ends:
- k = 1: perfect training score, huge gap. High variance.
- k = 700: both scores fall together, tiny gap. High bias.
- k = 30: the best validation score, 0.822, with a small gap. The bottom of the U.
Each number is an average over 5 cross-validation folds, so the choice does not depend on one lucky split.
How the data changes the answer
| Situation | Lean towards | Why |
|---|---|---|
| Small or noisy dataset | Simpler, regularised models | Variance is the bigger danger |
| Large, clean dataset | More complex models | Enough data to keep variance in check |
| Unstable model (high variance) | Bagging, random forests | Averaging cancels noise |
| Consistently weak model (high bias) | Boosting, richer features | Each step corrects remaining error |
| Explainability required | Simpler model, accept some bias | Stakeholders must follow the logic |
The dials you actually turn
Tree max_depth and min_samples_leaf, k in k-NN, regularisation strength (alpha, C), number of boosting rounds, network size, dropout rate. Each moves the model along the bias–variance axis. The hyperparameter tuning section covers searching them efficiently.
A real-life example
A lender compares three models for default prediction using 5-fold cross-validation on 40,000 loans: logistic regression (AUC 0.74, tiny gap, high bias), an unlimited decision tree (training AUC 1.00, validation 0.66, high variance), and gradient boosting with depth 4 and early stopping (validation 0.81, small gap). They choose boosting. Six months later, a new product line has only 3,000 loans. There, boosting overfits and the logistic regression wins. Same tradeoff, different data size, different answer.
Follow-up questions to expect
- "How do deep neural networks fit this picture?" — Very large networks sometimes generalise better as they grow past the point where they can memorise the data, an effect called double descent. The classic U-shape is still the right starting intuition, and regularisation and data size still matter.
- "Why do ensembles work?" — Bagging averages many high-variance models to cut variance; boosting adds weak, high-bias models one by one to cut bias.
- "Which metric do you use to choose?" — The validation metric that matches the business goal, averaged across cross-validation folds, never the training score.