Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How does the bias–variance tradeoff affect model selection?


Validation accuracy as k in k-NN grows0.7410.7870.8120.8220.8090.8010.7520123456k=1, highvariancek=30, bestk=700, high biask values: 1, 3, 10, 30, 100, 300, 700, each a 5-fold average.
Validation score rises while bias falls and drops once variance takes over; model selection is finding the top of that hill.

What you need to know

The U-shaped curve

As you turn up complexity:

  • Bias falls — the model can fit more shapes.
  • Variance rises — it also fits more noise.
  • Validation error first falls (bias shrinking), bottoms out, then rises (variance growing).

Model selection means finding the bottom of that U. A validation curve plots it directly.

A validation curve you can run

In k-nearest neighbours, a small k (copy the nearest customer) is very flexible; a large k (average hundreds of customers) is very smooth.

Python
from sklearn.datasets import make_classificationfrom sklearn.model_selection import validation_curvefrom sklearn.neighbors import KNeighborsClassifierfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScaler# Telecom churn-like data: 1,500 customers, 10 features, some noisy labelsX, y = make_classification(n_samples=1500, n_features=10, n_informative=4,                           flip_y=0.15, random_state=1)model = make_pipeline(StandardScaler(), KNeighborsClassifier())ks = [1, 3, 10, 30, 100, 300, 700]train, val = validation_curve(model, X, y, param_name="kneighborsclassifier__n_neighbors",                              param_range=ks, cv=5)for k, tr, va in zip(ks, train.mean(axis=1), val.mean(axis=1)):    print(f"k={k:>3}: train {tr:.3f} | validation {va:.3f} | gap {tr - va:.3f}")
Text
k=  1: train 1.000 | validation 0.741 | gap 0.259k=  3: train 0.877 | validation 0.787 | gap 0.090k= 10: train 0.847 | validation 0.812 | gap 0.035k= 30: train 0.834 | validation 0.822 | gap 0.012k=100: train 0.816 | validation 0.809 | gap 0.007k=300: train 0.799 | validation 0.801 | gap -0.002k=700: train 0.758 | validation 0.752 | gap 0.006

Read it from both ends:

  • k = 1: perfect training score, huge gap. High variance.
  • k = 700: both scores fall together, tiny gap. High bias.
  • k = 30: the best validation score, 0.822, with a small gap. The bottom of the U.

Each number is an average over 5 cross-validation folds, so the choice does not depend on one lucky split.

How the data changes the answer

SituationLean towardsWhy
Small or noisy datasetSimpler, regularised modelsVariance is the bigger danger
Large, clean datasetMore complex modelsEnough data to keep variance in check
Unstable model (high variance)Bagging, random forestsAveraging cancels noise
Consistently weak model (high bias)Boosting, richer featuresEach step corrects remaining error
Explainability requiredSimpler model, accept some biasStakeholders must follow the logic

The dials you actually turn

Tree max_depth and min_samples_leaf, k in k-NN, regularisation strength (alpha, C), number of boosting rounds, network size, dropout rate. Each moves the model along the bias–variance axis. The hyperparameter tuning section covers searching them efficiently.

A real-life example

A lender compares three models for default prediction using 5-fold cross-validation on 40,000 loans: logistic regression (AUC 0.74, tiny gap, high bias), an unlimited decision tree (training AUC 1.00, validation 0.66, high variance), and gradient boosting with depth 4 and early stopping (validation 0.81, small gap). They choose boosting. Six months later, a new product line has only 3,000 loans. There, boosting overfits and the logistic regression wins. Same tradeoff, different data size, different answer.

Follow-up questions to expect

  • "How do deep neural networks fit this picture?" — Very large networks sometimes generalise better as they grow past the point where they can memorise the data, an effect called double descent. The classic U-shape is still the right starting intuition, and regularisation and data size still matter.
  • "Why do ensembles work?" — Bagging averages many high-variance models to cut variance; boosting adds weak, high-bias models one by one to cut bias.
  • "Which metric do you use to choose?" — The validation metric that matches the business goal, averaged across cross-validation folds, never the training score.