Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How does bias differ from variance intuitively?


The dartboard, measured on delivery timescurve withkm squareddeep treestraight lineworst caseLow varianceHigh varianceLow biasHigh biasBias squared and variance: line 28.9 and 3.8, tree 0.3 and 65.9, curve 0.0 and 3.0.
Each dart is one model trained on one sample: bias is where the darts centre, variance is how widely they spread.

What you need to know

The dartboard, made precise

Think of each dart as one model trained on one possible training set. Throw many darts, meaning train the same kind of model on many different samples, and look at where they land.

Low variance (tight group)High variance (scattered)
Low bias (centred on bullseye)The goal: accurate and stableRight on average, unreliable each time
High bias (centred off target)Consistently wrong: underfittingWrong and unstable: the worst case

Putting numbers on the dartboard

We can literally measure both. Train each model on 200 different samples of 60 delivery orders, and compare its predictions with the true delivery time across a range of distances.

Python
import numpy as npfrom sklearn.linear_model import LinearRegressionfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import PolynomialFeaturesfrom sklearn.tree import DecisionTreeRegressorrng = np.random.default_rng(0)true_time = lambda km: 10 + 0.4 * km**2grid = np.linspace(1, 14, 50).reshape(-1, 1)          # the "bullseyes" we aim atmodels = {"straight line": lambda: LinearRegression(),          "curve (km^2)": lambda: make_pipeline(PolynomialFeatures(2), LinearRegression()),          "deep tree": lambda: DecisionTreeRegressor(random_state=0)}for name, make in models.items():    throws = []                                         # 200 training sets = 200 darts    for _ in range(200):        km = rng.uniform(0.5, 15, 60)        y = true_time(km) + rng.normal(0, 8, 60)        throws.append(make().fit(km.reshape(-1, 1), y).predict(grid))    throws = np.array(throws)    bias_sq = np.mean((throws.mean(axis=0) - true_time(grid.ravel())) ** 2)    variance = np.mean(throws.var(axis=0))    print(f"{name:<13}: bias^2 {bias_sq:5.1f} | variance {variance:5.1f} | total {bias_sq + variance:5.1f}")
Text
straight line: bias^2  28.9 | variance   3.8 | total  32.7curve (km^2) : bias^2   0.0 | variance   3.0 | total   3.0deep tree    : bias^2   0.3 | variance  65.9 | total  66.2
  • Straight line: tight group, wrong place. High bias (28.9), low variance (3.8).
  • Deep tree: centred on the truth (bias 0.3) but scattered everywhere (variance 65.9). It is the worst of the three.
  • Curve with km squared: centred and tight. Its shape matches the truth, so both are low.

On top of these, every model also suffers the irreducible noise in the data (here 8 squared = 64), which no model can remove. For squared error the full picture is:

Text
expected error = bias^2 + variance + irreducible noise

The diagnostic in one line

Look at training and validation error together. Both high and close: bias. Training low, validation much higher: variance.

A real-life example

Production scenario. Two teams build ETA models for the same delivery app. Team A's linear model is always about 6 minutes too optimistic for long orders; customers learn to add 6 minutes. Team B's deep tree is right on average but sometimes says 18 minutes and sometimes 35 for the same kind of order; customers stop trusting it. Product managers often prefer Team A's predictable error, which shows that the "right" balance also depends on how users experience mistakes.

Follow-up questions to expect

  • "Can a model have high bias and high variance at once?" — Yes, for example a badly specified model trained on very little, noisy data. It is systematically off and unstable.
  • "Which is easier to fix?" — Variance is often fixed with more data or averaging; bias needs a change to the model or features. Neither is always easier.
  • "What is irreducible error?" — Noise in the data itself, such as random traffic delays. No model can predict it, so it sets the floor on error.