Course Content
Machine Learning Essentials
6 sections · 16 lessons
Bias-Variance Tradeoff
Here is an experiment you can run in ten minutes, and it explains more about why models fail than any definition will.
Take 30 points sampled from a gentle curve, with a little random measurement noise added to each. Fit three models.
The first is a straight line. It cannot bend, so it slices through the curve — too high in the middle, too low at both ends. Training error: large. Test error: large, and about the same size.
The second is a degree-3 polynomial. It follows the curve closely, misses the noise, and lands close to the truth almost everywhere. Training error: small. Test error: small.
The third is a degree-15 polynomial. It bends to pass through or within a whisker of nearly every one of the 30 training points — training error is close to zero. Between the points it swings violently, shooting to +40 and back down to −25 in gaps where the true curve barely moves. Test error: enormous.
Now the part that matters. Throw those 30 points away, sample 30 fresh ones from the same source, and refit all three. The straight line barely moves — its slope shifts by a hair. The degree-3 curve moves a little. The degree-15 curve is unrecognisable: where it previously spiked upward it now dives, and its predictions for a given input have changed by 60 units.
You have just watched the two distinct ways a model can be wrong, and they behave in opposite directions.
Two different kinds of wrong
Bias is error from the model being too simple to represent the real pattern. It is a property of the model family, not of the particular dataset. A straight line fitted to a curve will be wrong in the same places every time, no matter which sample you give it. Bias is systematic, repeatable wrongness.
Variance is error from the model being too sensitive to the particular training sample. A high-variance model fits the noise as though it were signal, so when the noise changes — which it does with every new sample — the model changes with it. Variance is unstable, unrepeatable wrongness.
Bias is being consistently wrong in the same way. Variance is being wrong in a different way every time you retrain.
The classic picture is a dartboard. Four players throw five darts each:
| Low variance | High variance | |
|---|---|---|
| Low bias | Tight cluster on the bullseye — what you want | Scattered all round the bullseye; average is right, individual throws are not |
| High bias | Tight cluster, but three inches up and to the left — reliably wrong | Scattered and off-centre — the worst of both |
Each dart is one model trained on one sample of data. Bias is how far the centre of the cluster sits from the target. Variance is how spread out the cluster is. Note that this only makes sense across many training runs — which is why you cannot look at a single trained model and read its variance off the page. You infer it from behaviour.
The decomposition, and what it actually claims
Suppose the true relationship is y=f(x)+ϵ, where ϵ is random noise with mean 0 and variance σ2. You train a model f^ on a random dataset. The expected squared error at a point x, averaged over all possible training datasets, splits into exactly three pieces:
Read it in plain English:
- Bias² — take the average prediction your model family makes at x across all possible training sets, and see how far that average is from the truth. Squared.
- Variance — how much your model's prediction at x bounces around that average as the training set changes.
- Irreducible error — noise in the data itself. Two identical houses sell for different prices because one buyer was in a hurry. No model of any kind can predict that, and σ2 sets a floor beneath which no amount of cleverness will take you.
The practical content of this equation is the word "plus". These are separate contributions, and actions that reduce one commonly increase the other. That is the trade-off.
The irreducible floor is not a footnote
If your data has an inherent noise level equivalent to an RMSE of £15,000 on house prices, then a model reporting £15,200 is essentially perfect, and a team spending six months trying to reach £8,000 is chasing something that does not exist. Estimating that floor early — often by asking a domain expert how consistently a human could do the task — saves entire projects.
Complexity moves both terms, in opposite directions
Model complexity means capacity to represent patterns: polynomial degree, tree depth, number of features, number of parameters, number of neighbours in k-NN (inversely).
error ^ | \ / | \ / total test error | \ / | \ ~ ~ ~ ~ ~ ~ ~ ~ ~ ~ ~ ~ ~ ~ / | \ / | \___ ___/ <- variance rises | \___ ___/ | bias falls -> \_______/ | ^ | sweet spot +--------------------------------------------> complexity underfitting overfittingAt low complexity, bias dominates: the model cannot express the pattern. At high complexity, variance dominates: the model expresses the noise. Total error is U-shaped, and the minimum is what you are hunting for.
Concrete numbers
Running the polynomial experiment properly — fitting each degree on 100 different random samples of 30 points and measuring average behaviour — produces something like this:
| Degree | Train MSE | Test MSE | Bias² | Variance | Diagnosis |
|---|---|---|---|---|---|
| 1 | 28.4 | 29.9 | 28.1 | 0.8 | Underfitting |
| 2 | 11.2 | 12.5 | 10.4 | 1.1 | Still too simple |
| 3 | 2.1 | 3.4 | 1.2 | 1.2 | Sweet spot |
| 6 | 1.4 | 6.8 | 0.9 | 4.9 | Starting to overfit |
| 10 | 0.4 | 34.2 | 0.8 | 32.4 | Overfitting badly |
| 15 | 0.1 | 412.7 | 1.1 | 410.6 | Memorised the sample |
Three things to notice.
First, training error falls monotonically and never warns you about anything. It is close to zero at degree 15 while test error is a hundred times worse than at degree 3. Training error is not a measure of quality. It is a measure of how much capacity the model had.
Second, bias² stops improving after degree 3 — the true pattern was already captured. Everything after that is pure variance growth, paid for nothing.
Third, the sweet spot is not where either term is minimised. It is where their sum is minimised, and at degree 3 bias² and variance happen to be roughly equal. That balance is typical, not coincidental.
Diagnosing your own model in two numbers
You rarely get to compute bias and variance directly — that requires many training sets you do not have. Instead you read them off the gap between training and validation error.
| Training error | Validation error | Diagnosis | What is happening |
|---|---|---|---|
| High (18%) | High (19%) | High bias / underfitting | Model cannot even fit the data it has seen |
| Low (1%) | High (17%) | High variance / overfitting | Model memorised the training set |
| Low (3%) | Low (5%) | Good fit | Small gap, both acceptable |
| High (18%) | Higher (31%) | Both | Wrong model family and overfitting the noise |
| Low (2%) | Lower (1%) | Suspicious | Leakage, a tiny validation set, or a lucky split |
Two cautions on reading this table. "High" is relative to your irreducible floor and to what the task requires — 18% error is catastrophic for digit recognition and superb for predicting individual stock movements. And validation error below training error usually means something is broken rather than something is excellent.
Learning curves: the more informative diagnostic
Plot training and validation error as functions of training set size. The shape tells you not just what is wrong but whether more data would help.
1import numpy as np2import matplotlib.pyplot as plt3from sklearn.model_selection import learning_curve45sizes, train_scores, val_scores = learning_curve(6 model, X, y, cv=5,7 train_sizes=np.linspace(0.1, 1.0, 10),8 scoring="neg_mean_squared_error",9)1011train_err = -train_scores.mean(axis=1)12val_err = -val_scores.mean(axis=1)1314plt.plot(sizes, train_err, "o-", label="training error")15plt.plot(sizes, val_err, "s-", label="validation error")16plt.xlabel("training examples")17plt.ylabel("MSE")18plt.legend()High bias: the curves meet, high up
error | validation ~~~~~~~~~~~~~~~~~~~~~~~~ | / | training ~~~~~~~~~~~~~~~~~~~~~~~~~~ both plateau at 18% | +------------------------------------> training set sizeBoth curves flatten close together at an unacceptable error. The gap is small, so variance is not your problem. More data will not help at all — the curves are already flat, and adding examples to a model that cannot express the pattern just gives it more examples to be equally wrong about. This is the most valuable thing learning curves tell you, because "get more data" is an expensive instinct.
High variance: a persistent gap
error | validation \___ | \_____ | \______ still falling | | training _____---------------- near zero +------------------------------------> training set sizeTraining error is low, validation error is much higher, and the gap is closing slowly as data increases. Here more data does help, and the slope of the validation curve tells you roughly how much. If it is still dropping steeply at your current sample size, doubling the data is a sound investment. If it has nearly flattened with a gap remaining, you need regularisation instead.
Fixing each condition
The fixes are close to mirror images, which is why misdiagnosing costs you twice — you apply the opposite of what is needed.
| Fix | Effect on bias | Effect on variance | Use when |
|---|---|---|---|
| Add features / interactions | Down | Up | Underfitting |
| Increase model complexity (deeper tree, higher degree) | Down | Up | Underfitting |
| Reduce regularisation strength | Down | Up | Underfitting |
| Train longer / to convergence | Down | Up | Underfitting |
| Collect more training data | No effect | Down | Overfitting |
| Increase regularisation strength | Up | Down | Overfitting |
| Remove or select features | Up | Down | Overfitting |
| Simplify the model | Up | Down | Overfitting |
| Early stopping | Up | Down | Overfitting |
| Bagging / averaging models | Roughly unchanged | Down | Overfitting |
| Boosting | Down | Up | Underfitting |
The two bold rows are the interesting ones because they partly escape the trade-off rather than navigating it.
Why averaging reduces variance for free
This is worth doing with numbers. If you train n models on different random subsets of the data and average their predictions, and each model's prediction has variance σ2, then the average of n independent predictions has variance:
Average 100 models and variance drops by a factor of 100 — while bias stays put, because averaging many unbiased-ish estimates leaves the centre where it was. That is precisely what a random forest does: grow many deep trees, each of which individually overfits horribly, and average them. Each tree has low bias and enormous variance; the ensemble keeps the low bias and sheds most of the variance.
The catch is the word "independent". Real trees trained on overlapping data are correlated, and correlation limits the reduction. Random forests attack this by making each tree consider only a random subset of features at each split, deliberately weakening individual trees to make them disagree more. A slightly worse but more diverse set of models averages into a better ensemble.
Regularisation, concretely
Regularisation adds a penalty on parameter size to the loss, so the fit has to justify every bit of wiggle it uses:
At λ=0 you have the unpenalised model at full variance. As λ grows, coefficients shrink, the function becomes smoother, variance falls and bias rises. At λ→∞ all coefficients go to zero and you predict the mean — maximum bias, zero variance. Somewhere in between is the minimum of the sum, and cross-validation finds it.
Where the intuition breaks down
Two honest caveats, because the U-shaped curve is a model of reality, not reality.
The U is not always U-shaped. With very large modern networks, test error famously falls, rises around the point where the model has just enough parameters to interpolate the training data, and then falls again as parameters keep increasing — the "double descent" phenomenon. So "more parameters means more overfitting" is a reliable guide for the model families in classical machine learning and an unreliable one for very heavily overparameterised networks.
Bias and variance are properties of a procedure, not a fitted model. Strictly, they are defined by averaging over hypothetical training sets you never drew. When people say "this model has high variance", they mean "this training procedure, applied to samples of this size from this distribution, produces unstable models". That distinction matters when you change the sample size: the same algorithm can be high-variance on 200 rows and well-behaved on 200,000.
A worked diagnosis
A team is predicting house prices. Their gradient boosting model reports:
- Training RMSE: £4,100
- Validation RMSE: £38,600
A gap of nearly ten times. Clear high variance. Their instinct is to try a different algorithm; the table says otherwise. Cheapest interventions first:
1from sklearn.ensemble import GradientBoostingRegressor2from sklearn.model_selection import cross_val_score34for depth in [2, 3, 4, 6, 8]:5 m = GradientBoostingRegressor(max_depth=depth, n_estimators=500,6 learning_rate=0.05, subsample=0.8,7 random_state=42)8 s = -cross_val_score(m, X_train, y_train, cv=5,9 scoring="neg_root_mean_squared_error")10 print(f"depth {depth}: {s.mean():,.0f} +/- {s.std():,.0f}")Results: depth 8 gives £38,600; depth 6 gives £31,200; depth 4 gives £26,800; depth 3 gives £25,900; depth 2 gives £28,400. The curve is U-shaped in the hyperparameter, exactly as predicted, and depth 3 is the minimum. Reducing capacity — the opposite of what most people reach for when a model underperforms — cut the error by a third.
Then a learning curve at depth 3 shows the validation error still falling at the largest sample size, with a £9,000 gap remaining. That says more data would help further, and quantifies the argument for going and getting it.
What this means when you build something
Before you change anything about an underperforming model, get two numbers: training error and validation error. Almost every improvement decision follows from the relationship between them, and almost every wasted week comes from skipping this and reaching for the fix that feels most active.
If the two numbers are both bad and close together, stop collecting data and stop regularising. You need more capacity: better features, a more expressive model, less penalty. If training error is near zero and validation error is far above it, stop adding features and stop deepening the model. You need constraint: more data, more regularisation, fewer features, or an ensemble.
And keep the floor in mind. Every dataset carries noise that no model can predict, and knowing roughly where that floor sits turns "our model is not good enough" into a question with an answer. A model sitting just above the noise floor is finished, and the honest thing to do is say so rather than spend another month proving it.