Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What happens when a model has high bias and low variance?


What you need to know

What you see

  • Training error is high. The model cannot fit even the data it learned from.
  • Validation error is about the same as training error.
  • Cross-validation folds give similar scores (low variance).
  • Learning curves have flattened: more data does not move them.

Stable but wrong, in numbers

Here we train on five different samples of 500 orders and ask each model for the delivery time of a 14 km order, whose true expected time is 88.4 minutes.

Python
import numpy as npfrom sklearn.linear_model import LinearRegressionfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import PolynomialFeaturesrng = np.random.default_rng(1)def make_orders(n):    km = rng.uniform(0.5, 15, n)    minutes = 10 + 0.4 * km**2 + rng.normal(0, 2, n)      # true time for 14 km = 88.4 min    return km.reshape(-1, 1), minutesline_preds, curve_preds = [], []for sample in range(5):                                    # 5 different training sets    X, y = make_orders(500)    line = LinearRegression().fit(X, y)    curve = make_pipeline(PolynomialFeatures(2), LinearRegression()).fit(X, y)    line_preds.append(line.predict([[14]])[0])    curve_preds.append(curve.predict([[14]])[0])print("straight line:", np.round(line_preds, 1))print("with km^2    :", np.round(curve_preds, 1))
Text
straight line: [79.9 79.8 79.8 80.3 80.2]with km^2    : [88.5 88.6 88.7 88.1 88.3]

The straight line is remarkably consistent (79.8 to 80.3) and consistently 8 minutes too low. That is high bias, low variance. Adding one feature, distance squared, gives the model the right shape. Predictions stay just as stable and move onto the truth. We fixed bias without adding variance, because the new capacity matched the real pattern.

Fixes, in order of preference

FixExample
Add informative featuresDistance squared, peak-hour flag, rain flag
Use a more flexible modelGradient boosting instead of linear regression
Reduce regularisationLower alpha in ridge, raise C in logistic regression
Relax structural limitsIncrease max_depth, lower min_samples_leaf
Train longer (neural networks)More epochs, check the learning rate

What does not help: collecting more of the same data, bagging (it reduces variance, which is already low), or adding regularisation.

A real-life example

A car-insurance company prices premiums with a model that uses only the driver's age and the car's price, fitted as straight lines. Every year it is retrained, and every year the premiums come out almost the same, and every year the claims team says young drivers in big cities are under-priced. The model is stable and wrong: high bias. Adding city tier, an age-and-city interaction (young drivers in metros behave differently) and a gradient-boosting model cuts the pricing error for that segment in half.

Follow-up questions to expect

  • "Why does more data not help here?" — Because the error comes from the model's shape, not from noise in a small sample. The line learned from 50,000 rows is the same line it learned from 500.
  • "Could this be a data problem instead?" — Yes. If key information is missing from the features, even a flexible model underfits. Check whether a domain expert could predict the target from your features.
  • "When is high bias acceptable?" — When the data is small or noisy, when you need a stable and explainable model, or when the error it leaves is small enough for the business.