Course Content
Machine Learning Foundations
14 sections · 70 lessons
What happens when a model has high bias and low variance?
What you need to know
What you see
- Training error is high. The model cannot fit even the data it learned from.
- Validation error is about the same as training error.
- Cross-validation folds give similar scores (low variance).
- Learning curves have flattened: more data does not move them.
Stable but wrong, in numbers
Here we train on five different samples of 500 orders and ask each model for the delivery time of a 14 km order, whose true expected time is 88.4 minutes.
1import numpy as np2from sklearn.linear_model import LinearRegression3from sklearn.pipeline import make_pipeline4from sklearn.preprocessing import PolynomialFeatures56rng = np.random.default_rng(1)7def make_orders(n):8 km = rng.uniform(0.5, 15, n)9 minutes = 10 + 0.4 * km**2 + rng.normal(0, 2, n) # true time for 14 km = 88.4 min10 return km.reshape(-1, 1), minutes1112line_preds, curve_preds = [], []13for sample in range(5): # 5 different training sets14 X, y = make_orders(500)15 line = LinearRegression().fit(X, y)16 curve = make_pipeline(PolynomialFeatures(2), LinearRegression()).fit(X, y)17 line_preds.append(line.predict([[14]])[0])18 curve_preds.append(curve.predict([[14]])[0])1920print("straight line:", np.round(line_preds, 1))21print("with km^2 :", np.round(curve_preds, 1))straight line: [79.9 79.8 79.8 80.3 80.2]with km^2 : [88.5 88.6 88.7 88.1 88.3]The straight line is remarkably consistent (79.8 to 80.3) and consistently 8 minutes too low. That is high bias, low variance. Adding one feature, distance squared, gives the model the right shape. Predictions stay just as stable and move onto the truth. We fixed bias without adding variance, because the new capacity matched the real pattern.
Fixes, in order of preference
| Fix | Example |
|---|---|
| Add informative features | Distance squared, peak-hour flag, rain flag |
| Use a more flexible model | Gradient boosting instead of linear regression |
| Reduce regularisation | Lower alpha in ridge, raise C in logistic regression |
| Relax structural limits | Increase max_depth, lower min_samples_leaf |
| Train longer (neural networks) | More epochs, check the learning rate |
What does not help: collecting more of the same data, bagging (it reduces variance, which is already low), or adding regularisation.
A real-life example
A car-insurance company prices premiums with a model that uses only the driver's age and the car's price, fitted as straight lines. Every year it is retrained, and every year the premiums come out almost the same, and every year the claims team says young drivers in big cities are under-priced. The model is stable and wrong: high bias. Adding city tier, an age-and-city interaction (young drivers in metros behave differently) and a gradient-boosting model cuts the pricing error for that segment in half.
Follow-up questions to expect
- "Why does more data not help here?" — Because the error comes from the model's shape, not from noise in a small sample. The line learned from 50,000 rows is the same line it learned from 500.
- "Could this be a data problem instead?" — Yes. If key information is missing from the features, even a flexible model underfits. Check whether a domain expert could predict the target from your features.
- "When is high bias acceptable?" — When the data is small or noisy, when you need a stable and explainable model, or when the error it leaves is small enough for the business.