Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is bias in Machine Learning models?


What you need to know

This is statistical bias. It is different from fairness bias, where a model treats groups of people unfairly. Interviewers sometimes ask about both, so say which one you mean.

Where bias comes from

Every model assumes a shape. Linear regression assumes "one extra km adds the same minutes, always". A depth-2 tree assumes the world has at most four kinds of cases. If reality does not fit the assumed shape, the model is wrong in a consistent, repeatable way. That consistent error is bias.

More data does not fix bias

Delivery time grows faster than distance: a 14 km order takes much longer than twice a 7 km order, because long trips go through more traffic. Watch a straight line try to learn this with more and more data.

Python
import numpy as npfrom sklearn.linear_model import LinearRegressionfrom sklearn.metrics import mean_absolute_errorrng = np.random.default_rng(0)def make_orders(n):    km = rng.uniform(0.5, 15, n)                               # distance in km    minutes = 10 + 0.4 * km**2 + rng.normal(0, 2, n)           # time grows faster than distance    return km.reshape(-1, 1), minutesX_test, y_test = make_orders(5000)for n in [50, 500, 50000]:    X, y = make_orders(n)    line = LinearRegression().fit(X, y)    print(f"{n:>6} training orders: train MAE {mean_absolute_error(y, line.predict(X)):.2f}"          f" | test MAE {mean_absolute_error(y_test, line.predict(X_test)):.2f} minutes")
Text
    50 training orders: train MAE 5.59 | test MAE 5.70 minutes   500 training orders: train MAE 5.47 | test MAE 5.54 minutes 50000 training orders: train MAE 5.55 | test MAE 5.57 minutes

A thousand times more data, and the error stays near 5.5 minutes. Training and test error are almost the same, so this is not overfitting. The random noise in this data alone would cause an MAE of only about 1.6 minutes (a curved model reaches that), so roughly 4 minutes of the error is bias: the line's shape is wrong, and no amount of data can bend it.

How to reduce bias

  • Use a more flexible model (trees, boosting, a neural network).
  • Add features that describe the real shape, such as distance squared or "is peak hour".
  • Reduce regularisation strength.
  • Boosting reduces bias by adding models that correct the previous ones' errors.

A real-life example

Production scenario. A bank predicts credit-card spend with linear regression on income. It consistently under-predicts for high earners, whose spending jumps sharply with travel and luxury purchases. Adding a year of extra data changes nothing. Switching to gradient boosting, which can model the jump, reduces error for the top income band by 40%. The problem was the model's assumed shape.

Follow-up questions to expect

  • "How do you detect high bias?" — Training error itself is high, and validation error is close to it. The model cannot even fit data it has seen.
  • "Is bias always bad?" — Some bias is useful. A simpler, slightly biased model often has much lower variance and wins on total error, especially with little data.
  • "What is the difference between this bias and fairness bias?" — Statistical bias is systematic prediction error. Fairness bias is unequal treatment of groups, often learned from biased historical labels.