Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is variance in Machine Learning models?


What you need to know

A thought experiment you can actually run

Imagine collecting a fresh sample of 200 flat sales in Bengaluru, training a model, and asking it to price one flat: 1,200 sq ft, 10 km from the centre. Now do it 30 times, with 30 different samples. A stable model gives nearly the same answer each time. An unstable one jumps around.

Python
import numpy as npfrom sklearn.linear_model import LinearRegressionfrom sklearn.tree import DecisionTreeRegressorfrom sklearn.ensemble import RandomForestRegressorrng = np.random.default_rng(0)def sample_flats(n):                        # a fresh sample of Bengaluru sales    area, dist = rng.uniform(500, 2500, n), rng.uniform(2, 30, n)    price = 20 + 0.07 * area - 1.2 * dist + rng.normal(0, 10, n)    # lakh    return np.column_stack([area, dist]), priceflat = [[1200, 10]]                         # true expected price: 92 lakhmodels = {"linear regression": LinearRegression,          "deep tree": lambda: DecisionTreeRegressor(random_state=0),          "random forest": lambda: RandomForestRegressor(n_estimators=100, random_state=0)}for name, make in models.items():    preds = [make().fit(*sample_flats(200)).predict(flat)[0] for _ in range(30)]    print(f"{name:<17}: mean {np.mean(preds):5.1f} | spread (std) {np.std(preds):4.1f} lakh")
Text
linear regression: mean  92.0 | spread (std)  0.7 lakhdeep tree        : mean  92.7 | spread (std) 11.0 lakhrandom forest    : mean  92.4 | spread (std)  4.9 lakh

All three are right on average: about 92 lakh. The difference is the spread. The deep tree's answer for the same flat swings by about 11 lakh depending on which 200 sales it happened to see. It memorised each sample's noise. That swing is variance. A random forest averages 100 trees trained on different resamples, so their individual swings partly cancel, cutting the spread by more than half. The linear model is the most stable here because the true pattern really is linear.

Why variance hurts in production

In production you train once, on one sample. With a high-variance model, your one model may be one of the unlucky ones, and you cannot tell which from the training score. Retrain next month on slightly different data and predictions change noticeably, which confuses users and downstream teams.

How to reduce variance

  • More training data.
  • Simpler models or stronger regularisation.
  • Bagging, such as random forests: average many models trained on different samples.
  • Early stopping, dropout, and limits such as min_samples_leaf.

A real-life example

A telecom retrains its churn model every week on the latest 20,000 customers using an unconstrained gradient-boosting model. The marketing team notices that the "top 1,000 at-risk customers" list changes by 60% from week to week, even though customers have hardly changed. That instability is variance. The team adds early stopping, raises the minimum samples per leaf, and trains on 12 weeks of data instead of 1. The overlap between consecutive weekly lists rises to 85%, and validation AUC slightly improves.

Follow-up questions to expect

  • "How do you detect high variance?" — A large gap between training and validation scores, and large differences in scores across cross-validation folds.
  • "Why does bagging reduce variance but not bias?" — Averaging many noisy models cancels their random errors, but if every model makes the same systematic error, the average keeps it.
  • "Does k in k-NN affect variance?" — Yes. k = 1 copies the nearest single point, so it has very high variance; larger k averages more neighbours and is more stable but more biased.