Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is variance in Machine Learning models?
What you need to know
A thought experiment you can actually run
Imagine collecting a fresh sample of 200 flat sales in Bengaluru, training a model, and asking it to price one flat: 1,200 sq ft, 10 km from the centre. Now do it 30 times, with 30 different samples. A stable model gives nearly the same answer each time. An unstable one jumps around.
1import numpy as np2from sklearn.linear_model import LinearRegression3from sklearn.tree import DecisionTreeRegressor4from sklearn.ensemble import RandomForestRegressor56rng = np.random.default_rng(0)7def sample_flats(n): # a fresh sample of Bengaluru sales8 area, dist = rng.uniform(500, 2500, n), rng.uniform(2, 30, n)9 price = 20 + 0.07 * area - 1.2 * dist + rng.normal(0, 10, n) # lakh10 return np.column_stack([area, dist]), price1112flat = [[1200, 10]] # true expected price: 92 lakh13models = {"linear regression": LinearRegression,14 "deep tree": lambda: DecisionTreeRegressor(random_state=0),15 "random forest": lambda: RandomForestRegressor(n_estimators=100, random_state=0)}16for name, make in models.items():17 preds = [make().fit(*sample_flats(200)).predict(flat)[0] for _ in range(30)]18 print(f"{name:<17}: mean {np.mean(preds):5.1f} | spread (std) {np.std(preds):4.1f} lakh")linear regression: mean 92.0 | spread (std) 0.7 lakhdeep tree : mean 92.7 | spread (std) 11.0 lakhrandom forest : mean 92.4 | spread (std) 4.9 lakhAll three are right on average: about 92 lakh. The difference is the spread. The deep tree's answer for the same flat swings by about 11 lakh depending on which 200 sales it happened to see. It memorised each sample's noise. That swing is variance. A random forest averages 100 trees trained on different resamples, so their individual swings partly cancel, cutting the spread by more than half. The linear model is the most stable here because the true pattern really is linear.
Why variance hurts in production
In production you train once, on one sample. With a high-variance model, your one model may be one of the unlucky ones, and you cannot tell which from the training score. Retrain next month on slightly different data and predictions change noticeably, which confuses users and downstream teams.
How to reduce variance
- More training data.
- Simpler models or stronger regularisation.
- Bagging, such as random forests: average many models trained on different samples.
- Early stopping, dropout, and limits such as
min_samples_leaf.
A real-life example
A telecom retrains its churn model every week on the latest 20,000 customers using an unconstrained gradient-boosting model. The marketing team notices that the "top 1,000 at-risk customers" list changes by 60% from week to week, even though customers have hardly changed. That instability is variance. The team adds early stopping, raises the minimum samples per leaf, and trains on 12 weeks of data instead of 1. The overlap between consecutive weekly lists rises to 85%, and validation AUC slightly improves.
Follow-up questions to expect
- "How do you detect high variance?" — A large gap between training and validation scores, and large differences in scores across cross-validation folds.
- "Why does bagging reduce variance but not bias?" — Averaging many noisy models cancels their random errors, but if every model makes the same systematic error, the average keeps it.
- "Does k in k-NN affect variance?" — Yes. k = 1 copies the nearest single point, so it has very high variance; larger k averages more neighbours and is more stable but more biased.