Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
How does high variance in data affect model training and generalization?
What you need to know
Meaning 1: spread in the input data
A wide-ranging feature is not a problem by itself. The trouble is mixing scales. If "income" ranges over lakhs of rupees and "age" over tens of years:
- Gradient descent takes a step size that is too big for one weight and too small for the other, so it zig-zags and converges slowly.
- k-NN, k-means and SVMs compute distances, and income differences swamp age differences.
- Regularisation penalises coefficients unevenly, because a coefficient's size depends on the feature's units.
The fix is standardising (z-scores) or min-max scaling. Tree models do not need it, because they only compare values to thresholds.
Large variance in the target, especially noise the features cannot explain, sets a floor on the error any model can reach. That part is called irreducible error.
Meaning 2: variance of the model
The bias–variance trade-off splits expected error into three parts:
expected error = bias^2 + variance + irreducible noiseA very simple model has high bias (it misses the pattern) and low variance. A very flexible model has low bias and high variance. You want the depth or complexity that minimises the total.
1from sklearn.datasets import make_classification2from sklearn.model_selection import train_test_split3from sklearn.tree import DecisionTreeClassifier45X, y = make_classification(n_samples=1000, n_features=20, n_informative=5,6 flip_y=0.1, random_state=42)7X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=42)89for depth in [None, 4]:10 tree = DecisionTreeClassifier(max_depth=depth, random_state=42).fit(X_tr, y_tr)11 print(f"max_depth={depth}: train={tree.score(X_tr, y_tr):.2f} test={tree.score(X_te, y_te):.2f}")max_depth=None: train=1.00 test=0.81max_depth=4: train=0.88 test=0.85The unlimited tree memorises the training set, including the 10% of labels that flip_y made noisy, and scores 100% on it. It drops to 81% on new data. Limiting depth to 4 lowers training accuracy but raises test accuracy: less variance, better generalisation.
Ways to reduce model variance
- More training data, or data augmentation.
- Regularisation: L1/L2 penalties, dropout, weight decay.
- A simpler model: shallower trees, fewer features.
- Bagging and random forests, which average many high-variance trees.
- Early stopping when validation loss starts rising.
A real-life example
Think of a student who memorises last year's exam paper word for word. They score 100% on that paper and fail a new one. That is high model variance. Where the analogy breaks: a model does not "know" it is memorising; it simply has enough capacity to fit every quirk, so you must measure the gap on held-out data.
Production case: a delivery ETA model built with deep gradient-boosted trees scores an MAE of 1.5 minutes on training orders and 7 minutes on last week's orders. The engineer limits tree depth, adds a minimum number of samples per leaf, and uses early stopping on a validation set. Training MAE rises to 3.5 minutes; last week's MAE falls to 4.2 minutes.
Follow-up questions to expect
- "How do you tell high variance from high bias?" — High variance: training error low, validation error much higher. High bias: both errors high and close together.
- "Why does bagging reduce variance?" — Averaging many models trained on different bootstrap samples cancels out their individual, partly independent errors.
- "Does more data fix high bias?" — Usually not; a model too simple for the pattern stays too simple. More data mainly helps high variance.