Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

How does high variance in data affect model training and generalization?


Two meanings of high varianceSpread in the input data• A feature spans a huge range• Mixed scales makegradient descent zig-zag• One column dominates k-NN distances• Fix: standardise or min-max scaleVariance of the model• Predictions swing with the sample• Deep tree: train 1.00, test 0.81• It has memorised the label noise• Fix: limit depth, regularise, bag
The same word names two problems with different fixes, so say which one you mean before you answer.

What you need to know

Meaning 1: spread in the input data

A wide-ranging feature is not a problem by itself. The trouble is mixing scales. If "income" ranges over lakhs of rupees and "age" over tens of years:

  • Gradient descent takes a step size that is too big for one weight and too small for the other, so it zig-zags and converges slowly.
  • k-NN, k-means and SVMs compute distances, and income differences swamp age differences.
  • Regularisation penalises coefficients unevenly, because a coefficient's size depends on the feature's units.

The fix is standardising (z-scores) or min-max scaling. Tree models do not need it, because they only compare values to thresholds.

Large variance in the target, especially noise the features cannot explain, sets a floor on the error any model can reach. That part is called irreducible error.

Meaning 2: variance of the model

The bias–variance trade-off splits expected error into three parts:

Text
expected error = bias^2 + variance + irreducible noise

A very simple model has high bias (it misses the pattern) and low variance. A very flexible model has low bias and high variance. You want the depth or complexity that minimises the total.

Python
from sklearn.datasets import make_classificationfrom sklearn.model_selection import train_test_splitfrom sklearn.tree import DecisionTreeClassifierX, y = make_classification(n_samples=1000, n_features=20, n_informative=5,                           flip_y=0.1, random_state=42)X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=42)for depth in [None, 4]:    tree = DecisionTreeClassifier(max_depth=depth, random_state=42).fit(X_tr, y_tr)    print(f"max_depth={depth}: train={tree.score(X_tr, y_tr):.2f}  test={tree.score(X_te, y_te):.2f}")
Text
max_depth=None: train=1.00  test=0.81max_depth=4: train=0.88  test=0.85

The unlimited tree memorises the training set, including the 10% of labels that flip_y made noisy, and scores 100% on it. It drops to 81% on new data. Limiting depth to 4 lowers training accuracy but raises test accuracy: less variance, better generalisation.

Ways to reduce model variance

  • More training data, or data augmentation.
  • Regularisation: L1/L2 penalties, dropout, weight decay.
  • A simpler model: shallower trees, fewer features.
  • Bagging and random forests, which average many high-variance trees.
  • Early stopping when validation loss starts rising.

A real-life example

Think of a student who memorises last year's exam paper word for word. They score 100% on that paper and fail a new one. That is high model variance. Where the analogy breaks: a model does not "know" it is memorising; it simply has enough capacity to fit every quirk, so you must measure the gap on held-out data.

Production case: a delivery ETA model built with deep gradient-boosted trees scores an MAE of 1.5 minutes on training orders and 7 minutes on last week's orders. The engineer limits tree depth, adds a minimum number of samples per leaf, and uses early stopping on a validation set. Training MAE rises to 3.5 minutes; last week's MAE falls to 4.2 minutes.

Follow-up questions to expect

  • "How do you tell high variance from high bias?" — High variance: training error low, validation error much higher. High bias: both errors high and close together.
  • "Why does bagging reduce variance?" — Averaging many models trained on different bootstrap samples cancels out their individual, partly independent errors.
  • "Does more data fix high bias?" — Usually not; a model too simple for the pattern stays too simple. More data mainly helps high variance.