Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

What does variance represent in a dataset used for training ML models?


Variance of five delivery times, mean 3025-52530003552528-243224minutesminus meansquaredSquares sum to 58: divide by 5 for 11.6, or by 4 for the sample variance 14.5.
Squaring stops the deviations cancelling to zero and lets the two 5-minute misses supply 50 of the 58.

What you need to know

The idea before the formula

Two delivery partners both average 30 minutes. One always arrives in 28–32 minutes. The other arrives anywhere from 10 to 50. The mean is the same; the spread is not. Variance is one number that captures that spread.

To build it, measure how far each value is from the mean. Some distances are negative and some positive, and they always cancel to zero, so square them first. Then average the squares.

Text
population variance = sum of (x - mean)^2 / nsample variance     = sum of (x - mean)^2 / (n - 1)

Worked example by hand

Five delivery times: 25, 30, 35, 28, 32 minutes. The mean is 150 / 5 = 30.

xx - mean(x - mean)^2
25-525
3000
35525
28-24
3224
sum = 58

Population variance = 58 / 5 = 11.6. Sample variance = 58 / 4 = 14.5.

Why divide by n - 1?

When you only have a sample, you measure distances from the sample mean, which is always a bit closer to your data than the true mean is. That makes the spread look smaller than it really is. Dividing by n - 1 instead of n corrects this. It is called Bessel's correction, and it matters for small samples only.

Python
import numpy as npimport pandas as pdtimes = np.array([25, 30, 35, 28, 32])      # delivery minutesprint("numpy var (ddof=0):", times.var())print("numpy var (ddof=1):", times.var(ddof=1))print("pandas var        :", pd.Series(times).var())
Text
numpy var (ddof=0): 11.6numpy var (ddof=1): 14.5pandas var        : 14.5

NumPy and pandas use different defaults. NumPy divides by n (ddof=0) and pandas divides by n - 1 (ddof=1). This is a classic source of "why don't my numbers match?".

What variance tells you about training data

  • Near-zero variance: the feature is almost constant, such as a "country" column that is "India" for 99.99% of rows. Usually drop it (scikit-learn's VarianceThreshold does this). Check first, though: a rare binary flag with low variance can still be very predictive.
  • Very different variances across features: income in rupees (variance in the billions) next to age (variance around 100). k-NN, k-means, SVMs and gradient descent will be dominated by income until you standardise.
  • PCA keeps the directions with the most variance, on the idea that spread is where the information is.

A real-life example

A food-delivery company compares two cloud kitchens with the same mean preparation time of 15 minutes. Kitchen A has a variance of 4 (minutes squared); kitchen B has a variance of 64. Customers of kitchen B see wildly unreliable ETAs, even though "the average is fine". The operations team targets kitchen B's variance, not its mean — they standardise the menu and batch orders, and ETA complaints fall.

In the ETA model's training data, a feature "is_monsoon_month" turns out to be 0 for every row because the data only covers October to March. It has zero variance, so the model cannot learn anything from it — and in July the model will be surprised. The fix is more representative data, not a different model.

Follow-up questions to expect

  • "Why square the deviations instead of taking absolute values?" — Squares are smooth and easy to differentiate, variances of independent variables add up, and squaring links to the normal distribution. Mean absolute deviation exists and is more robust, but less convenient mathematically.
  • "What is the difference between population and sample variance?" — Population divides by n and is used when you have every value; sample divides by n - 1 to correct for estimating the mean from the same data.
  • "How does variance relate to PCA?" — PCA finds the directions of largest variance and keeps the top few, so standardise first or large-unit features win just because of their units.