Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
How is standard deviation interpreted in the context of feature distribution?
What you need to know
From variance back to real units
Variance in the delivery example was 14.5 minutes squared, which is hard to picture. Take the square root and you get 3.8 minutes. Now you can say: a typical delivery is about 4 minutes away from the 30-minute average.
SD = square root of variancez = (x - mean) / SDHow to read it
- Relative to the mean. An SD of 5 minutes on a 30-minute mean is modest. An SD of 5 minutes on a 6-minute mean is huge. The ratio SD / mean is the coefficient of variation, useful for comparing spread across features with different scales.
- As a ruler. If a feature is roughly bell-shaped, about two-thirds of values fall within 1 SD of the mean and about 95% within 2 SD. For skewed data this rule does not hold, so check a histogram first.
- As a z-score. A z-score says how many SDs a value is from its mean. It lets you compare values measured on different scales.
Worked example: z-scores on exam results
A student scores 80 in Maths, where the class mean is 60 and the SD is 10. They score 82 in Physics, where the mean is 75 and the SD is 5.
Maths: z = (80 - 60) / 10 = 2.0Physics: z = (82 - 75) / 5 = 1.4The Physics mark is higher in raw points but the Maths result is more exceptional relative to the class. This is exactly what standardisation does to features before a model sees them: it puts everything on the same "how unusual is this" scale.
A real-life example
Two batters both average 40 runs over their last seven innings.
1import numpy as np23batter_a = np.array([35, 42, 38, 45, 40, 36, 44]) # steady4batter_b = np.array([0, 5, 110, 2, 90, 8, 65]) # boom or bust5for name, runs in [("A", batter_a), ("B", batter_b)]:6 print(name, "mean:", round(runs.mean(), 1), " sd:", round(runs.std(ddof=1), 1))A mean: 40.0 sd: 3.9B mean: 40.0 sd: 47.1A selector who only looks at the mean sees two identical players. The SD shows that batter A reliably gives you about 40 and batter B gives you either a big score or almost nothing. Neither is "better" — it depends on whether the team needs stability or match-winning upside.
The ML version: a demand-forecasting feature "orders per hour" has a mean of 120 and an SD of 3 at one store and an SD of 60 at another. The second store's demand is far harder to predict, so its forecast needs wider error bands and more safety stock.
Follow-up questions to expect
- "What does an SD of zero mean?" — Every value is identical, so the feature is constant. Standardising it would divide by zero; drop it.
- "How do you flag outliers with SD?" — Mark values with |z| above 3. This works for roughly normal data; for skewed data, use the IQR or median absolute deviation, because the outliers themselves inflate the SD.
- "Why standardise features?" — So features on large scales do not dominate distance calculations and gradient steps, and so regularisation penalises every coefficient fairly.