Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
What is a normal distribution, and why does it frequently appear in ML datasets?
What you need to know
The shape
Most values sit near the mean. The further you go from the mean, the rarer values become, and they thin out at the same rate on both sides. Mean, median and mode are all equal at the centre.
X ~ Normal(mean, SD) e.g. exam scores ~ Normal(65, 12)standard normal: mean 0, SD 1 (this is what z-scores follow)Why it appears: the Central Limit Theorem
The CLT says: if you add up (or average) many independent random values with a finite variance, the result is approximately normal — even if the individual values are not.
Intuition. Your delivery time is the sum of many small delays: kitchen queue, traffic lights, finding parking, the lift. Each is lopsided, but when many are added, the ups and downs partly cancel, and the total piles up in the middle.
The CLT also says how narrow the average gets:
SD of a sample mean = SD of one value / square root of nHere is the CLT in action. Single delays follow a very skewed distribution. Averages of 50 of them are almost symmetric.
1import numpy as np2from scipy.stats import skew34rng = np.random.default_rng(0)5# single delivery delays: right-skewed, mean 10 minutes6delays = rng.exponential(scale=10, size=(100_000, 50))78print("skew of single delays :", round(skew(delays[:, 0]), 2))9print("skew of means of 50 delays :", round(skew(delays.mean(axis=1)), 2))10print("sd of means (theory 10/sqrt(50) = 1.41):", round(delays.mean(axis=1).std(), 2))skew of single delays : 2.03skew of means of 50 delays : 0.28sd of means (theory 10/sqrt(50) = 1.41): 1.41Skew drops from 2.0 to 0.3, and the spread of the averages matches the formula. This is why confidence intervals and A/B tests on averages work even when the raw data is skewed.
Where it shows up in ML
- Noise and errors: measurement noise and the residuals of a good regression model are often close to normal.
- Sample statistics: means, and therefore metrics like average accuracy across bootstrap samples, are approximately normal.
- Model design: weight initialisation, Gaussian noise in data augmentation and diffusion models, and the latent space of a VAE all use normal distributions on purpose.
A real-life example
A university has 5,000 students take the same exam. Each score is the result of many small factors — sleep, each topic studied, luck on specific questions — so the histogram comes out close to a bell with mean 65 and SD 12. That lets the exam board say "a score of 89 is 2 SD above the mean, roughly the top 2–3%".
Compare that with order values on a shopping app: most orders are ₹300–800, a few are ₹50,000. That histogram has a long right tail, and assuming normality would badly misjudge how often large orders happen. The CLT applies to averages of many orders, not to single orders.
Follow-up questions to expect
- "What does the CLT need to hold?" — Independent (or weakly dependent) values, a finite variance, and a large enough n; very heavy tails or strong dependence slow it down or break it.
- "How do you check whether data is normal?" — Plot a histogram and a Q-Q plot; formal tests like Shapiro–Wilk exist but flag tiny, harmless deviations on large datasets.
- "Name distributions that are not normal in ML data." — Counts (Poisson), waiting times (exponential), incomes and prices (roughly log-normal), popularity (power law).