Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
Why do many ML algorithms assume data is normally distributed?
What you need to know
Who assumes what
| Method | What must be (roughly) normal | What breaks if not |
|---|---|---|
| Linear regression (OLS) | The residuals, for exact small-sample inference | p-values and intervals; coefficients stay unbiased |
| Logistic regression | Nothing — errors are Bernoulli | Not applicable |
| Gaussian Naive Bayes | Each feature within each class | Probabilities and accuracy degrade |
| LDA / QDA | Features within each class | Decision boundary is off |
| Gaussian mixture models | Each cluster | Clusters split or merge wrongly |
| t-tests, z-tests | The sample mean (the CLT often provides this) | Only with small, skewed samples |
| Trees, random forests, boosting, neural nets | Nothing | Not applicable |
Why the normal is so convenient
- Two numbers describe it. Estimate the mean and variance and you have the whole distribution.
- Sums of normals are normal. Maths stays in one family, so formulas have closed forms.
- Least squares = maximum likelihood. If the errors are normal, the most likely line is the one that minimises the sum of squared errors. That is why MSE is the default regression loss.
if y = prediction + error, error ~ Normal(0, SD)then maximising the likelihood ⇔ minimising sum of (y - prediction)^2- Maximum entropy. Among all distributions with a given mean and variance, the normal assumes the least extra structure, so it is a "safe default" for noise.
Why trees and neural networks do not care
A decision tree only asks "is this feature above a threshold?" — the shape of the distribution is irrelevant. A neural network learns whatever mapping reduces the loss. Neither computes a likelihood from a normal formula, which is part of why gradient-boosted trees work so well on messy tabular data.
A real-life example
A data scientist predicts delivery time with linear regression and reports "each extra kilometre adds 2.1 minutes, p < 0.001". A reviewer asks about assumptions. The residual histogram has a long right tail from rainy-day delays. With 50,000 orders, the coefficient is still fine and the CLT keeps the standard errors close to right. With only 60 orders from a new city, the same skewed residuals would make that p-value unreliable, so the analyst would bootstrap the confidence interval instead.
Meanwhile, the team's production ETA model is gradient-boosted trees, which needs no normality at all. The linear model is kept only for explaining effects to operations managers.
Follow-up questions to expect
- "Does linear regression need normally distributed features?" — No. It needs a linear relationship, independent errors and constant error variance; normal residuals matter only for exact small-sample inference.
- "Why is MSE the default regression loss?" — It is the maximum-likelihood loss when errors are normal, and it is smooth and easy to optimise.
- "What loss would you use for heavy-tailed errors?" — MAE or Huber loss, which grow linearly for large errors, so outliers do not dominate.