Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

Why do many ML algorithms assume data is normally distributed?


What you need to know

Who assumes what

MethodWhat must be (roughly) normalWhat breaks if not
Linear regression (OLS)The residuals, for exact small-sample inferencep-values and intervals; coefficients stay unbiased
Logistic regressionNothing — errors are BernoulliNot applicable
Gaussian Naive BayesEach feature within each classProbabilities and accuracy degrade
LDA / QDAFeatures within each classDecision boundary is off
Gaussian mixture modelsEach clusterClusters split or merge wrongly
t-tests, z-testsThe sample mean (the CLT often provides this)Only with small, skewed samples
Trees, random forests, boosting, neural netsNothingNot applicable

Why the normal is so convenient

  • Two numbers describe it. Estimate the mean and variance and you have the whole distribution.
  • Sums of normals are normal. Maths stays in one family, so formulas have closed forms.
  • Least squares = maximum likelihood. If the errors are normal, the most likely line is the one that minimises the sum of squared errors. That is why MSE is the default regression loss.
Text
if  y = prediction + error,   error ~ Normal(0, SD)then maximising the likelihood  ⇔  minimising  sum of (y - prediction)^2
  • Maximum entropy. Among all distributions with a given mean and variance, the normal assumes the least extra structure, so it is a "safe default" for noise.

Why trees and neural networks do not care

A decision tree only asks "is this feature above a threshold?" — the shape of the distribution is irrelevant. A neural network learns whatever mapping reduces the loss. Neither computes a likelihood from a normal formula, which is part of why gradient-boosted trees work so well on messy tabular data.

A real-life example

A data scientist predicts delivery time with linear regression and reports "each extra kilometre adds 2.1 minutes, p < 0.001". A reviewer asks about assumptions. The residual histogram has a long right tail from rainy-day delays. With 50,000 orders, the coefficient is still fine and the CLT keeps the standard errors close to right. With only 60 orders from a new city, the same skewed residuals would make that p-value unreliable, so the analyst would bootstrap the confidence interval instead.

Meanwhile, the team's production ETA model is gradient-boosted trees, which needs no normality at all. The linear model is kept only for explaining effects to operations managers.

Follow-up questions to expect

  • "Does linear regression need normally distributed features?" — No. It needs a linear relationship, independent errors and constant error variance; normal residuals matter only for exact small-sample inference.
  • "Why is MSE the default regression loss?" — It is the maximum-likelihood loss when errors are normal, and it is smooth and easy to optimise.
  • "What loss would you use for heavy-tailed errors?" — MAE or Huber loss, which grow linearly for large errors, so outliers do not dominate.