Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

What are the consequences if training data does not follow a normal distribution?


What you need to know

Separate three questions

  1. Does my model assume normality? Trees and neural nets: no. Linear regression: only the residuals, and only for inference. Gaussian NB, LDA, GMMs: yes, for the features.
  2. Do my statistics assume normality? z-score outlier rules, "68% within 1 SD", and small-sample t-tests do.
  3. Do extreme values hurt my loss? MSE squares errors, so a heavy tail lets a few points dominate training, whatever the model.

What goes wrong, concretely

  • Heavy tails and MSE. One order predicted 400 minutes off contributes 160,000 to the squared error; a thousand orders each 5 minutes off contribute 25,000. The model bends toward the one outlier.
  • Scaling. Standardisation still runs fine, but on a skewed feature most values get squashed near zero while a few get huge z-scores. Distance-based models then see little variation among typical users.
  • Wrong thresholds. A "3 SD" alert on log-normal latency fires far too often, as the code below shows.

Transform when it helps

  • Log (np.log1p for data with zeros): for right-skewed positive data like prices, counts and latency.
  • Box–Cox: searches for the power transform that makes positive data most normal. A lambda of 0 means log, 0.5 means square root.
  • Yeo–Johnson: like Box–Cox but allows zero and negative values (sklearn.preprocessing.PowerTransformer).
  • Quantile transform: maps any distribution to a normal or uniform shape by rank.
Python
import numpy as npfrom scipy.stats import skew, boxcoxrng = np.random.default_rng(0)latency = rng.lognormal(mean=5, sigma=0.8, size=100_000)   # ms, heavy right tailcut = latency.mean() + 3 * latency.std()print("share above mean + 3 SD:", round((latency > cut).mean(), 4), "(normal would give 0.0013)")transformed, lam = boxcox(latency)print("skew before:", round(skew(latency), 2), " after Box-Cox:", round(skew(transformed), 2),      " lambda:", round(lam, 2))
Text
share above mean + 3 SD: 0.0184 (normal would give 0.0013)skew before: 3.93  after Box-Cox: 0.0  lambda: 0.0

On skewed latency, a 3 SD rule flags 1.8% of requests — 14 times more than the normal rule promises. Box–Cox picks lambda = 0, a log transform, and the skew disappears.

A real-life example

A property portal predicts flat prices in Mumbai with linear regression. Prices range from ₹40 lakh to ₹40 crore, a long right tail. Trained on raw prices, the model's errors are dominated by a handful of luxury flats and it underprices ordinary 1BHKs.

The team trains on log(price) instead. Now the model predicts percentage differences — "a sea view adds about 18%" — which is how prices actually behave, and residuals look roughly normal. A gradient-boosted model trained on raw prices does about as well on ranking flats, but the log model gives interpretable coefficients for the pricing team. The choice follows the goal, not a rule that "data must be normal".

Follow-up questions to expect

  • "Do you need to transform features for a random forest?" — No. Trees split on order, so any monotonic transform such as log leaves the splits unchanged.
  • "What is a caveat of predicting log(target)?" — Converting back with exp gives something closer to the median than the mean, so totals come out too low unless you apply a correction.
  • "What if the data is bimodal?" — Transforms will not fix it; it usually means two groups are mixed, so find the variable that separates them or model them separately.