Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
What are the consequences if training data does not follow a normal distribution?
What you need to know
Separate three questions
- Does my model assume normality? Trees and neural nets: no. Linear regression: only the residuals, and only for inference. Gaussian NB, LDA, GMMs: yes, for the features.
- Do my statistics assume normality? z-score outlier rules, "68% within 1 SD", and small-sample t-tests do.
- Do extreme values hurt my loss? MSE squares errors, so a heavy tail lets a few points dominate training, whatever the model.
What goes wrong, concretely
- Heavy tails and MSE. One order predicted 400 minutes off contributes 160,000 to the squared error; a thousand orders each 5 minutes off contribute 25,000. The model bends toward the one outlier.
- Scaling. Standardisation still runs fine, but on a skewed feature most values get squashed near zero while a few get huge z-scores. Distance-based models then see little variation among typical users.
- Wrong thresholds. A "3 SD" alert on log-normal latency fires far too often, as the code below shows.
Transform when it helps
- Log (
np.log1pfor data with zeros): for right-skewed positive data like prices, counts and latency. - Box–Cox: searches for the power transform that makes positive data most normal. A lambda of 0 means log, 0.5 means square root.
- Yeo–Johnson: like Box–Cox but allows zero and negative values (
sklearn.preprocessing.PowerTransformer). - Quantile transform: maps any distribution to a normal or uniform shape by rank.
1import numpy as np2from scipy.stats import skew, boxcox34rng = np.random.default_rng(0)5latency = rng.lognormal(mean=5, sigma=0.8, size=100_000) # ms, heavy right tail67cut = latency.mean() + 3 * latency.std()8print("share above mean + 3 SD:", round((latency > cut).mean(), 4), "(normal would give 0.0013)")910transformed, lam = boxcox(latency)11print("skew before:", round(skew(latency), 2), " after Box-Cox:", round(skew(transformed), 2),12 " lambda:", round(lam, 2))share above mean + 3 SD: 0.0184 (normal would give 0.0013)skew before: 3.93 after Box-Cox: 0.0 lambda: 0.0On skewed latency, a 3 SD rule flags 1.8% of requests — 14 times more than the normal rule promises. Box–Cox picks lambda = 0, a log transform, and the skew disappears.
A real-life example
A property portal predicts flat prices in Mumbai with linear regression. Prices range from ₹40 lakh to ₹40 crore, a long right tail. Trained on raw prices, the model's errors are dominated by a handful of luxury flats and it underprices ordinary 1BHKs.
The team trains on log(price) instead. Now the model predicts percentage differences — "a sea view adds about 18%" — which is how prices actually behave, and residuals look roughly normal. A gradient-boosted model trained on raw prices does about as well on ranking flats, but the log model gives interpretable coefficients for the pricing team. The choice follows the goal, not a rule that "data must be normal".
Follow-up questions to expect
- "Do you need to transform features for a random forest?" — No. Trees split on order, so any monotonic transform such as log leaves the splits unchanged.
- "What is a caveat of predicting log(target)?" — Converting back with exp gives something closer to the median than the mean, so totals come out too low unless you apply a correction.
- "What if the data is bimodal?" — Transforms will not fix it; it usually means two groups are mixed, so find the variable that separates them or model them separately.