Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

How do mean, median, and mode behave in skewed data commonly found in user-generated datasets?


Which way the tail pullsRight skew — minutes watched• Long tail of a few heavy users• mode 10, median 13.5, mean 45.2• Report and impute with the median• log1p cuts skew from 2.65 to 1.27Left skew — an easy exam• Long tail of a few low scores• mean below median below mode• Median is still the honest centre• Tree models ignore the skew either way
The mean follows the tail because it adds up sizes, so the gap between mean and median is a free skew detector.

What you need to know

What skew means

A distribution is skewed when one tail is longer than the other. Right skew (positive skew) has a long tail of large values. Most people watch a few minutes of video, and a few people watch for hours. Left skew (negative skew) has a long tail of small values, like exam scores on an easy test where most students score 85–95 and a few score 30.

Why the three measures separate

  • The mode sits at the peak, where most users are.
  • The median sits at the middle position. The tail adds values on one side, so the middle shifts a little toward it.
  • The mean adds up the actual sizes, so the tail's large values pull it the furthest.
Text
right skew:  mode < median < meanleft skew:   mean < median < modesymmetric:   mean ≈ median ≈ mode

This ordering is a reliable rule of thumb for smooth, single-peaked data. It can fail for discrete data or distributions with several peaks, so treat it as a check, not a law.

Python
import numpy as npfrom scipy.stats import skew# minutes watched per user on a streaming app (made-up)watch = np.array([5, 8, 10, 10, 10, 12, 15, 18, 25, 40, 90, 300])print("mean  :", round(watch.mean(), 1))print("median:", np.median(watch))print("skew  :", round(skew(watch), 2))logged = np.log1p(watch)print("skew after log1p:", round(skew(logged), 2))
Text
mean  : 45.2median: 13.5skew  : 2.65skew after log1p: 1.27

The mode is 10, the median 13.5 and the mean 45.2 — the right-skew order. One 300-minute user more than triples the mean. log1p computes log(1 + x), which is safe for zeros, and it cuts the skew in half by squeezing the long tail.

What to do about it in ML

  • Reporting: use the median and percentiles.
  • Imputation: use the median, not the mean.
  • Linear and logistic regression, k-NN, neural networks: a log or Box–Cox transform often helps, because a few huge values otherwise dominate the fit.
  • Tree models: they split on thresholds, so a monotonic transform like log changes nothing for them.

A real-life example

An e-commerce company predicts next month's spend per customer. The target is heavily right-skewed: median ₹800, mean ₹3,100, a few wholesale buyers spending ₹5 lakh. A linear regression trained on raw spend chases the wholesale buyers and gives poor predictions for everyone else.

The engineer trains on log1p(spend) and converts predictions back with expm1. Errors for typical customers drop a lot. The trade-off: the model now predicts something closer to the median than the mean, so if finance needs total revenue, a correction is needed when converting back.

Follow-up questions to expect

  • "Give an example of left skew in ML data." — Age at retirement, exam scores on an easy test, or a model's confidence scores when most predictions are near 1.0.
  • "Does skew matter for tree-based models?" — Much less. Trees split on value order, so monotonic transforms have no effect; skew mainly matters for linear, distance-based and neural models.
  • "How do you detect skew quickly?" — Compare mean and median, compute scipy.stats.skew, and plot a histogram.