Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
How do mean, median, and mode behave in skewed data commonly found in user-generated datasets?
What you need to know
What skew means
A distribution is skewed when one tail is longer than the other. Right skew (positive skew) has a long tail of large values. Most people watch a few minutes of video, and a few people watch for hours. Left skew (negative skew) has a long tail of small values, like exam scores on an easy test where most students score 85–95 and a few score 30.
Why the three measures separate
- The mode sits at the peak, where most users are.
- The median sits at the middle position. The tail adds values on one side, so the middle shifts a little toward it.
- The mean adds up the actual sizes, so the tail's large values pull it the furthest.
right skew: mode < median < meanleft skew: mean < median < modesymmetric: mean ≈ median ≈ modeThis ordering is a reliable rule of thumb for smooth, single-peaked data. It can fail for discrete data or distributions with several peaks, so treat it as a check, not a law.
1import numpy as np2from scipy.stats import skew34# minutes watched per user on a streaming app (made-up)5watch = np.array([5, 8, 10, 10, 10, 12, 15, 18, 25, 40, 90, 300])6print("mean :", round(watch.mean(), 1))7print("median:", np.median(watch))8print("skew :", round(skew(watch), 2))910logged = np.log1p(watch)11print("skew after log1p:", round(skew(logged), 2))mean : 45.2median: 13.5skew : 2.65skew after log1p: 1.27The mode is 10, the median 13.5 and the mean 45.2 — the right-skew order. One 300-minute user more than triples the mean. log1p computes log(1 + x), which is safe for zeros, and it cuts the skew in half by squeezing the long tail.
What to do about it in ML
- Reporting: use the median and percentiles.
- Imputation: use the median, not the mean.
- Linear and logistic regression, k-NN, neural networks: a log or Box–Cox transform often helps, because a few huge values otherwise dominate the fit.
- Tree models: they split on thresholds, so a monotonic transform like log changes nothing for them.
A real-life example
An e-commerce company predicts next month's spend per customer. The target is heavily right-skewed: median ₹800, mean ₹3,100, a few wholesale buyers spending ₹5 lakh. A linear regression trained on raw spend chases the wholesale buyers and gives poor predictions for everyone else.
The engineer trains on log1p(spend) and converts predictions back with expm1. Errors for typical customers drop a lot. The trade-off: the model now predicts something closer to the median than the mean, so if finance needs total revenue, a correction is needed when converting back.
Follow-up questions to expect
- "Give an example of left skew in ML data." — Age at retirement, exam scores on an easy test, or a model's confidence scores when most predictions are near 1.0.
- "Does skew matter for tree-based models?" — Much less. Trees split on value order, so monotonic transforms have no effect; skew mainly matters for linear, distance-based and neural models.
- "How do you detect skew quickly?" — Compare mean and median, compute
scipy.stats.skew, and plot a histogram.