Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
When is the median a better measure than the mean in real-world ML datasets?
What you need to know
How to find the median
Sort the values. If the count is odd, the median is the middle one. If the count is even, it is the mean of the two middle values.
odd n: median = value at position (n + 1) / 2even n: median = mean of the values at positions n/2 and n/2 + 1Worked example. Seven delivery times in minutes: 22, 25, 27, 30, 31, 35, 180. The middle (4th) value is 30. The mean is 350 / 7 = 50. One stuck order, 180 minutes, moved the mean by 20 minutes but did not move the median at all.
When the median wins
- Skewed data. Income, prices, order values, session length, items per user. A few very large values create a long right tail.
- Data with errors. A sensor that sometimes reports 99999, or a typo that adds a zero.
- Latency and SLAs. Users feel the typical and the worst cases, not the average. Teams track p50 (the median), p95 and p99.
- Imputation. Filling blanks with the median keeps the centre of a skewed column where it was.
When the mean is still right
The mean is better when the data is roughly symmetric, when you need a total (total revenue = mean × count), or when the maths needs it — loss functions and gradients are built on means. For a symmetric distribution the two are almost equal anyway.
1import numpy as np23orders = np.array([400, 450, 500, 520, 550, 600, 50_000]) # rupees4print("mean :", round(orders.mean(), 1))5print("median:", np.median(orders))6print("mean without the bulk order:", round(orders[:-1].mean(), 1))mean : 7574.3median: 520.0mean without the bulk order: 503.3Removing one order changes the mean from 7,574 to 503. The median barely notices. Also note the percentile call for latency work: np.percentile(latencies, [50, 95, 99]) gives p50, p95 and p99 in one line.
A real-life example
A food-delivery app shows "average delivery time: 34 minutes". Most orders arrive in 25–30 minutes, but on rainy evenings a few take over two hours. Those few pull the mean up to 34, which is too pessimistic for most customers and hides how bad the bad cases are.
The team switches the dashboard to "median 28 minutes, p95 52 minutes". Now the typical experience and the tail are both visible, and an alert on p95 catches the rainy-evening problem that the mean smoothed over.
For the ETA model's training data, the same team fills missing "distance to restaurant" values with the median distance. Using the mean would have pushed every filled value toward the few very long-distance orders.
Follow-up questions to expect
- "What are p95 and p99?" — The values below which 95% or 99% of observations fall. p99 latency of 800 ms means 1 request in 100 is slower than 800 ms.
- "Can you average medians?" — No. The median of a whole dataset is not the mean of the medians of its parts. Keep raw data or histograms if you need to combine groups.
- "Why not always use the median?" — It ignores the size of values beyond their order, cannot be summed into totals, and is harder to optimise directly, so means stay central to training.