Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
What is the 68–95–99.7 rule, and where is it practically useful in ML?
What you need to know
The rule with a worked example
Delivery times at one restaurant are roughly normal with mean 30 minutes and SD 5.
| Range | Minutes | Share of orders |
|---|---|---|
| mean ± 1 SD | 25 to 35 | about 68% |
| mean ± 2 SD | 20 to 40 | about 95% |
| mean ± 3 SD | 15 to 45 | about 99.7% |
So an order taking 47 minutes is more than 3 SD above the mean — expected in only about 1 in 740 orders (half of the 0.27% outside ±3 SD, because only the slow side counts).
The exact values come from the normal curve:
1from scipy.stats import norm23for k in [1, 2, 3]:4 inside = norm.cdf(k) - norm.cdf(-k)5 print(f"within {k} SD: {inside:.4f} outside: {1 - inside:.4f}")within 1 SD: 0.6827 outside: 0.3173within 2 SD: 0.9545 outside: 0.0455within 3 SD: 0.9973 outside: 0.0027norm.cdf(k) gives the share of values below k SDs. The difference between +k and -k is the share inside the band.
Practical uses
- Outlier flags: |z| greater than 3 as a first filter.
- Monitoring and drift: alert when today's average prediction, input value or error moves more than 3 SD from its historical baseline.
- Control charts: the ±3 SD bands on quality dashboards.
- Normality sanity check: if 12% of points sit outside ±2 SD instead of about 5%, the tails are heavier than normal.
Two traps
Big data makes "rare" common. 0.27% of 1,000,000 transactions is 2,700 values beyond 3 SD, even in perfectly normal data. Those are not errors.
Outliers hide themselves. A huge outlier inflates the SD, which shrinks its own z-score. With n values and SD computed by dividing by n, no z-score can exceed the square root of n - 1, however extreme the value.
1import numpy as np23times = np.array([28, 30, 31, 29, 32, 30, 27, 33, 31, 29, 300]) # one broken log4z = (times - times.mean()) / times.std()5print("z of 300 using mean and SD:", round(z[-1], 2))67median = np.median(times)8mad = np.median(np.abs(times - median))9robust_z = 0.6745 * (times - median) / mad10print("robust z of 300 :", round(robust_z[-1], 1))z of 300 using mean and SD: 3.16robust z of 300 : 182.1A 300-minute "delivery" among 30-minute ones is absurd, yet its ordinary z-score is only 3.16 — the maximum possible with 11 points. The robust z-score uses the median and the median absolute deviation (MAD), which the outlier cannot inflate, and it screams 182. The 0.6745 factor makes robust z comparable to ordinary z on normal data.
A real-life example
An ML platform team monitors the average fraud score their model outputs each hour. Over three months the hourly average is 0.042 with an SD of 0.004. They alert at mean ± 3 SD: below 0.030 or above 0.054. One night it reads 0.071 — more than 7 SD above. On-call finds that a mobile app release sent a new device field as null, and the model treated every null as risky. The rule turned "is this number weird?" into one line of alerting logic.
Where it breaks: the same team tried 3 SD alerts on raw API latency. Latency has a long right tail, so the alert fired dozens of times a day on normal traffic. They moved to percentile-based alerts (p99 above a fixed limit) instead.
Follow-up questions to expect
- "What if the data is not normal?" — Use percentiles, the IQR rule (below Q1 - 1.5 × IQR or above Q3 + 1.5 × IQR), or robust z-scores based on the median and MAD.
- "What does Chebyshev's inequality say?" — For any distribution with a finite variance, at least 75% of values are within 2 SD and at least 89% within 3 SD; weaker, but it needs no normality.
- "Where does the 95% in a confidence interval come from?" — The sample mean is roughly normal by the CLT, and exactly 95% of a normal lies within 1.96 SD of its centre — the precise version of "about 2 SD".