Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

What is the 68–95–99.7 rule, and where is it practically useful in ML?


Delivery times with mean 30 and SD 525 to 35 minutes: about 68 percent20 to 40 minutes: about 95 percent15 to 45 minutes: about 99.7 percentSlower than 45: about 1 order in 740
On a million normal rows, beyond 3 SD still means 2,700 legitimate values, so a flag is a reason to look, not to delete.

What you need to know

The rule with a worked example

Delivery times at one restaurant are roughly normal with mean 30 minutes and SD 5.

RangeMinutesShare of orders
mean ± 1 SD25 to 35about 68%
mean ± 2 SD20 to 40about 95%
mean ± 3 SD15 to 45about 99.7%

So an order taking 47 minutes is more than 3 SD above the mean — expected in only about 1 in 740 orders (half of the 0.27% outside ±3 SD, because only the slow side counts).

The exact values come from the normal curve:

Python
from scipy.stats import normfor k in [1, 2, 3]:    inside = norm.cdf(k) - norm.cdf(-k)    print(f"within {k} SD: {inside:.4f}   outside: {1 - inside:.4f}")
Text
within 1 SD: 0.6827   outside: 0.3173within 2 SD: 0.9545   outside: 0.0455within 3 SD: 0.9973   outside: 0.0027

norm.cdf(k) gives the share of values below k SDs. The difference between +k and -k is the share inside the band.

Practical uses

  • Outlier flags: |z| greater than 3 as a first filter.
  • Monitoring and drift: alert when today's average prediction, input value or error moves more than 3 SD from its historical baseline.
  • Control charts: the ±3 SD bands on quality dashboards.
  • Normality sanity check: if 12% of points sit outside ±2 SD instead of about 5%, the tails are heavier than normal.

Two traps

Big data makes "rare" common. 0.27% of 1,000,000 transactions is 2,700 values beyond 3 SD, even in perfectly normal data. Those are not errors.

Outliers hide themselves. A huge outlier inflates the SD, which shrinks its own z-score. With n values and SD computed by dividing by n, no z-score can exceed the square root of n - 1, however extreme the value.

Python
import numpy as nptimes = np.array([28, 30, 31, 29, 32, 30, 27, 33, 31, 29, 300])   # one broken logz = (times - times.mean()) / times.std()print("z of 300 using mean and SD:", round(z[-1], 2))median = np.median(times)mad = np.median(np.abs(times - median))robust_z = 0.6745 * (times - median) / madprint("robust z of 300           :", round(robust_z[-1], 1))
Text
z of 300 using mean and SD: 3.16robust z of 300           : 182.1

A 300-minute "delivery" among 30-minute ones is absurd, yet its ordinary z-score is only 3.16 — the maximum possible with 11 points. The robust z-score uses the median and the median absolute deviation (MAD), which the outlier cannot inflate, and it screams 182. The 0.6745 factor makes robust z comparable to ordinary z on normal data.

A real-life example

An ML platform team monitors the average fraud score their model outputs each hour. Over three months the hourly average is 0.042 with an SD of 0.004. They alert at mean ± 3 SD: below 0.030 or above 0.054. One night it reads 0.071 — more than 7 SD above. On-call finds that a mobile app release sent a new device field as null, and the model treated every null as risky. The rule turned "is this number weird?" into one line of alerting logic.

Where it breaks: the same team tried 3 SD alerts on raw API latency. Latency has a long right tail, so the alert fired dozens of times a day on normal traffic. They moved to percentile-based alerts (p99 above a fixed limit) instead.

Follow-up questions to expect

  • "What if the data is not normal?" — Use percentiles, the IQR rule (below Q1 - 1.5 × IQR or above Q3 + 1.5 × IQR), or robust z-scores based on the median and MAD.
  • "What does Chebyshev's inequality say?" — For any distribution with a finite variance, at least 75% of values are within 2 SD and at least 89% within 3 SD; weaker, but it needs no normality.
  • "Where does the 95% in a confidence interval come from?" — The sample mean is roughly normal by the CLT, and exactly 95% of a normal lies within 1.96 SD of its centre — the precise version of "about 2 SD".