Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

Why is sampling necessary when working with large-scale AI datasets?


Estimating mean delivery time from 5 million orders10033.53.11,00031.61.010,00031.80.3rows sampledestimate, min95 percent marginTrue mean of all 5 million: 32.0 minutes.
The margin shrinks with the square root of the sample, not with the size of the table you sampled from.

What you need to know

Why a small sample can be enough

To estimate an average, what matters is how many rows you sample, not how many exist. The uncertainty of a sample mean is its standard error:

Text
standard error of a mean       = SD / square root of nstandard error of a proportion = square root of [p × (1 - p) / n]95% margin of error            ≈ 1.96 × standard error

The margin shrinks with the square root of n: four times the data halves the error. The population size barely appears, as long as it is much larger than the sample.

Worked example: estimating delivery time from 5 million orders

Python
import numpy as nprng = np.random.default_rng(0)all_orders = rng.gamma(shape=4, scale=8, size=5_000_000)   # 5 million delivery timesprint("true mean (all 5M)  :", round(all_orders.mean(), 2))for n in [100, 1_000, 10_000]:    sample = rng.choice(all_orders, size=n, replace=False)    se = sample.std(ddof=1) / np.sqrt(n)    print(f"sample of {n:>6}: mean = {sample.mean():.2f}  ± {1.96 * se:.2f}")
Text
true mean (all 5M)  : 32.0sample of    100: mean = 33.48  ± 3.07sample of   1000: mean = 31.63  ± 1.01sample of  10000: mean = 31.81  ± 0.31

With 10,000 rows — 0.2% of the data — the estimate is within a third of a minute, and the true value sits inside every stated range.

Worked example: how many conversations to label?

You want to measure a chatbot's answer accuracy to within ±5 percentage points at 95% confidence. Use p = 0.5, the worst case:

Text
n = 1.96^2 × 0.5 × 0.5 / 0.05^2 ≈ 384

About 400 randomly sampled conversations are enough, whether the bot had 50,000 or 50 million conversations.

Common sampling methods

  • Simple random: every row has an equal chance. The default.
  • Stratified: sample separately within groups (class, region, device) so each is represented in the right proportion. Essential for rare classes.
  • Systematic: every k-th row. Simple, but dangerous if the data has a repeating pattern.
  • Cluster: sample whole groups, such as whole stores or whole users. Cheaper, but less precise.
  • Reservoir sampling: keep a fixed-size uniform sample from a stream whose length you do not know in advance.
  • Deliberate re-balancing: under-sample the majority class or over-sample the minority for training — but evaluate on the true distribution.

Stratification matters most with rare classes:

Python
import numpy as npfrom sklearn.model_selection import train_test_splity = np.array([1] * 20 + [0] * 980)          # 1,000 payments, 2% fraudX = np.arange(len(y)).reshape(-1, 1)for name, strat in [("random", None), ("stratified", y)]:    counts = []    for seed in range(8):        _, _, _, y_test = train_test_split(X, y, test_size=100, random_state=seed, stratify=strat)        counts.append(int(y_test.sum()))    print(f"{name:<10} fraud cases in each 100-row test set: {counts}")
Text
random     fraud cases in each 100-row test set: [2, 2, 1, 1, 1, 3, 4, 1]stratified fraud cases in each 100-row test set: [2, 2, 2, 2, 2, 2, 2, 2]

Random splits give anywhere from 1 to 4 fraud cases, so the recall you measure jumps around by chance. Stratified splits always keep the 2% rate.

A real-life example

A railway-booking company has 3 billion search logs and wants to test 20 feature ideas for a demand model. Training on all of it takes 14 hours per run. The team trains on a 2% stratified sample by route and month, gets results in 20 minutes, shortlists the three best ideas, and only then trains on the full data. Iteration went from two ideas a day to twenty.

The trap they avoided: an earlier attempt sampled "the last 30 days" because it was easy to query. It contained no festival season, and the model had no idea what Diwali traffic looks like. The sample was large but not representative.

Follow-up questions to expect

  • "When should you not sample?" — When you need exact totals (billing, finance), when you are hunting very rare events that a sample might miss, or when the full data is cheap to process.
  • "What is the difference between over-sampling and SMOTE?" — Over-sampling duplicates minority rows; SMOTE creates synthetic minority rows by interpolating between neighbours. Both change the class balance, so probabilities need recalibrating afterwards.
  • "How does sample size affect a confidence interval?" — Its width shrinks with the square root of n, so quadrupling the sample halves the interval.