Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
Why is sampling necessary when working with large-scale AI datasets?
What you need to know
Why a small sample can be enough
To estimate an average, what matters is how many rows you sample, not how many exist. The uncertainty of a sample mean is its standard error:
standard error of a mean = SD / square root of nstandard error of a proportion = square root of [p × (1 - p) / n]95% margin of error ≈ 1.96 × standard errorThe margin shrinks with the square root of n: four times the data halves the error. The population size barely appears, as long as it is much larger than the sample.
Worked example: estimating delivery time from 5 million orders
1import numpy as np23rng = np.random.default_rng(0)4all_orders = rng.gamma(shape=4, scale=8, size=5_000_000) # 5 million delivery times56print("true mean (all 5M) :", round(all_orders.mean(), 2))7for n in [100, 1_000, 10_000]:8 sample = rng.choice(all_orders, size=n, replace=False)9 se = sample.std(ddof=1) / np.sqrt(n)10 print(f"sample of {n:>6}: mean = {sample.mean():.2f} ± {1.96 * se:.2f}")true mean (all 5M) : 32.0sample of 100: mean = 33.48 ± 3.07sample of 1000: mean = 31.63 ± 1.01sample of 10000: mean = 31.81 ± 0.31With 10,000 rows — 0.2% of the data — the estimate is within a third of a minute, and the true value sits inside every stated range.
Worked example: how many conversations to label?
You want to measure a chatbot's answer accuracy to within ±5 percentage points at 95% confidence. Use p = 0.5, the worst case:
n = 1.96^2 × 0.5 × 0.5 / 0.05^2 ≈ 384About 400 randomly sampled conversations are enough, whether the bot had 50,000 or 50 million conversations.
Common sampling methods
- Simple random: every row has an equal chance. The default.
- Stratified: sample separately within groups (class, region, device) so each is represented in the right proportion. Essential for rare classes.
- Systematic: every k-th row. Simple, but dangerous if the data has a repeating pattern.
- Cluster: sample whole groups, such as whole stores or whole users. Cheaper, but less precise.
- Reservoir sampling: keep a fixed-size uniform sample from a stream whose length you do not know in advance.
- Deliberate re-balancing: under-sample the majority class or over-sample the minority for training — but evaluate on the true distribution.
Stratification matters most with rare classes:
1import numpy as np2from sklearn.model_selection import train_test_split34y = np.array([1] * 20 + [0] * 980) # 1,000 payments, 2% fraud5X = np.arange(len(y)).reshape(-1, 1)67for name, strat in [("random", None), ("stratified", y)]:8 counts = []9 for seed in range(8):10 _, _, _, y_test = train_test_split(X, y, test_size=100, random_state=seed, stratify=strat)11 counts.append(int(y_test.sum()))12 print(f"{name:<10} fraud cases in each 100-row test set: {counts}")random fraud cases in each 100-row test set: [2, 2, 1, 1, 1, 3, 4, 1]stratified fraud cases in each 100-row test set: [2, 2, 2, 2, 2, 2, 2, 2]Random splits give anywhere from 1 to 4 fraud cases, so the recall you measure jumps around by chance. Stratified splits always keep the 2% rate.
A real-life example
A railway-booking company has 3 billion search logs and wants to test 20 feature ideas for a demand model. Training on all of it takes 14 hours per run. The team trains on a 2% stratified sample by route and month, gets results in 20 minutes, shortlists the three best ideas, and only then trains on the full data. Iteration went from two ideas a day to twenty.
The trap they avoided: an earlier attempt sampled "the last 30 days" because it was easy to query. It contained no festival season, and the model had no idea what Diwali traffic looks like. The sample was large but not representative.
Follow-up questions to expect
- "When should you not sample?" — When you need exact totals (billing, finance), when you are hunting very rare events that a sample might miss, or when the full data is cheap to process.
- "What is the difference between over-sampling and SMOTE?" — Over-sampling duplicates minority rows; SMOTE creates synthetic minority rows by interpolating between neighbours. Both change the class balance, so probabilities need recalibrating afterwards.
- "How does sample size affect a confidence interval?" — Its width shrinks with the square root of n, so quadrupling the sample halves the interval.