Course Content
Mathematics for Machine Learning
5 sections · 13 lessons
Hypothesis Testing — Deciding If a Difference Is Real
You have two versions of a model. On a held-out test set of 2,000 examples, model A gets 86.9% accuracy and model B gets 87.3%. Model B wins. You ship model B.
Except that a 0.4 percentage point gap on 2,000 examples is eight examples. Eight. If eight borderline cases had gone the other way, the ranking would flip. So the real question is not "which number is bigger" — you can see that — but "would this ranking survive if I had drawn a different test set from the same distribution?"
You can put a number on that. The uncertainty in an accuracy estimate from n examples is roughly p(1−p)/n. With p≈0.87 and n=2000, that is 0.87×0.13/2000=0.0075, or about 0.75 percentage points. Each of your two accuracy figures wobbles by about three quarters of a point just from the luck of which examples landed in the test set. The gap between them is 0.4 points. It is comfortably inside the noise.
Hypothesis testing is the machinery for making that argument precisely, for a wide range of situations, instead of by eye. It answers one question: is this difference larger than what chance alone would routinely produce?
The framework
Every test starts by writing down two competing statements about the world.
The null hypothesis H0 is the boring one: there is no effect, no difference, nothing going on. The observed gap is entirely chance. The alternative hypothesis H1 is the interesting one: there really is a difference.
The null is always the "nothing happening" statement, and that asymmetry is deliberate. The null is a specific, complete claim, so you can compute what data it would produce. "There is a difference" is vague — a difference of how much? — and you cannot compute anything from it. So the strategy is always: assume the boring thing is true, work out how surprising your data would be under that assumption, and if it is surprising enough, abandon the boring thing.
You never prove the alternative. You only find the null too implausible to keep.
The procedure
- State H0 and H1 before looking at the data.
- Choose a significance level α — conventionally 0.05 — which is the false-alarm rate you are willing to tolerate.
- Compute a test statistic: a single number summarising how far the data sits from what the null predicts, measured in units of its own noise.
- Compute the p-value: the probability, if the null were true, of getting a test statistic at least this extreme.
- If p<α, reject H0. Otherwise, fail to reject it.
What a p-value is, and the four things it is not
A p-value is the answer to exactly one question: supposing there is genuinely no effect, how often would I see data at least this extreme? A p-value of 0.03 means that a world with no real difference would produce a gap this big or bigger 3% of the time.
| People say | Actually |
|---|---|
| "p=0.03, so there's a 3% chance the null is true" | No. It is the probability of the data given the null, not the probability of the null given the data. Those are different quantities, and swapping them is the base-rate fallacy. |
| "p=0.03, so there's a 97% chance my result is real" | No. That number does not appear anywhere in the calculation. |
| "p=0.001, so the effect is large" | No. A p-value conflates effect size with sample size. A trivial effect becomes highly significant with enough data. |
| "p=0.09, so there is no effect" | No. Failing to reject is not evidence of absence. It usually means you did not have enough data to detect anything. |
Two ways to be wrong
| H0 is actually true | H0 is actually false | |
|---|---|---|
| You reject H0 | Type I error (false positive), rate α | Correct, probability 1−β (the power) |
| You keep H0 | Correct | Type II error (false negative), rate β |
A Type I error means shipping a change that does nothing. A Type II error means discarding a change that would have worked. The two trade off: lowering α to 0.01 makes false alarms rarer but makes you miss more real effects. The only way to reduce both at once is to gather more data.
Power 1−β is the probability of detecting an effect that is genuinely there. It depends on the true effect size, the sample size, the noise level, and α. Working out the sample size you need before running the experiment is the single most valuable statistical habit, because a test with 30% power will fail to detect a real improvement seven times out of ten, and you will conclude "no effect" when the honest conclusion is "no idea".
The t-test: comparing means
The t-test asks whether an observed mean, or a difference between two means, is larger than the sampling noise would explain. Every version has the same shape:
A t of 3 means the effect is three times bigger than its own uncertainty. A t of 0.4 means it is swamped by noise.
One sample against a claimed value
Here xˉ is the sample mean, μ0 is the value the null claims, s is the sample standard deviation, and n is the sample size. The denominator s/n is the standard error — how much the sample mean itself bounces around from sample to sample. Note the n: quadrupling your data halves the noise, which is why improvements get expensive.
Worked example. Your service is supposed to respond in 200 ms on average. You sample 25 requests and get xˉ=214 ms with s=30 ms. Is it genuinely slower, or did you sample a bad 25?
H0:μ=200. The standard error is 30/25=30/5=6 ms. So
The observed 14 ms overshoot is 2.33 standard errors from the claim. With n=25 there are n−1=24 degrees of freedom, and the two-sided critical value at α=0.05 is 2.064. Since 2.33>2.064, reject: the service really is slower than 200 ms. The two-sided p-value is about 0.028.
Degrees of freedom deserve a word. You used the same 25 numbers to estimate both the mean and the standard deviation. Once the mean is fixed, only 24 of the deviations are free to vary — the last is determined. That lost degree of freedom is why s slightly underestimates the true spread and why the t-distribution has fatter tails than the normal for small samples.
Two independent samples
Two model variants evaluated on separate data. Variant A: n=30, mean F1 of 0.812, s=0.045. Variant B: n=30, mean F1 of 0.834, s=0.052.
Welch's version, which does not assume the two groups have equal variance:
Computing the standard error: 300.0452=0.0000675 and 300.0522=0.0000901. Their sum is 0.0001576, whose square root is 0.01256. Then
With roughly 57 degrees of freedom the critical value is about 2.00, so 1.75 does not clear the bar. The p-value is about 0.085. B looks better, but not convincingly so on this evidence.
| Student's pooled t-test | Welch's t-test | |
|---|---|---|
| Assumes equal variances? | Yes | No |
| Cost when variances are equal | — | Negligible loss of power |
| Cost when variances differ | Error rates become wrong, sometimes badly | — |
| Recommendation | Only when you have a real reason | The default; SciPy uses it unless told otherwise |
Paired data — the one that catches practitioners out
Suppose you evaluate both models with 10-fold cross-validation on the same dataset. You now have ten pairs of scores, and within each pair both models saw exactly the same data. Fold 3 might be intrinsically hard, dragging both scores down; fold 7 might be easy, lifting both.
Treating those twenty numbers as two independent samples of ten throws away the pairing and buries the real signal under between-fold variation. The paired t-test instead computes the ten differences and tests whether their mean is zero. Fold difficulty cancels out entirely, and the test becomes far more sensitive.
If the same data produced both numbers, the samples are paired. Using an independent-samples test there is throwing away most of your statistical power.
There is a caveat even so: cross-validation folds overlap in their training sets, so the differences are not fully independent of each other, and a naive paired t-test on CV folds is known to be optimistic. Corrected variants exist, and for high-stakes comparisons a held-out test set with a proper interval is safer than any CV-fold test.
Assumptions, and what to do when they fail
| Assumption | How it breaks | Alternative |
|---|---|---|
| Roughly normal data (or large n) | Heavy skew with a small sample | Mann-Whitney U (independent), Wilcoxon signed-rank (paired) |
| Independent observations | Repeated measures, time series, users appearing twice | Paired test, or a mixed-effects model |
| No extreme outliers | One enormous value dominates the mean and inflates s | Trimmed means, rank-based tests, or a bootstrap |
The normality requirement is milder than people fear. Because sample means tend towards a normal distribution as n grows regardless of the shape of the underlying data, a t-test on 50 or more observations is fairly robust to non-normality. Small samples of visibly skewed data are where it genuinely misleads.
Confidence intervals say more than p-values
A p-value gives you a binary verdict. A confidence interval gives you a range of plausible values, which is almost always the more useful object.
For the F1 comparison above: the difference is 0.022, the standard error 0.01256, and the critical value 2.00. So the 95% interval is
This says far more than "p=0.085, not significant". It says the true improvement could plausibly be anywhere from slightly negative to nearly 5 points. That is a useful, honest statement: the experiment simply was not big enough to pin down the effect, and you should either collect more data or accept the uncertainty explicitly.
Interpreting it correctly. A 95% confidence interval does not mean "there is a 95% probability the true value is in this interval". The true value is a fixed number; it either is or is not in your interval. What is random is the interval, because it was built from a random sample. The correct statement is about the procedure: if you repeated this experiment many times, 95% of the intervals you construct would contain the true value.
The bridge to hypothesis testing is exact: a 95% confidence interval for a difference excludes zero precisely when the corresponding two-sided test rejects at α=0.05. Our interval contains zero; our test failed to reject. Same information, but the interval also tells you the size of the effect you are uncertain about.
Other tests you will meet
Chi-square: are these categories related?
where O is what you observed and E is what the null predicts. Each term asks how far a cell deviates from expectation, scaled by how big that expectation was — being 15 off from an expected 35 is far more striking than being 15 off from an expected 1000.
Worked example. Does your classifier perform differently for two user segments?
| Correct | Incorrect | Total | |
|---|---|---|---|
| Segment X | 180 | 20 | 200 |
| Segment Y | 150 | 50 | 200 |
| Total | 330 | 70 | 400 |
Under the null that segment and correctness are unrelated, each cell's expected count is (row total x column total) / grand total. For "X and correct": 200×330/400=165. For "X and incorrect": 200×70/400=35. By symmetry the Y row expects the same.
With (2−1)(2−1)=1 degree of freedom the critical value at α=0.05 is 3.84. We got 15.58, so reject firmly: the model genuinely performs worse for segment Y. That is a fairness finding, not a statistical curiosity, and it is exactly the kind of thing a single aggregate accuracy number hides.
The same statistic in goodness-of-fit form checks whether observed counts match a claimed distribution. If your model should output four classes uniformly on 400 examples (100 each) and you observe 90, 110, 85, 115, then χ2=(100+100+225+225)/100=6.5 on 3 degrees of freedom, against a critical value of 7.81 — consistent with uniform.
ANOVA: more than two groups
With three model variants, the tempting move is three pairwise t-tests: A vs B, A vs C, B vs C. Do not. Each test carries its own 5% false-positive risk, and running several inflates the overall risk well past 5%.
One-way ANOVA tests all groups at once via an F-statistic that compares variation between group means against variation within groups. If the between-group variation is large relative to the within-group noise, at least one group differs. A significant F tells you that something differs but not what; you then run pairwise comparisons with a correction such as Tukey's.
Testing a correlation
A sample correlation of r=0.3 from 10 points is unremarkable; the same r from 1000 points is overwhelming evidence of a real relationship. The test statistic t=r(n−2)/(1−r2) formalises that, on n−2 degrees of freedom. Note that "statistically significant correlation" and "useful correlation" are different things — with a million rows, r=0.01 is significant and explains 0.01% of the variance.
The mistakes that quietly invalidate results
Testing many things at once
Run 20 independent tests at α=0.05 on data where nothing is real. The probability of at least one false positive is
Nearly 64%. Comparing 20 hyperparameter settings and celebrating the one with p<0.05 is, more likely than not, celebrating noise. Two standard corrections:
- Bonferroni: test each at α/m. For 20 tests, use 0.0025. Simple and strictly correct, but conservative — it will make you miss real effects when m is large.
- Benjamini-Hochberg: controls the false discovery rate, the expected proportion of your rejections that are false, rather than the probability of any false rejection. Sort the p-values ascending and find the largest i with p(i)≤miα; reject everything up to it. Much more powerful for large-scale screening.
Peeking
Checking an A/B test every hour and stopping when p first dips below 0.05 will produce a "significant" result on pure noise with near-certainty if you check often enough. The p-value wanders randomly and will eventually dip below any threshold. Fix the sample size in advance, or use a sequential testing method designed for continuous monitoring.
Confusing significance with importance
With ten million rows, a 0.001% accuracy improvement will have p<0.001. It is real. It is also worthless. Always report the effect size and its confidence interval alongside the p-value, and decide on the basis of the effect size.
Choosing the hypothesis after seeing the data
Deciding to run a one-sided test because the difference happened to go that way doubles your false-positive rate. Deciding which subgroup to analyse after noticing it looks promising does far worse. Write the hypothesis down first.
Running these tests
1import numpy as np2from scipy import stats34rng = np.random.default_rng(0)5a = rng.normal(0.812, 0.045, 30) # variant A scores6b = rng.normal(0.834, 0.052, 30) # variant B scores78# Independent samples, unequal variances (Welch) -- the safe default.9t, p = stats.ttest_ind(a, b, equal_var=False)10print(f"Welch t={t:.3f} p={p:.4f}")1112# Paired: same folds evaluated by both models.13fold_a = np.array([0.81, 0.79, 0.84, 0.82, 0.80, 0.83, 0.78, 0.85, 0.81, 0.82])14fold_b = np.array([0.83, 0.81, 0.86, 0.83, 0.82, 0.85, 0.80, 0.87, 0.84, 0.84])15t, p = stats.ttest_rel(fold_a, fold_b)16print(f"paired t={t:.3f} p={p:.5f}") # far smaller p than the unpaired version1718# Contingency table: is accuracy independent of segment?19table = np.array([[180, 20], [150, 50]])20chi2, p, dof, expected = stats.chi2_contingency(table, correction=False)21print(f"chi2={chi2:.2f} dof={dof} p={p:.6f}")2223# Three or more groups at once.24c = rng.normal(0.828, 0.048, 30)25print("ANOVA:", stats.f_oneway(a, b, c))2627# Bootstrap CI -- no distributional assumption at all.28def bootstrap_ci(x, y, n_boot=10000, alpha=0.05):29 diffs = np.empty(n_boot)30 for i in range(n_boot):31 xs = rng.choice(x, size=len(x), replace=True)32 ys = rng.choice(y, size=len(y), replace=True)33 diffs[i] = ys.mean() - xs.mean()34 lo, hi = np.percentile(diffs, [100 * alpha / 2, 100 * (1 - alpha / 2)])35 return lo, hi3637lo, hi = bootstrap_ci(a, b)38print(f"95% bootstrap CI for the difference: ({lo:.4f}, {hi:.4f})")The bootstrap deserves particular attention. It makes no assumption about the shape of the data. You resample your own observations with replacement thousands of times, recompute the statistic each time, and read the interval straight off the percentiles of the resulting distribution. It works for medians, for F1 scores, for AUC, for any statistic where the analytic sampling distribution is unknown or intractable — which covers most of the metrics you actually care about in machine learning.
Using this on your own results
Go back to the 86.9% versus 87.3%. Now you can answer properly. Bootstrap the test set: resample 2,000 predictions with replacement, recompute both accuracies and their difference, repeat 10,000 times, and read off the 2.5th and 97.5th percentiles of the difference. If that interval comfortably excludes zero, the improvement is real. If it straddles zero — which, at this sample size and gap, it almost certainly does — you have not yet earned the right to say B is better.
Three habits follow from this, and they matter more than knowing which test to reach for.
First, report intervals rather than point estimates. "87.3%" invites false precision. "87.3% (95% CI: 85.8–88.8%)" tells the reader exactly how much to trust it, and makes it obvious when two models are indistinguishable.
Second, work out the sample size you need before you run the comparison. If you want to reliably detect a 1-point improvement in accuracy around 87%, the standard error must be well under half a point, which means tens of thousands of test examples. Discovering afterwards that your test set was never large enough to answer your question is a painful and entirely avoidable waste.
Third, treat every model comparison as an experiment with a pre-registered plan. Decide the metric, the test, and the sample size in advance. The moment you start choosing the analysis after seeing the numbers, the p-values stop meaning what they claim to mean — and the whole apparatus, carefully built to protect you from fooling yourself, quietly stops working.