Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
What common sampling biases can negatively impact AI models and lead to misleading conclusions?
What you need to know
The biases, each with an ML example
| Bias | What happens | ML example |
|---|---|---|
| Selection | The sample comes from a different population | A credit model trained only on approved applicants, then used on all applicants |
| Survivorship | Only "survivors" of some filter are visible | Studying only apps still on the store to learn what makes apps succeed |
| Under-coverage | A group is barely represented | A speech model trained mostly on metro accents performs worse on rural speakers |
| Temporal | Training period differs from serving period | A demand model trained on non-festival months, used during Diwali |
| Label | Labels reflect past human decisions | A hiring model learns the preferences of past recruiters |
| Self-selection / non-response | Only some people choose to respond | App-store ratings come mostly from delighted or furious users |
Selection bias in credit has its own name, the reject inference problem: you never observe whether rejected applicants would have repaid, so the model only learns about people the old system already liked.
Survivorship bias has a famous historical example. In the Second World War, analysts studied returning bombers to decide where to add armour. Statistician Abraham Wald pointed out that the planes hit in other places never came back — so the armour belonged where the surviving planes had no holes.
Why overall metrics hide it
A model can look good overall while failing a group that is small in the test set.
1import pandas as pd23results = pd.DataFrame({4 "device": ["android"] * 900 + ["ios"] * 100,5 "correct": [1] * 855 + [0] * 45 + [1] * 70 + [0] * 30,6})7print("overall accuracy:", results["correct"].mean())8print(results.groupby("device")["correct"].agg(["mean", "size"]))overall accuracy: 0.925 mean sizedeviceandroid 0.95 900ios 0.70 10092.5% overall hides 70% on iOS. If iOS is 10% of the evaluation data but 40% of the users who will actually see the model, real-world accuracy will be much lower than reported.
How to reduce sampling bias
- Sample from the serving population, including traffic the old system filtered out (for example, a small randomised "explore" slice).
- Stratify by the groups that matter so each has enough rows.
- Split by time for anything time-dependent, and include seasonal events.
- Report metrics per segment — region, device, language, age band — not only the average.
- Reweight samples to match the known population mix when you cannot re-collect.
- Document how the data was collected, so the next team knows its limits.
A real-life example
A food-delivery app trains a "will this user reorder?" model using survey answers from an in-app pop-up. Only 4% of users answer, mostly frequent customers who like the app. The model predicts high reorder rates for everyone and the marketing team under-spends on win-back offers.
The fix is to stop using survey responders as the population. The team labels reorder behaviour directly from order logs for a random sample of all users, stratified by city and order frequency, and checks accuracy per segment. Predicted reorder rates for occasional users drop sharply, and win-back campaigns are rebuilt around them.
A spam-filter version of the same trap: training only on emails users reported as spam misses the spam that was never reported, so the filter is weakest exactly on the subtle spam people did not notice.
Follow-up questions to expect
- "How do you detect sampling bias?" — Compare the distribution of key features in training data with production data; large differences in region, device or time mean the sample does not match who you serve.
- "Can more data fix sampling bias?" — Not if it comes through the same biased process; you need different data or reweighting, not more of the same.
- "How is this related to fairness?" — Under-coverage and label bias are the most common routes by which models perform worse for, or discriminate against, particular groups.