Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

Why is correlation not sufficient to establish causation in AI systems?


The confounder behind a 0.72 correlationHot weatherIce-cream sales riseMore swimming,more drownings
The two children correlate at 0.72, yet after removing what temperature explains, the correlation falls to about zero.

What you need to know

The four alternative explanations

  • Confounding. A third factor causes both. Ice-cream sales and drownings rise together because hot weather drives both.
  • Reverse causation. "Customers who contact support churn more" — maybe unhappy customers contact support, rather than support causing churn.
  • Selection bias. The data only includes certain cases. If you study only apps that survived, you may conclude that some practice causes success when failed apps did it too.
  • Coincidence. Test 1,000 random features against a target and about 50 will look "significant" at the 5% level by pure chance.

Watching a confounder at work

Python
import numpy as nprng = np.random.default_rng(0)temp = rng.uniform(20, 42, 365)                          # daily max temperature, °Cice_cream = 50 * temp + rng.normal(0, 150, 365)          # cones solddrownings = 0.3 * temp + rng.normal(0, 1.5, 365)         # incidentsprint("raw correlation:", round(np.corrcoef(ice_cream, drownings)[0, 1], 2))# remove what temperature explains from each, then correlate what is leftres_ice = ice_cream - np.polyval(np.polyfit(temp, ice_cream, 1), temp)res_drown = drownings - np.polyval(np.polyfit(temp, drownings, 1), temp)print("after controlling for temperature:", round(np.corrcoef(res_ice, res_drown)[0, 1], 2))
Text
raw correlation: 0.72after controlling for temperature: -0.06

In this simulation neither causes the other; both depend only on temperature. The raw correlation is a strong 0.72. After removing the part of each explained by temperature, it vanishes. This "control for the confounder" step only works if you know and measured the confounder, which is the hard part in real data.

How to get causal evidence

  • Randomised experiments (A/B tests). Randomly assign who gets the change. Randomisation balances every confounder, known and unknown, between the groups.
  • Natural experiments and causal inference when you cannot randomise: difference-in-differences, instrumental variables, regression discontinuity, propensity-score matching. These rely on assumptions you must state and defend.

A real-life example

An e-commerce analyst finds that users who add items to a wishlist spend three times more than users who do not. The product team wants to show a big "Add to wishlist" pop-up to everyone.

The analyst warns that engaged shoppers both use wishlists and spend more: engagement is the confounder. The team runs an A/B test instead: half the users see the pop-up. Wishlist use doubles in the test group, but spend per user rises by less than 1%, which is not statistically significant. The correlation was real; the causal effect was close to zero. The pop-up is dropped before it annoyed everyone.

Follow-up questions to expect

  • "When is correlation enough?" — When you only need to predict, not intervene. A churn predictor can use any signal that forecasts churn, as long as it is available at prediction time.
  • "Why does randomisation solve confounding?" — Random assignment makes the groups the same on average in every respect except the treatment, so any outcome difference is caused by the treatment.
  • "What is Simpson's paradox?" — A trend that appears in combined data reverses inside every subgroup, usually because a confounder is distributed unevenly across groups.