Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

What does a correlation coefficient close to zero imply for feature selection?


Age against churn risk, a U-shapeWhat the coefficients say• Pearson r: -0.014• Spearman: -0.015• Both measure one direction only• Verdict: nothing hereWhat is really there• Young and old customers churn• Mutual information: 0.787• A tree splits both ends off• Verdict: keep the feature
A straight line through a U is flat, so a zero correlation can hide the strongest feature in the table.

What you need to know

Why r can be zero for a strong relationship

Pearson r only measures how well a straight line fits. Imagine customer age and churn risk: very young customers churn (they switch for offers) and very old customers churn (they find the app hard), while middle-aged customers stay. On a scatter plot this is a clear U. A straight line through a U is flat, so r ≈ 0.

Python
import numpy as npfrom scipy.stats import pearsonr, spearmanrfrom sklearn.feature_selection import mutual_info_regressionrng = np.random.default_rng(0)age = rng.uniform(18, 70, 2_000)churn_risk = (age - 44) ** 2 / 100 + rng.normal(0, 1, 2_000)   # U-shapeprint("pearson r :", round(pearsonr(age, churn_risk)[0], 3))print("spearman  :", round(spearmanr(age, churn_risk)[0], 3))print("mutual info:", round(mutual_info_regression(age.reshape(-1, 1), churn_risk, random_state=0)[0], 3))
Text
pearson r : -0.014spearman  : -0.015mutual info: 0.787

Both Pearson and Spearman say "nothing here", because the relationship is not monotonic either. Mutual information, which detects any kind of dependence, is clearly positive. A decision tree would simply split young and old customers away from the middle-aged ones and use this feature heavily.

Other reasons a zero-correlation feature can matter

  • Interactions. "Is weekend" might have zero correlation with delivery time on its own, but combined with "is raining" it predicts long delays. Tree models and neural networks find such combinations.
  • Categorical features. City, device type or payment method encoded as numbers 1, 2, 3 give a meaningless Pearson r. Use group means, Cramér's V or MI.
  • Rare but decisive values. A flag that is on for 0.1% of rows can have a tiny r and still identify most of the fraud.

Better ways to decide

  • Plot the feature against the target.
  • Use mutual information or Spearman as a screen.
  • Use permutation importance: shuffle the feature and measure how much validation performance drops.
  • Train with and without the feature and compare on a validation set.

A near-zero variance is a stronger reason to drop a feature than a near-zero correlation, though even then check rare flags.

A real-life example

An AC maker predicts service calls from outdoor temperature. The correlation across the year is close to zero, and a junior analyst suggests removing temperature. Plotting it shows a U: calls peak in the hottest weeks (units overworked) and in the coldest weeks (units switched on after months unused, with dust and gas issues). A gradient-boosted model using temperature cuts forecast error substantially compared with the model without it. The feature was essential; the statistic was the wrong tool.

Follow-up questions to expect

  • "Can Spearman catch a U-shape?" — No. Spearman catches monotonic relationships only; a U goes down then up, so its rank correlation is also near zero.
  • "What is permutation importance?" — Shuffle one feature's values in the validation set and measure how much the model's score drops; a big drop means the model relies on that feature.
  • "Does zero correlation mean independence?" — No. Independence implies zero correlation, but zero correlation does not imply independence, as the U-shape shows.