Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

What is correlation, and how is it quantified in ML workflows?


What you need to know

The intuition

When study hours are above average, is the exam score usually above average too? If yes, the two move together: positive correlation. If high hours go with low scores, negative correlation. If knowing one tells you nothing about the other in a straight-line way, the correlation is near zero.

The formula

Covariance multiplies each pair of deviations from the mean. Dividing by both standard deviations removes the units, so r is always between -1 and +1.

Text
r = sum of (x - mean_x)(y - mean_y) / square root of [sum of (x - mean_x)^2 × sum of (y - mean_y)^2]

Worked example by hand

Five students: study hours 1, 2, 3, 4, 5 and exam scores 52, 55, 61, 64, 68. Means are 3 and 60.

Text
x - mean_x:   -2, -1,  0,  1,  2y - mean_y:   -8, -5,  1,  4,  8sum of products        = 16 + 5 + 0 + 4 + 16 = 41sum of (x - mean_x)^2  = 4 + 1 + 0 + 1 + 4   = 10sum of (y - mean_y)^2  = 64 + 25 + 1 + 16 + 64 = 170r = 41 / square root of (10 × 170) = 41 / 41.23 ≈ 0.994
Python
import pandas as pddf = pd.DataFrame({    "study_hours": [1, 2, 3, 4, 5],    "exam_score":  [52, 55, 61, 64, 68],})print(round(df["study_hours"].corr(df["exam_score"]), 3))                     # Pearsonprint(round(df["study_hours"].corr(df["exam_score"], method="spearman"), 3))  # Spearman
Text
0.9941.0

Spearman is exactly 1 because the ranks match perfectly: more hours always means a higher score. Pearson is slightly below 1 because the points are not on a perfect straight line. On a whole DataFrame, df.corr() gives every pair at once.

Which coefficient to use

MeasureCapturesUse when
PearsonStraight-line relationshipsNumeric data, roughly linear, few outliers
SpearmanAny monotonic relationship (always up or always down)Curved but one-directional, or outliers
KendallMonotonic, based on pairsSmall samples, many ties
Cramér's V / mutual informationAny associationCategorical variables, non-linear patterns

How it is used in ML

  • Feature screening: which features move with the target?
  • Redundancy: pairs with |r| above about 0.9 carry nearly the same information; keep one, or combine them.
  • Leakage check: a feature with r = 0.99 to the target is suspicious — it may be computed from the target.
  • Always plot. Anscombe's quartet is four small datasets with almost identical r (about 0.82) and completely different shapes, including a curve and a single outlier creating the whole correlation.

A real-life example

A food-delivery team explores its ETA data. Correlation of delivery time with distance: r = 0.78. With number of items: r = 0.12. With "restaurant rating": r = -0.05. They also find "distance_km" and "distance_road_km" have r = 0.97, so they keep only the road distance.

Then one feature, "minutes_since_order_placed_at_delivery", shows r = 0.99 with delivery time. It is the target in disguise, recorded after the delivery. Correlation analysis caught a leak before it reached training.

Follow-up questions to expect

  • "What is the difference between covariance and correlation?" — Covariance has units and its size depends on scale; correlation divides by both SDs, so it is unit-free and bounded between -1 and +1.
  • "Is r = 0.5 strong?" — It depends on the field; r squared = 0.25 means a straight line explains 25% of the variance in the other variable.
  • "How do outliers affect Pearson r?" — A single extreme point can create or destroy a correlation; Spearman, which uses ranks, is far more robust.