Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

What is the difference between independent and dependent events in ML pipelines?


What you need to know

The intuition

Flip a coin twice. The first flip does not change the second, so the flips are independent. Now draw two cards from a deck without putting the first back. If the first card was an ace, there are fewer aces left, so the second draw is dependent on the first.

Text
independent:  P(A and B) = P(A) × P(B)dependent:    P(A and B) = P(A) × P(B | A)

P(B | A) is read "probability of B given A" — the next section covers it in depth.

Worked example. Chance of drawing two aces:

Text
with replacement (independent):     4/52 × 4/52 = 1/169  ≈ 0.0059without replacement (dependent):    4/52 × 3/51 = 12/2652 ≈ 0.0045

Where it matters in a pipeline

  • Train/test splits. Random splitting assumes rows are independent. Rows from the same user, patient, device or document family are dependent, because they share hidden traits. The model can "recognise" the user instead of learning the pattern. This is a form of data leakage.
  • Time series. Today's sales depend on yesterday's. A random split lets the model peek at the future. Split by time: train on the past, test on later data.
  • Model assumptions. Naive Bayes assumes features are independent given the class. Many statistical tests assume independent observations; with repeated measures from the same user, their p-values are too optimistic.

The fix for grouped data is a group-aware split:

Python
import numpy as npfrom sklearn.model_selection import GroupShuffleSplitemails = np.arange(10)                          # 10 emailscampaign = np.array([1, 1, 1, 2, 2, 3, 3, 3, 4, 4])   # which campaign sent each onesplitter = GroupShuffleSplit(n_splits=1, test_size=0.3, random_state=0)train_idx, test_idx = next(splitter.split(emails, groups=campaign))print("train campaigns:", sorted(set(campaign[train_idx].tolist())))print("test campaigns :", sorted(set(campaign[test_idx].tolist())))
Text
train campaigns: [1, 2]test campaigns : [3, 4]

Every campaign lands entirely in train or entirely in test, so the test set measures performance on campaigns the model has never seen. GroupKFold does the same for cross-validation.

A real-life example

A team builds a spam filter. Spammers send the same message with tiny changes to thousands of addresses. With a random split, near-copies of each campaign sit in both train and test. The filter scores 98% on test. In production it catches far less, because each new campaign is genuinely new.

After switching to a split grouped by campaign, the offline score drops to 91% — and that number matches what production sees. The lower number is the honest one.

An everyday version: a teacher sets a test using questions the class already practised word for word. The scores are high, but they measure memory, not understanding. The questions and the practice were dependent.

Follow-up questions to expect

  • "How do you test whether two events are independent?" — Check whether P(A and B) is close to P(A) × P(B), or whether P(B | A) is close to P(B). For categorical data, a chi-square test of independence does this formally.
  • "What is leakage?" — Information in training that would not be available at prediction time, including overlap between train and test through dependent rows; it makes offline scores look better than reality.
  • "Is uncorrelated the same as independent?" — No. Correlation only measures linear relationships; two variables can be uncorrelated and still strongly dependent, such as x and x squared for x centred on zero.