Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

How do mutually exclusive and independent events differ, and why does it matter in modeling?


Same three logits, two output layersSoftmax — classes exclusive• action 0.60, comedy 0.37, horror 0.03• Always sums to exactly 1• Labels compete for one unit• Cannot say both, cannot say neitherSigmoid per label• action 0.88, comedy 0.82, horror 0.27• Sums to 1.97 here• Each label scored on its own• Can say both, or neither
The output layer is an assumption about the world: softmax insists exactly one label is true.

What you need to know

The intuition

One coin flip: "heads" and "tails" are mutually exclusive — the same flip cannot be both. Two separate flips: "first is heads" and "second is heads" are independent — both can happen, and one does not affect the other.

Text
mutually exclusive:  P(A and B) = 0          P(A or B) = P(A) + P(B)independent:         P(A and B) = P(A) × P(B)

Why they conflict. Suppose P(A) = 0.3 and P(B) = 0.5. If they were independent, P(A and B) would be 0.15. If they are mutually exclusive, it is 0. Both cannot be true unless one of the events has probability 0.

Mutually exclusive

  • Cannot happen together
  • P(A and B) = 0
  • Knowing A tells you B is false
  • Example: one email is "spam" or "not spam"

Independent

  • Can happen together
  • P(A and B) = P(A) × P(B)
  • Knowing A tells you nothing about B
  • Example: two separate coin flips

Why it matters in modelling

The choice of output layer encodes this assumption.

  • Multi-class, single label (a digit is 0–9, a ticket goes to exactly one team): the classes are mutually exclusive and cover all cases. Use softmax, which forces probabilities to sum to 1, with categorical cross-entropy.
  • Multi-label (a movie can be action and comedy): labels can co-occur. Use a sigmoid per label, each giving its own probability from 0 to 1, with binary cross-entropy per label.
Python
import numpy as nplogits = np.array([2.0, 1.5, -1.0])       # action, comedy, horrorsoftmax = np.exp(logits) / np.exp(logits).sum()sigmoid = 1 / (1 + np.exp(-logits))print("softmax:", softmax.round(2), "sum =", softmax.sum().round(2))print("sigmoid:", sigmoid.round(2), "sum =", sigmoid.sum().round(2))
Text
softmax: [0.6  0.37 0.03] sum = 1.0sigmoid: [0.88 0.82 0.27] sum = 1.97

With softmax, "action" and "comedy" compete for the same probability mass, so the model cannot say both are likely. With sigmoids, it can say 0.88 action and 0.82 comedy at once. Note that separate sigmoids do not assume the labels are independent in the data; they just stop the output layer from forcing a choice.

A real-life example

A news app tags articles with topics. The first version used softmax over "cricket", "business" and "politics". An article about a cricket board's broadcasting deal is about both cricket and business, but the model had to split one unit of probability between them, so it gave 0.5 and 0.45, and neither passed the 0.6 display threshold. The article got no tag.

Switching to one sigmoid per topic let the model output 0.9 cricket and 0.8 business. The single-label assumption — mutual exclusivity — was wrong for the problem, and no amount of training could fix it.

Follow-up questions to expect

  • "Can two events be both mutually exclusive and independent?" — Only if at least one has probability zero. Otherwise P(A and B) would have to be 0 and positive at the same time.
  • "What are collectively exhaustive events?" — Events that between them cover every possible outcome. Softmax classes are assumed mutually exclusive and collectively exhaustive, which is why an "other" class is often needed.
  • "How do you handle 'none of the above' in multi-label?" — With sigmoids it is natural: every label can be below threshold. With softmax you would need an explicit "none" class.