Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
How is probability formally defined, and why is it fundamental to AI models?
What you need to know
The intuition
Probability is a way to put a number on uncertainty. 0 means "impossible", 1 means "certain", and everything between is "how sure". A weather app that says "70% chance of rain" means: on days that look like today, it rained about 7 times out of 10.
The formal rules (Kolmogorov's axioms)
Every probability you will ever use follows three rules.
1. P(A) >= 0 for any event A2. P(all possible outcomes) = 13. If A and B cannot both happen, P(A or B) = P(A) + P(B)Everything else follows from these. For example, the complement rule: P(not A) = 1 - P(A).
Worked example. Roll a fair die. The event "even" has 3 favourable outcomes out of 6, so P(even) = 3/6 = 0.5. P(not even) = 1 - 0.5 = 0.5.
Two ways to read a probability
- Frequentist: the long-run fraction of times something happens if you repeat it. "The model's precision is 0.9" means 9 in 10 flagged items were right over many items.
- Bayesian: a degree of belief given the information you have, updated as data arrives. "I am 80% sure this email is spam."
ML uses both. Evaluation metrics are frequentist counts; model outputs and Bayesian methods are beliefs.
Why AI models are built on it
A neural classifier produces raw scores called logits. Softmax turns them into numbers that follow the axioms: each is positive and they sum to 1.
1import numpy as np23logits = np.array([2.0, 1.0, 0.1]) # raw scores for spam, promo, normal4probs = np.exp(logits) / np.exp(logits).sum()5print(probs.round(3), "sum =", probs.sum().round(3))[0.659 0.242 0.099] sum = 1.0The output looks like a probability, but it is only trustworthy if the model is calibrated: when it says 0.66, it should be right about 66% of the time. Many modern networks are overconfident, so teams check calibration and fix it with temperature scaling or similar methods.
Beyond classification, probability is the core of language models (next-token distributions), generative models, Naive Bayes, uncertainty estimates, and loss functions like cross-entropy, which is the negative log of the probability the model gave to the correct answer.
A real-life example
A spam filter scores an email 0.92. The product decides: above 0.9, send to spam; between 0.5 and 0.9, show a warning banner; below 0.5, deliver normally. None of this works if 0.92 is just an arbitrary score. The team checks calibration on last month's mail: of all emails scored 0.9–0.95, 93% were actually spam. The scores behave like probabilities, so the thresholds mean what the product team thinks they mean.
Follow-up questions to expect
- "Is a softmax output a true probability?" — It satisfies the axioms mathematically, but it only matches real-world frequencies if the model is calibrated; check with a reliability diagram.
- "What is cross-entropy in probability terms?" — The negative log of the probability the model assigned to the true class, averaged over examples. Confident wrong answers cost the most.
- "Frequentist or Bayesian?" — Frequentist treats probability as long-run frequency and parameters as fixed; Bayesian treats parameters as uncertain and updates a belief with data.