Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
How do you compute the joint probability of independent events?
What you need to know
The intuition
If half the time A happens, and independently half the time B happens, then both happen in half of a half — a quarter — of cases. Multiplying probabilities is "taking a fraction of a fraction".
P(A and B) = P(A) × P(B)P(A and B and C) = P(A) × P(B) × P(C)P(at least one A) = 1 - P(no A at all)Worked examples
A request passes through three independent services: model endpoint (99%), vector database (99.5%) and an LLM API (99.9%).
P(all up) = 0.99 × 0.995 × 0.999 ≈ 0.984So about 1.6% of requests hit at least one failure — worse than any single component. Chaining services always lowers availability.
The complement trick for "at least one". Roll a die four times. What is the chance of at least one six?
P(no six in one roll) = 5/6P(no six in four rolls) = (5/6)^4 ≈ 0.482P(at least one six) = 1 - 0.482 = 0.518A quick simulation to check the availability number:
1import numpy as np23rng = np.random.default_rng(1)4n = 1_000_0005model_up = rng.random(n) < 0.996vector_db_up = rng.random(n) < 0.99578print("formula :", 0.99 * 0.995)9print("simulated :", (model_up & vector_db_up).mean().round(4))formula : 0.98505simulated : 0.985Why ML adds logs instead of multiplying
A language model scores a sentence by multiplying the probabilities of each token. Multiply hundreds of small numbers and the result becomes too small for a computer to store — it rounds to zero (underflow). So we add logarithms instead, because log(a × b) = log(a) + log(b).
1import numpy as np23p = np.full(400, 0.01) # 400 tokens, each with probability 0.014print("product :", np.prod(p))5print("sum of logs:", np.log(p).sum().round(2))product : 0.0sum of logs: -1842.07This is why training uses log-likelihood and why Naive Bayes implementations sum log-probabilities.
A real-life example
A fraud team has two independent checks on a UPI payment: a device-fingerprint rule that catches 60% of fraud and a velocity rule that catches 50%. If they are truly independent, the chance a fraud slips past both is 0.4 × 0.5 = 0.2, so together they catch 80%.
In reality, both rules tend to fire on the same kind of fraud (new device plus burst of payments), so they are dependent. Measured on labelled data, the pair catches only 68%. The engineer reports the measured number and looks for a third signal that fails on different cases.
Follow-up questions to expect
- "What if the events are not independent?" — Use the chain rule, P(A and B) = P(A) × P(B | A), with the conditional probability measured from data.
- "How do you compute P(A or B)?" — P(A) + P(B) - P(A and B). The subtraction removes the overlap that was counted twice.
- "Why do language models use log-probabilities?" — Products of many small probabilities underflow to zero; sums of logs stay in a safe numeric range and are easier to optimise.