Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

How do you compute the joint probability of independent events?


One request, three independent servicesModelendpointup: 0.99Vectordatabaseup: 0.995LLM APIup: 0.999All three up:about 0.984Only valid if they share no failure cause, such as one cloud region.
Every service you chain multiplies in another number below one, so the whole is always less available than its weakest part.

What you need to know

The intuition

If half the time A happens, and independently half the time B happens, then both happen in half of a half — a quarter — of cases. Multiplying probabilities is "taking a fraction of a fraction".

Text
P(A and B)         = P(A) × P(B)P(A and B and C)   = P(A) × P(B) × P(C)P(at least one A)  = 1 - P(no A at all)

Worked examples

A request passes through three independent services: model endpoint (99%), vector database (99.5%) and an LLM API (99.9%).

Text
P(all up) = 0.99 × 0.995 × 0.999 ≈ 0.984

So about 1.6% of requests hit at least one failure — worse than any single component. Chaining services always lowers availability.

The complement trick for "at least one". Roll a die four times. What is the chance of at least one six?

Text
P(no six in one roll)   = 5/6P(no six in four rolls) = (5/6)^4 ≈ 0.482P(at least one six)     = 1 - 0.482 = 0.518

A quick simulation to check the availability number:

Python
import numpy as nprng = np.random.default_rng(1)n = 1_000_000model_up = rng.random(n) < 0.99vector_db_up = rng.random(n) < 0.995print("formula   :", 0.99 * 0.995)print("simulated :", (model_up & vector_db_up).mean().round(4))
Text
formula   : 0.98505simulated : 0.985

Why ML adds logs instead of multiplying

A language model scores a sentence by multiplying the probabilities of each token. Multiply hundreds of small numbers and the result becomes too small for a computer to store — it rounds to zero (underflow). So we add logarithms instead, because log(a × b) = log(a) + log(b).

Python
import numpy as npp = np.full(400, 0.01)          # 400 tokens, each with probability 0.01print("product   :", np.prod(p))print("sum of logs:", np.log(p).sum().round(2))
Text
product   : 0.0sum of logs: -1842.07

This is why training uses log-likelihood and why Naive Bayes implementations sum log-probabilities.

A real-life example

A fraud team has two independent checks on a UPI payment: a device-fingerprint rule that catches 60% of fraud and a velocity rule that catches 50%. If they are truly independent, the chance a fraud slips past both is 0.4 × 0.5 = 0.2, so together they catch 80%.

In reality, both rules tend to fire on the same kind of fraud (new device plus burst of payments), so they are dependent. Measured on labelled data, the pair catches only 68%. The engineer reports the measured number and looks for a third signal that fails on different cases.

Follow-up questions to expect

  • "What if the events are not independent?" — Use the chain rule, P(A and B) = P(A) × P(B | A), with the conditional probability measured from data.
  • "How do you compute P(A or B)?" — P(A) + P(B) - P(A and B). The subtraction removes the overlap that was counted twice.
  • "Why do language models use log-probabilities?" — Products of many small probabilities underflow to zero; sums of logs stay in a safe numeric range and are easier to optimise.