Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Implement A/B testing for prompts or models.
What you need to know
Offline evals tell you a new prompt is not broken; an online A/B test tells you whether real users are better off. Three parts:
- Assignment. Each user is placed in variant A or B. It must be sticky (the same user always gets the same variant, or a conversation flips mid-way) and random with respect to the metric. Hashing
experiment:user_idgives both, and needs no database. Including the experiment name means a user in B for one experiment is not automatically in B for every experiment. - Metrics. One primary metric decided in advance (task success, thumbs-up rate), plus guardrail metrics that must not get worse (p95 latency, cost per request, escalation rate).
- Significance. A difference in rates might be noise. The two-proportion z-test asks how many standard errors apart the two rates are:
p_a = wins_a / n_a, p_b = wins_b / n_bp = (wins_a + wins_b) / (n_a + n_b) pooled ratese = sqrt(p × (1 − p) × (1/n_a + 1/n_b))z = (p_b − p_a) / setwo-sided p-value = 1 − erf(|z| / sqrt(2))A p-value below 0.05 is the usual bar. Peeking — checking every day and stopping the first time p dips under 0.05 — makes false winners far more likely, because you are giving noise many chances to look significant.
1import hashlib, math2from collections import defaultdict34def assign(user_id: str, experiment: str, variants: list[str],5 weights: list[float] | None = None) -> str:6 """Sticky, stateless assignment: the same user always gets the same variant."""7 digest = hashlib.sha256(f"{experiment}:{user_id}".encode()).hexdigest()8 point = int(digest[:8], 16) / 2**32 # uniform in [0, 1)9 weights = weights or [1 / len(variants)] * len(variants)10 cumulative = 0.011 for variant, weight in zip(variants, weights):12 cumulative += weight13 if point < cumulative:14 return variant15 return variants[-1] # guards against rounding in the weights1617class ExperimentLog:18 def __init__(self) -> None:19 self.n, self.wins = defaultdict(int), defaultdict(int)20 self.latency, self.usd = defaultdict(list), defaultdict(float)2122 def record(self, variant: str, success: bool, latency_ms: float, usd: float) -> None:23 self.n[variant] += 124 self.wins[variant] += bool(success)25 self.latency[variant].append(latency_ms)26 self.usd[variant] += usd2728def two_proportion_test(wins_a: int, n_a: int, wins_b: int, n_b: int) -> tuple[float, float]:29 """(z, two-sided p-value) for 'B's success rate differs from A's'."""30 if n_a == 0 or n_b == 0:31 return 0.0, 1.032 pooled = (wins_a + wins_b) / (n_a + n_b)33 se = math.sqrt(pooled * (1 - pooled) * (1 / n_a + 1 / n_b))34 if se == 0:35 return 0.0, 1.036 z = (wins_b / n_b - wins_a / n_a) / se37 return z, 1 - math.erf(abs(z) / math.sqrt(2))The tricky parts:
/ 2**32, not/ 0xFFFFFFFF: dividing by 2³² keeps the point strictly below 1, so it always falls inside the last bucket.1 − erf(|z| / √2)is the two-sided p-value using only the standard library:erfgives the normal distribution's area within |z| standard deviations.se == 0happens when both variants have 0% or both 100% success; there is no evidence either way.
Complexity: assign is one hash, O(length of the ids), plus O(variants). record is O(1). The test is O(1) from the counters. Memory is O(requests) only if you keep raw latencies; keep a histogram instead at scale.
A real-life example
1print(assign("user_42", "refund_prompt_v2", ["A", "B"]) ==2 assign("user_42", "refund_prompt_v2", ["A", "B"])) # True: sticky3split = [assign(f"user_{i}", "refund_prompt_v2", ["A", "B"]) for i in range(10_000)]4print(split.count("A"), split.count("B")) # 4954 504656z, p = two_proportion_test(wins_a=120, n_a=1000, wins_b=150, n_b=1000)7print(round(z, 3), round(p, 4)) # 1.963 0.0496| quantity | value |
|---|---|
| p_a, p_b | 0.120, 0.150 |
| pooled p | 270 / 2000 = 0.135 |
| se | sqrt(0.135 × 0.865 × 0.002) = 0.01528 |
| z | 0.030 / 0.01528 = 1.963 |
| p-value | 0.0496 — just under 0.05 |
The 10,000 synthetic users split 4,954 / 5,046, within one percent of 50/50. The test result is the teaching point: a 3-point lift on 1,000 users each is barely significant. Had the team checked daily and stopped the moment p crossed 0.05, a result this marginal is exactly what peeking produces by chance.
A quick-commerce app testing a new substitution-suggestion prompt runs this for a fixed two weeks, checks that cost per order and complaint rate did not rise, and only then ships variant B.
Follow-up questions to expect
- "How many users do you need?" — Run a power calculation before starting: detecting a 3-point lift from a 12% base at 80% power and 5% significance needs roughly 2,000 users per variant. With less traffic, rely on offline evals.
- "What if B is better but costs twice as much?" — That is what guardrail metrics are for; the decision weighs quality against cost and latency, agreed before the test.
- "How do you test a risky change safely?" — Shadow mode: serve A, run B in the background on the same requests, and compare B's outputs offline with no user exposure.