Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Implement A/B testing for prompts or models.


Task success for two prompts100012012.0%100015015.0%userssuccessesratevariant Avariant Bz = 1.963, two-sided p = 0.0496: only just under 0.05.
A three-point lift on a thousand users each is barely significant — exactly the result that daily peeking manufactures.

What you need to know

Offline evals tell you a new prompt is not broken; an online A/B test tells you whether real users are better off. Three parts:

  • Assignment. Each user is placed in variant A or B. It must be sticky (the same user always gets the same variant, or a conversation flips mid-way) and random with respect to the metric. Hashing experiment:user_id gives both, and needs no database. Including the experiment name means a user in B for one experiment is not automatically in B for every experiment.
  • Metrics. One primary metric decided in advance (task success, thumbs-up rate), plus guardrail metrics that must not get worse (p95 latency, cost per request, escalation rate).
  • Significance. A difference in rates might be noise. The two-proportion z-test asks how many standard errors apart the two rates are:
Text
p_a = wins_a / n_a,   p_b = wins_b / n_bp   = (wins_a + wins_b) / (n_a + n_b)                 pooled ratese  = sqrt(p × (1 − p) × (1/n_a + 1/n_b))z   = (p_b − p_a) / setwo-sided p-value = 1 − erf(|z| / sqrt(2))

A p-value below 0.05 is the usual bar. Peeking — checking every day and stopping the first time p dips under 0.05 — makes false winners far more likely, because you are giving noise many chances to look significant.

Python
import hashlib, mathfrom collections import defaultdictdef assign(user_id: str, experiment: str, variants: list[str],           weights: list[float] | None = None) -> str:    """Sticky, stateless assignment: the same user always gets the same variant."""    digest = hashlib.sha256(f"{experiment}:{user_id}".encode()).hexdigest()    point = int(digest[:8], 16) / 2**32                # uniform in [0, 1)    weights = weights or [1 / len(variants)] * len(variants)    cumulative = 0.0    for variant, weight in zip(variants, weights):        cumulative += weight        if point < cumulative:            return variant    return variants[-1]                                # guards against rounding in the weightsclass ExperimentLog:    def __init__(self) -> None:        self.n, self.wins = defaultdict(int), defaultdict(int)        self.latency, self.usd = defaultdict(list), defaultdict(float)    def record(self, variant: str, success: bool, latency_ms: float, usd: float) -> None:        self.n[variant] += 1        self.wins[variant] += bool(success)        self.latency[variant].append(latency_ms)        self.usd[variant] += usddef two_proportion_test(wins_a: int, n_a: int, wins_b: int, n_b: int) -> tuple[float, float]:    """(z, two-sided p-value) for 'B's success rate differs from A's'."""    if n_a == 0 or n_b == 0:        return 0.0, 1.0    pooled = (wins_a + wins_b) / (n_a + n_b)    se = math.sqrt(pooled * (1 - pooled) * (1 / n_a + 1 / n_b))    if se == 0:        return 0.0, 1.0    z = (wins_b / n_b - wins_a / n_a) / se    return z, 1 - math.erf(abs(z) / math.sqrt(2))

The tricky parts:

  • / 2**32, not / 0xFFFFFFFF: dividing by 2³² keeps the point strictly below 1, so it always falls inside the last bucket.
  • 1 − erf(|z| / √2) is the two-sided p-value using only the standard library: erf gives the normal distribution's area within |z| standard deviations.
  • se == 0 happens when both variants have 0% or both 100% success; there is no evidence either way.

Complexity: assign is one hash, O(length of the ids), plus O(variants). record is O(1). The test is O(1) from the counters. Memory is O(requests) only if you keep raw latencies; keep a histogram instead at scale.

A real-life example

Python
print(assign("user_42", "refund_prompt_v2", ["A", "B"]) ==      assign("user_42", "refund_prompt_v2", ["A", "B"]))          # True: stickysplit = [assign(f"user_{i}", "refund_prompt_v2", ["A", "B"]) for i in range(10_000)]print(split.count("A"), split.count("B"))                        # 4954 5046z, p = two_proportion_test(wins_a=120, n_a=1000, wins_b=150, n_b=1000)print(round(z, 3), round(p, 4))                                   # 1.963 0.0496
quantityvalue
p_a, p_b0.120, 0.150
pooled p270 / 2000 = 0.135
sesqrt(0.135 × 0.865 × 0.002) = 0.01528
z0.030 / 0.01528 = 1.963
p-value0.0496 — just under 0.05

The 10,000 synthetic users split 4,954 / 5,046, within one percent of 50/50. The test result is the teaching point: a 3-point lift on 1,000 users each is barely significant. Had the team checked daily and stopped the moment p crossed 0.05, a result this marginal is exactly what peeking produces by chance.

A quick-commerce app testing a new substitution-suggestion prompt runs this for a fixed two weeks, checks that cost per order and complaint rate did not rise, and only then ships variant B.

Follow-up questions to expect

  • "How many users do you need?" — Run a power calculation before starting: detecting a 3-point lift from a 12% base at 80% power and 5% significance needs roughly 2,000 users per variant. With less traffic, rely on offline evals.
  • "What if B is better but costs twice as much?" — That is what guardrail metrics are for; the decision weighs quality against cost and latency, agreed before the test.
  • "How do you test a risky change safely?" — Shadow mode: serve A, run B in the background on the same requests, and compare B's outputs offline with no user exposure.