LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you implement A/B testing for LLM systems?


What you need to know

Offline eval versus online A/B

Offline eval

  • Fixed golden set, run in CI
  • Fast and cheap; catches regressions
  • Measures what you thought to test
  • Gate before any user sees a change

Online A/B test

  • Real users, random split
  • Slow; needs thousands of sessions
  • Measures real outcomes and surprises
  • Decides whether a change is worth keeping

Design rules

  • Unit of randomisation — user or session, via a stable hash. Per-request randomisation mixes arms inside one conversation.
  • One change per experiment — prompt, model, or retrieval config. Change two and you cannot tell which caused the result.
  • Primary metric — a business or product outcome: resolution without a human, task completion, conversion, thumbs-up rate.
  • Guardrail metrics — must not regress: cost per request, p95 latency, error rate, guardrail block rate, complaint rate.
  • Duration — at least one full week; novelty can inflate early results.

Sample size, with numbers

Small improvements need big samples. Suppose the escalation-to-human rate is 18% and you want to detect a drop to 16.5% with the usual 5% significance level and 80% power:

Python
from statistics import NormalDistdef sample_size_per_arm(p_base, p_new, alpha=0.05, power=0.80):    z = NormalDist().inv_cdf    z_a, z_b = z(1 - alpha / 2), z(power)    var = p_base * (1 - p_base) + p_new * (1 - p_new)    return int((z_a + z_b) ** 2 * var / (p_base - p_new) ** 2) + 1print(sample_size_per_arm(0.18, 0.165))   # 9955

You need about 10,000 conversations per arm. At 4,000 conversations a day in the treatment arm, that is 2–3 days of data — but you still run a full week for weekly patterns. Because conversations from the same user are correlated, randomising by user means analysing by user too, or using methods that account for clustering.

LLM-specific pitfalls

  • Judge drift — if the metric is scored by an LLM judge, pin the judge model and prompt for the whole test.
  • Cost differences — a "winning" arm that costs 3× more may not be worth it; report cost per successful outcome.
  • Length bias — longer answers often get more thumbs-up but lower completion; watch several metrics.

A real-life example

A fintech support bot tests prompt v15, which asks one clarifying question before answering refund queries. Offline, it passes the eval set. Online, 10% of users get v15 for two weeks; about 56,000 conversations reach each arm.

Results: escalation to humans falls from 18.1% to 16.4%, which is significant. Average conversation length rises from 3.8 to 4.3 turns, so cost per conversation rises 11% — within the agreed 15% guardrail. Complaint rate is unchanged. The first three days showed a much bigger drop (to 14.9%), which faded — a novelty effect the team would have mistaken for the real result if they had stopped early. Since each avoided escalation saves about Rs 60 of agent time, the change pays for its extra tokens many times over, and it ships.

Follow-up questions to expect

  • "What if the LLM judge and user metrics disagree?" — Trust the user outcome; then investigate whether the judge rubric measures the wrong thing.
  • "Can you A/B test a model change?" — Yes, the same way; pin both snapshots, and remember that cost and latency differences are part of the result.
  • "What is an interleaving or side-by-side test?" — Showing outputs from both arms (or mixing ranked results) and collecting preferences; faster for ranking-like tasks, but less realistic.