Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Online evaluation and A/B tests
The offline suite said PolicyPal v2.4 was three points better than v2.3 on key-fact correctness. The product manager asked the right question: "Does that make any difference to employees?" The eval set is 250 cases chosen by the team and frozen in time. Real traffic is 1,200 questions a day, shifts with every policy announcement, and contains questions nobody on the team has imagined.
Offline evaluation tells you whether a change is safe to try. Online evaluation tells you what it did to real people. It is harder, because production has no gold answers. You have to infer quality from what people do, and you have to be careful not to see effects that are not there.
This lesson covers the signals worth collecting, the arithmetic of A/B tests at PolicyPal's scale, and the rollout process that keeps a bad change from reaching 2,000 people at once.
Signals from production
No single signal measures answer quality. Each one sees part of it and misses the rest.
| Signal | What it suggests | Blind spot |
|---|---|---|
| Thumbs up or down | Direct opinion | Only 8% of answers are rated, mostly by unhappy or very happy people |
| Re-asked within 2 minutes | The answer did not help | Some people re-ask to get more detail |
| HR or IT ticket on the same topic within 24 hours | The question was not resolved | Some topics always need a ticket |
| "Talk to HR" after an answer | The answer was not enough | Sometimes the right outcome |
| Citation clicked | The user checked the source | Could mean trust or doubt |
Share of not_in_policy and needs_human | Coverage and caution | Rises with genuinely new topics |
PolicyPal's primary online metric is the re-ask rate, because it is recorded for every question, not only the 8% that get a rating, and a rephrased question shortly after an answer is a strong sign the answer failed. The 24-hour ticket rate is the second metric, because deflecting helpdesk tickets is what the project promised.
Behavioural signals still need a human anchor. Every week, 50 random answers are graded by an HR reviewer with the same rubric as the offline suite. That produces a number directly comparable to the offline one, and it catches the kind of problem no behaviour reveals: answers that are wrong but believed. Reviewers see answers with names and IDs redacted.
A/B tests, and the arithmetic that limits them
In an A/B test, users are split at random between the current version and the new one, and you compare a metric. Two rules apply before any maths. Randomise by user, not by message, so each person sees one consistent version and their repeated questions do not leak between arms. And decide the metric and the test length before starting, not after looking at the numbers.
How many questions do you need? For a metric that is a rate, a standard approximation gives the sample size per arm for a 5% false-alarm rate and 80% power.
1from math import ceil23def n_per_arm(p1: float, p2: float, z_alpha: float = 1.96, z_power: float = 0.84) -> int:4 """Samples per arm to detect a change from rate p1 to p2."""5 variance = p1 * (1 - p1) + p2 * (1 - p2)6 return ceil((z_alpha + z_power) ** 2 * variance / (p1 - p2) ** 2)78print(n_per_arm(0.12, 0.10)) # re-ask rate 12% -> 10%: 3,834 questions per arm9print(n_per_arm(0.70, 0.73)) # thumbs-up share 70% -> 73%: 3,547 ratings per armNow apply PolicyPal's traffic. With 1,200 questions a day split in half, each arm gets 600 questions a day. Detecting a re-ask rate change from 12% to 10% needs 3,834 per arm: about six and a half working days. Round it up to two full weeks, so both arms see the same mix of Mondays and Fridays. That is feasible.
The thumbs metric is different. Only 8% of answers are rated, about 48 ratings per arm per day. Detecting a three-point change in the thumbs-up share needs 3,547 ratings per arm: about 74 days. That is not an experiment; it is a season, during which policies and traffic will change underneath it. At PolicyPal's scale, thumbs are useful for finding bad answers to read, and useless for deciding between versions.
Three habits prevent fooling yourself. Do not stop early because the numbers look good on day three; checking repeatedly and stopping at the first good-looking result greatly inflates false alarms. Pick one primary metric, because testing ten metrics will find one that "moved" by chance. And watch for novelty: a visible change can shift behaviour for a few days simply because it is new.
When you cannot run an A/B test
Many changes cannot, or should not, be A/B tested at Harbourline's size. Effects smaller than about two points on the re-ask rate need more traffic than two weeks provides. Safety fixes, like a stronger PII filter, ship to everyone immediately; you do not leave half the company unprotected to measure the difference.
For these, the evidence comes from the offline suite with flips, a week of shadow traffic where the new version runs in the background and its outputs are compared and sampled for review, and a careful staged rollout.
A staged rollout with guardrail metrics
- Offline gate — the eval suite passes, with no broken cases in the sensitive, adversarial or tool groups and no unconfirmed side effects.
- Shadow — the new version answers real traffic in the background for a few days; disagreements are sampled and reviewed.
- Canary — 5% of users get the new version for two days, watched on guardrail metrics: error rate, p95 latency, cost per question and handover rate.
- A/B at 50% — when the effect is large enough to measure, for two full weeks on the primary metric.
- Full rollout — with automatic rollback if any guardrail crosses its limit.
Guardrail metrics are not what you are trying to improve. They are what must not get worse: latency, errors, cost and the handover rate. A change that improves the re-ask rate but doubles p95 latency is not an improvement.
Check your understanding
0 of 3 answered
1.Why can PolicyPal not use the thumbs-up rate to decide between two versions?
2.An A/B test shows a significant improvement on day 3 of a planned 14-day run. What should the team do?
3.A stronger PII filter is ready. Should it go through a 50% A/B test?