Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Online evaluation and A/B tests


The offline suite said PolicyPal v2.4 was three points better than v2.3 on key-fact correctness. The product manager asked the right question: "Does that make any difference to employees?" The eval set is 250 cases chosen by the team and frozen in time. Real traffic is 1,200 questions a day, shifts with every policy announcement, and contains questions nobody on the team has imagined.

Offline evaluation tells you whether a change is safe to try. Online evaluation tells you what it did to real people. It is harder, because production has no gold answers. You have to infer quality from what people do, and you have to be careful not to see effects that are not there.

This lesson covers the signals worth collecting, the arithmetic of A/B tests at PolicyPal's scale, and the rollout process that keeps a bad change from reaching 2,000 people at once.

How a change reaches 2,000 peopleOffline gate: suite and flip listShadow: runs silently on real trafficCanary: 5% for two days, guardrailsA/B: 50% for two weeks, re-ask rateFull rollout, automatic rollback
At 1,200 questions a day the re-ask rate can decide a test in a week, while thumbs at 8% coverage would need 74 days.

Signals from production

No single signal measures answer quality. Each one sees part of it and misses the rest.

SignalWhat it suggestsBlind spot
Thumbs up or downDirect opinionOnly 8% of answers are rated, mostly by unhappy or very happy people
Re-asked within 2 minutesThe answer did not helpSome people re-ask to get more detail
HR or IT ticket on the same topic within 24 hoursThe question was not resolvedSome topics always need a ticket
"Talk to HR" after an answerThe answer was not enoughSometimes the right outcome
Citation clickedThe user checked the sourceCould mean trust or doubt
Share of not_in_policy and needs_humanCoverage and cautionRises with genuinely new topics

PolicyPal's primary online metric is the re-ask rate, because it is recorded for every question, not only the 8% that get a rating, and a rephrased question shortly after an answer is a strong sign the answer failed. The 24-hour ticket rate is the second metric, because deflecting helpdesk tickets is what the project promised.

Behavioural signals still need a human anchor. Every week, 50 random answers are graded by an HR reviewer with the same rubric as the offline suite. That produces a number directly comparable to the offline one, and it catches the kind of problem no behaviour reveals: answers that are wrong but believed. Reviewers see answers with names and IDs redacted.

A/B tests, and the arithmetic that limits them

In an A/B test, users are split at random between the current version and the new one, and you compare a metric. Two rules apply before any maths. Randomise by user, not by message, so each person sees one consistent version and their repeated questions do not leak between arms. And decide the metric and the test length before starting, not after looking at the numbers.

How many questions do you need? For a metric that is a rate, a standard approximation gives the sample size per arm for a 5% false-alarm rate and 80% power.

Python
from math import ceildef n_per_arm(p1: float, p2: float, z_alpha: float = 1.96, z_power: float = 0.84) -> int:    """Samples per arm to detect a change from rate p1 to p2."""    variance = p1 * (1 - p1) + p2 * (1 - p2)    return ceil((z_alpha + z_power) ** 2 * variance / (p1 - p2) ** 2)print(n_per_arm(0.12, 0.10))    # re-ask rate 12% -> 10%: 3,834 questions per armprint(n_per_arm(0.70, 0.73))    # thumbs-up share 70% -> 73%: 3,547 ratings per arm

Now apply PolicyPal's traffic. With 1,200 questions a day split in half, each arm gets 600 questions a day. Detecting a re-ask rate change from 12% to 10% needs 3,834 per arm: about six and a half working days. Round it up to two full weeks, so both arms see the same mix of Mondays and Fridays. That is feasible.

The thumbs metric is different. Only 8% of answers are rated, about 48 ratings per arm per day. Detecting a three-point change in the thumbs-up share needs 3,547 ratings per arm: about 74 days. That is not an experiment; it is a season, during which policies and traffic will change underneath it. At PolicyPal's scale, thumbs are useful for finding bad answers to read, and useless for deciding between versions.

Three habits prevent fooling yourself. Do not stop early because the numbers look good on day three; checking repeatedly and stopping at the first good-looking result greatly inflates false alarms. Pick one primary metric, because testing ten metrics will find one that "moved" by chance. And watch for novelty: a visible change can shift behaviour for a few days simply because it is new.

When you cannot run an A/B test

Many changes cannot, or should not, be A/B tested at Harbourline's size. Effects smaller than about two points on the re-ask rate need more traffic than two weeks provides. Safety fixes, like a stronger PII filter, ship to everyone immediately; you do not leave half the company unprotected to measure the difference.

For these, the evidence comes from the offline suite with flips, a week of shadow traffic where the new version runs in the background and its outputs are compared and sampled for review, and a careful staged rollout.

A staged rollout with guardrail metrics

  1. Offline gate — the eval suite passes, with no broken cases in the sensitive, adversarial or tool groups and no unconfirmed side effects.
  2. Shadow — the new version answers real traffic in the background for a few days; disagreements are sampled and reviewed.
  3. Canary — 5% of users get the new version for two days, watched on guardrail metrics: error rate, p95 latency, cost per question and handover rate.
  4. A/B at 50% — when the effect is large enough to measure, for two full weeks on the primary metric.
  5. Full rollout — with automatic rollback if any guardrail crosses its limit.

Guardrail metrics are not what you are trying to improve. They are what must not get worse: latency, errors, cost and the handover rate. A change that improves the re-ask rate but doubles p95 latency is not an improvement.

Check your understanding

0 of 3 answered

1.Why can PolicyPal not use the thumbs-up rate to decide between two versions?

2.An A/B test shows a significant improvement on day 3 of a planned 14-day run. What should the team do?

3.A stronger PII filter is ready. Should it go through a 50% A/B test?