Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Preparing data and comparing against the prompted baseline


By the time the router was trained, PolicyPal had eleven weeks of logs: about 66,000 messages, each labelled by the prompted router from Section 4. The tempting shortcut was to train on all of them as they were. It would be fast, and the labels were free.

The problem is that the prompted router was right 94.8% of the time. Training on its labels directly, a form of distillation, would teach the new model the teacher's 5.2% of mistakes as if they were correct. Worse, those mistakes were not random. The prompted router was weakest exactly on the rare, important classes: multi-intent messages and sensitive HR topics.

A fine-tuned model is only as good as the data it learns from, and the honesty of its evaluation depends on how that data is split. This lesson builds the dataset, runs the comparison that decides whether the router ships, and sets the rules that keep it good after launch.

From logs to an honest test set66,000 logged,teacher-labelled3,000sampled by route820 reviewed,131 labels fixedGroupnear-duplicatesat 0.95Split bygroup:80, 10, 10Held-out result: 95.6% against 94.8%, 15 ms against 850 ms.
Splitting by near-duplicate group keeps paraphrases on one side, so the test measures skill rather than memory.

Sampling: what goes into the dataset

Sixty-six thousand messages are far more than the router needs, and most of them are the same few questions. "How many leaves do I have?" and its variants appeared about 900 times. PolicyPal's dataset took 3,000 messages, sampled to cover what matters.

  • Stratified by route, so rare routes are well represented: sensitive HR topics are 3% of traffic but 400 examples, 13% of the dataset.
  • Hard cases added on purpose: every message where the prompted mid-size and small models disagreed.
  • Redacted before storage: names, employee IDs and phone numbers replaced with placeholders, using the PII tools from Section 9, because training data is kept far longer than logs.

Labelling: where the truth comes from

Starting labels came from the prompted router. Then people checked the places where it was most likely to be wrong: every disagreement between the two prompted models, every sensitive_hr and multi_step example, and a random 10% of the rest. That was 820 messages, reviewed by two HR and IT staff over about 25 hours, using a two-page labelling guide with a definition and three edge cases per route.

They corrected 131 labels, 16% of those reviewed. Most corrections moved messages into multi_step or sensitive_hr, which confirmed the worry about the teacher's blind spots. A message like "My manager keeps commenting on my health in meetings, what's the sick leave policy?" had been labelled policy_question. It is a sensitive HR matter first.

Deduplication and a leak-free split

If a message and its near-copy land on both sides of the split, one in training and one in test, the test measures memory, not skill. With support messages this happens constantly, because people ask the same thing in almost the same words. So PolicyPal groups near-duplicates first and splits by group, not by message.

Python
# policypal/router/split.pyimport jsonimport randomfrom sentence_transformers import SentenceTransformerrows = [json.loads(line) for line in open("router/labelled.jsonl")]emb = SentenceTransformer("BAAI/bge-small-en-v1.5").encode(    [r["text"] for r in rows], normalize_embeddings=True)leaders, group = [], []for i, vec in enumerate(emb):    sims = emb[leaders] @ vec if leaders else []    best = max(range(len(leaders)), key=lambda g: sims[g], default=None)    if best is None or sims[best] < 0.95:        # not a near-duplicate of any leader        leaders.append(i)        best = len(leaders) - 1    group.append(best)ids = list(range(len(leaders)))random.Random(7).shuffle(ids)cut1, cut2 = int(0.8 * len(ids)), int(0.9 * len(ids))split_of = {g: "train" if k < cut1 else "val" if k < cut2 else "test" for k, g in enumerate(ids)}for name in ("train", "val", "test"):    with open(f"router/{name}.jsonl", "w") as f:        for row, g in zip(rows, group):            if split_of[g] == name:                f.write(json.dumps({"text": row["text"], "label": row["label"]}) + "\n")

Each message joins the group of the first earlier message it is at least 95% similar to, or starts a new group. Whole groups go to train, validation or test, so paraphrases never straddle the boundary. The fixed random seed makes the split reproducible, which matters when you retrain later and want to compare fairly. On 3,000 messages this runs in a few seconds.

The test file is then put away. Nobody looks at it while choosing ranks, epochs or thresholds. It is opened once, for the decision.

The comparison that decides

All three candidates were run on the same 300 test messages. Accuracy alone would hide what matters, so the table shows the two classes with the highest cost of error.

RouterAccuracyMacro-F1Sensitive HR recallMulti-step recallMedian latency
Prompted mid-size model94.8%0.9239 of 4084%850 ms
Prompted small hosted model91.2%0.8738 of 4076%520 ms
Fine-tuned 0.5B with LoRA95.6%0.9439 of 4088%15 ms

Macro-F1 averages the F1 score of each class equally, so a rare class counts as much as a common one. It rose more than accuracy did, which means the fine-tuned router improved mostly on the rare classes, exactly where the corrected labels were.

But look at the sensitive column. All three routers missed one of 40, and PolicyPal's target is 99% recall. Forty examples are also far too few to prove 99% of anything; one miss moves the number by 2.5 points. So HR helped assemble a separate sensitive test set of 200 real, redacted messages from their case records. On that set the fine-tuned router's recall was 97.0%. Not good enough.

A threshold for the class that must not be missed

The classification head gives a probability for every route. Instead of always taking the highest, PolicyPal biases the decision towards safety.

Python
def route(probs: dict[str, float]) -> str:    if probs["sensitive_hr"] >= 0.15:          # any real chance of a sensitive topic        return "sensitive_hr"    best = max(probs, key=probs.get)    if probs[best] < 0.60:                     # unsure: the agent can handle anything        return "multi_step"    return best

With the 0.15 threshold, sensitive recall on the 200-message set rose to 99.0%. The cost: about 2% of ordinary messages are now sent to an HR partner unnecessarily. HR accepted that, because dismissing a false handover takes about two minutes, and a missed sensitive case can take months to repair. The second rule sends low-confidence messages to the agent, which is slower but can handle any mix of intents. It applies to about 4% of traffic.

Thresholds like these are product decisions with a price, not tuning tricks. Write down the trade-off and who agreed to it.

Shadow first, then switch, then watch

Before switching, the fine-tuned router ran in shadow for a week: both routers classified every message, only the prompted one's decision was used, and every disagreement was logged. They disagreed on 1.9% of messages. On a reviewed sample of 200 disagreements, the fine-tuned router was right 58% of the time, the prompted one 36%, and both were wrong 6%.

After switching, the prompted router stays as a fallback for when the GPU service is down. Retraining is triggered by any of three signals: a new route, the share of low-confidence messages rising above 6% for a week, or the monthly review of 100 random messages finding accuracy below 94%.

Check your understanding

0 of 3 answered

1.Why does PolicyPal split data by near-duplicate group rather than by individual message?

2.On 40 sensitive test messages, the router missed one. Why did the team build a separate 200-message sensitive test set?

3.The 0.15 threshold on sensitive_hr sends about 2% of ordinary messages to HR unnecessarily. Why is that acceptable?