Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Preparing data and comparing against the prompted baseline
By the time the router was trained, PolicyPal had eleven weeks of logs: about 66,000 messages, each labelled by the prompted router from Section 4. The tempting shortcut was to train on all of them as they were. It would be fast, and the labels were free.
The problem is that the prompted router was right 94.8% of the time. Training on its labels directly, a form of distillation, would teach the new model the teacher's 5.2% of mistakes as if they were correct. Worse, those mistakes were not random. The prompted router was weakest exactly on the rare, important classes: multi-intent messages and sensitive HR topics.
A fine-tuned model is only as good as the data it learns from, and the honesty of its evaluation depends on how that data is split. This lesson builds the dataset, runs the comparison that decides whether the router ships, and sets the rules that keep it good after launch.
Sampling: what goes into the dataset
Sixty-six thousand messages are far more than the router needs, and most of them are the same few questions. "How many leaves do I have?" and its variants appeared about 900 times. PolicyPal's dataset took 3,000 messages, sampled to cover what matters.
- Stratified by route, so rare routes are well represented: sensitive HR topics are 3% of traffic but 400 examples, 13% of the dataset.
- Hard cases added on purpose: every message where the prompted mid-size and small models disagreed.
- Redacted before storage: names, employee IDs and phone numbers replaced with placeholders, using the PII tools from Section 9, because training data is kept far longer than logs.
Labelling: where the truth comes from
Starting labels came from the prompted router. Then people checked the places where it was most likely to be wrong: every disagreement between the two prompted models, every sensitive_hr and multi_step example, and a random 10% of the rest. That was 820 messages, reviewed by two HR and IT staff over about 25 hours, using a two-page labelling guide with a definition and three edge cases per route.
They corrected 131 labels, 16% of those reviewed. Most corrections moved messages into multi_step or sensitive_hr, which confirmed the worry about the teacher's blind spots. A message like "My manager keeps commenting on my health in meetings, what's the sick leave policy?" had been labelled policy_question. It is a sensitive HR matter first.
Deduplication and a leak-free split
If a message and its near-copy land on both sides of the split, one in training and one in test, the test measures memory, not skill. With support messages this happens constantly, because people ask the same thing in almost the same words. So PolicyPal groups near-duplicates first and splits by group, not by message.
1# policypal/router/split.py2import json3import random45from sentence_transformers import SentenceTransformer67rows = [json.loads(line) for line in open("router/labelled.jsonl")]8emb = SentenceTransformer("BAAI/bge-small-en-v1.5").encode(9 [r["text"] for r in rows], normalize_embeddings=True)1011leaders, group = [], []12for i, vec in enumerate(emb):13 sims = emb[leaders] @ vec if leaders else []14 best = max(range(len(leaders)), key=lambda g: sims[g], default=None)15 if best is None or sims[best] < 0.95: # not a near-duplicate of any leader16 leaders.append(i)17 best = len(leaders) - 118 group.append(best)1920ids = list(range(len(leaders)))21random.Random(7).shuffle(ids)22cut1, cut2 = int(0.8 * len(ids)), int(0.9 * len(ids))23split_of = {g: "train" if k < cut1 else "val" if k < cut2 else "test" for k, g in enumerate(ids)}2425for name in ("train", "val", "test"):26 with open(f"router/{name}.jsonl", "w") as f:27 for row, g in zip(rows, group):28 if split_of[g] == name:29 f.write(json.dumps({"text": row["text"], "label": row["label"]}) + "\n")Each message joins the group of the first earlier message it is at least 95% similar to, or starts a new group. Whole groups go to train, validation or test, so paraphrases never straddle the boundary. The fixed random seed makes the split reproducible, which matters when you retrain later and want to compare fairly. On 3,000 messages this runs in a few seconds.
The test file is then put away. Nobody looks at it while choosing ranks, epochs or thresholds. It is opened once, for the decision.
The comparison that decides
All three candidates were run on the same 300 test messages. Accuracy alone would hide what matters, so the table shows the two classes with the highest cost of error.
| Router | Accuracy | Macro-F1 | Sensitive HR recall | Multi-step recall | Median latency |
|---|---|---|---|---|---|
| Prompted mid-size model | 94.8% | 0.92 | 39 of 40 | 84% | 850 ms |
| Prompted small hosted model | 91.2% | 0.87 | 38 of 40 | 76% | 520 ms |
| Fine-tuned 0.5B with LoRA | 95.6% | 0.94 | 39 of 40 | 88% | 15 ms |
Macro-F1 averages the F1 score of each class equally, so a rare class counts as much as a common one. It rose more than accuracy did, which means the fine-tuned router improved mostly on the rare classes, exactly where the corrected labels were.
But look at the sensitive column. All three routers missed one of 40, and PolicyPal's target is 99% recall. Forty examples are also far too few to prove 99% of anything; one miss moves the number by 2.5 points. So HR helped assemble a separate sensitive test set of 200 real, redacted messages from their case records. On that set the fine-tuned router's recall was 97.0%. Not good enough.
A threshold for the class that must not be missed
The classification head gives a probability for every route. Instead of always taking the highest, PolicyPal biases the decision towards safety.
1def route(probs: dict[str, float]) -> str:2 if probs["sensitive_hr"] >= 0.15: # any real chance of a sensitive topic3 return "sensitive_hr"4 best = max(probs, key=probs.get)5 if probs[best] < 0.60: # unsure: the agent can handle anything6 return "multi_step"7 return bestWith the 0.15 threshold, sensitive recall on the 200-message set rose to 99.0%. The cost: about 2% of ordinary messages are now sent to an HR partner unnecessarily. HR accepted that, because dismissing a false handover takes about two minutes, and a missed sensitive case can take months to repair. The second rule sends low-confidence messages to the agent, which is slower but can handle any mix of intents. It applies to about 4% of traffic.
Thresholds like these are product decisions with a price, not tuning tricks. Write down the trade-off and who agreed to it.
Shadow first, then switch, then watch
Before switching, the fine-tuned router ran in shadow for a week: both routers classified every message, only the prompted one's decision was used, and every disagreement was logged. They disagreed on 1.9% of messages. On a reviewed sample of 200 disagreements, the fine-tuned router was right 58% of the time, the prompted one 36%, and both were wrong 6%.
After switching, the prompted router stays as a fallback for when the GPU service is down. Retraining is triggered by any of three signals: a new route, the share of low-confidence messages rising above 6% for a week, or the monthly review of 100 random messages finding accuracy below 94%.
Check your understanding
0 of 3 answered
1.Why does PolicyPal split data by near-duplicate group rather than by individual message?
2.On 40 sensitive test messages, the router missed one. Why did the team build a separate 200-message sensitive test set?
3.The 0.15 threshold on sensitive_hr sends about 2% of ordinary messages to HR unnecessarily. Why is that acceptable?