Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your model fine-tuned on support transcripts now emits real customer emails when prompted in specific ways. How do you prevent and detect PII memorization in fine-tuned models?
What you need to know
What memorisation is
Two facts drive the fix. First, repetition drives memorisation: text that appears many times in training (signature blocks, quoted email threads) is far more likely to be reproduced word for word. Second, more training on the same data means more memorisation: extra epochs and high learning rates push the model toward copying.
Before training: scrub and deduplicate
Replace PII with consistent placeholders rather than deleting it. The model still learns that an email address goes here, and a conversation that mentions the same email twice keeps its structure.
1from presidio_analyzer import AnalyzerEngine23analyzer = AnalyzerEngine()45ENTITIES = ["EMAIL_ADDRESS", "PHONE_NUMBER", "PERSON", "CREDIT_CARD"]67def pseudonymise(text: str, mapping: dict) -> str:8 results = analyzer.analyze(text=text, language="en", entities=ENTITIES)9 end = len(text) + 110 for r in sorted(results, key=lambda r: r.start, reverse=True):11 if r.end > end: # skip a span that overlaps one already replaced12 continue13 end = r.start14 key = (r.entity_type, text[r.start:r.end])15 if key not in mapping:16 n = sum(1 for k in mapping if k[0] == r.entity_type) + 117 mapping[key] = f"<{r.entity_type}_{n}>"18 text = text[:r.start] + mapping[key] + text[r.end:]19 return text2021# "Mail rahul.k@example.com" -> "Mail <EMAIL_ADDRESS_1>"Presidio is a common open-source detector. Test it on your own data: Indian formats (Aadhaar, PAN, UPI IDs, +91 numbers) often need extra recognisers or regex patterns. Then near-deduplicate with MinHash so that repeated signatures and quoted threads appear once.
During and after training
| Technique | Effect on memorisation | Cost |
|---|---|---|
| PII scrubbing | Removes the target of the leak | Detector misses some PII; measure its recall |
| Near-deduplication | Large drop, since repetition drives memorising | Slight data loss |
| Fewer epochs, lower learning rate | Less copying | May underfit |
| LoRA instead of full fine-tuning | Fewer trainable parameters, usually less memorisation | Not a guarantee |
| DP-SGD (differentially private training) | Formal bound on what one record can influence | Real accuracy loss, slower training |
Then test for it as a release gate:
- Plant canaries — insert 50 fake records with unique fake emails into training data before training.
- Prefix attack — prompt with the start of real training records and check if the model completes the real continuation.
- Targeted probes — "what is the email of the customer in ticket 4821?", role-play and repeated-token prompts.
- Measure — leak rate on real PII and canary recall; block release above zero on canaries.
In production, a PII filter on the output path catches what slips through, and logging each redaction shows you attack patterns.
A real-life example
Scenario, numbers made up. A telecom company fine-tunes a 7B model on 400,000 support transcripts for five epochs. A red-team run of 2,000 probes gets 37 real customer emails out of it. Investigation shows one agent's signature block, with a customer escalation address, appears 9,000 times.
The team retrains: Presidio plus custom Indian-ID patterns replace PII with placeholders, MinHash deduplication removes about 18% of tokens, and training uses LoRA for two epochs. Fifty planted canaries go in first. The new model leaks 0 of 2,000 real emails and 0 of 50 canaries, with a small drop in answer quality that the team judges acceptable. The old model and its checkpoints are deleted, and the output filter stays on.
Follow-up questions to expect
- "Isn't an output PII filter enough?" — It is a backstop, not a fix. The data is still in the weights, filters miss unusual formats, and regulators expect you to remove the data at the source.
- "What if a customer asks for their data to be deleted?" — Reliable "unlearning" of one record from a trained model is still immature. Keep data lineage so you know which models saw which records, and retrain from scrubbed data on a schedule.
- "How do you know the scrubber worked?" — Hand-label a sample of a few hundred transcripts and measure the detector's recall per PII type before training.