Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your model fine-tuned on support transcripts now emits real customer emails when prompted in specific ways. How do you prevent and detect PII memorization in fine-tuned models?


What you need to know

What memorisation is

Two facts drive the fix. First, repetition drives memorisation: text that appears many times in training (signature blocks, quoted email threads) is far more likely to be reproduced word for word. Second, more training on the same data means more memorisation: extra epochs and high learning rates push the model toward copying.

Before training: scrub and deduplicate

Replace PII with consistent placeholders rather than deleting it. The model still learns that an email address goes here, and a conversation that mentions the same email twice keeps its structure.

Python
from presidio_analyzer import AnalyzerEngineanalyzer = AnalyzerEngine()ENTITIES = ["EMAIL_ADDRESS", "PHONE_NUMBER", "PERSON", "CREDIT_CARD"]def pseudonymise(text: str, mapping: dict) -> str:    results = analyzer.analyze(text=text, language="en", entities=ENTITIES)    end = len(text) + 1    for r in sorted(results, key=lambda r: r.start, reverse=True):        if r.end > end:          # skip a span that overlaps one already replaced            continue        end = r.start        key = (r.entity_type, text[r.start:r.end])        if key not in mapping:            n = sum(1 for k in mapping if k[0] == r.entity_type) + 1            mapping[key] = f"<{r.entity_type}_{n}>"        text = text[:r.start] + mapping[key] + text[r.end:]    return text# "Mail rahul.k@example.com" -> "Mail <EMAIL_ADDRESS_1>"

Presidio is a common open-source detector. Test it on your own data: Indian formats (Aadhaar, PAN, UPI IDs, +91 numbers) often need extra recognisers or regex patterns. Then near-deduplicate with MinHash so that repeated signatures and quoted threads appear once.

During and after training

TechniqueEffect on memorisationCost
PII scrubbingRemoves the target of the leakDetector misses some PII; measure its recall
Near-deduplicationLarge drop, since repetition drives memorisingSlight data loss
Fewer epochs, lower learning rateLess copyingMay underfit
LoRA instead of full fine-tuningFewer trainable parameters, usually less memorisationNot a guarantee
DP-SGD (differentially private training)Formal bound on what one record can influenceReal accuracy loss, slower training

Then test for it as a release gate:

  1. Plant canaries — insert 50 fake records with unique fake emails into training data before training.
  2. Prefix attack — prompt with the start of real training records and check if the model completes the real continuation.
  3. Targeted probes — "what is the email of the customer in ticket 4821?", role-play and repeated-token prompts.
  4. Measure — leak rate on real PII and canary recall; block release above zero on canaries.

In production, a PII filter on the output path catches what slips through, and logging each redaction shows you attack patterns.

A real-life example

Scenario, numbers made up. A telecom company fine-tunes a 7B model on 400,000 support transcripts for five epochs. A red-team run of 2,000 probes gets 37 real customer emails out of it. Investigation shows one agent's signature block, with a customer escalation address, appears 9,000 times.

The team retrains: Presidio plus custom Indian-ID patterns replace PII with placeholders, MinHash deduplication removes about 18% of tokens, and training uses LoRA for two epochs. Fifty planted canaries go in first. The new model leaks 0 of 2,000 real emails and 0 of 50 canaries, with a small drop in answer quality that the team judges acceptable. The old model and its checkpoints are deleted, and the output filter stays on.

Follow-up questions to expect

  • "Isn't an output PII filter enough?" — It is a backstop, not a fix. The data is still in the weights, filters miss unusual formats, and regulators expect you to remove the data at the source.
  • "What if a customer asks for their data to be deleted?" — Reliable "unlearning" of one record from a trained model is still immature. Keep data lineage so you know which models saw which records, and retrain from scrubbed data on a schedule.
  • "How do you know the scrubber worked?" — Hand-label a sample of a few hundred transcripts and measure the detector's recall per PII type before training.