Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How should AI systems handle personally identifiable information (PII)?
What you need to know
Where PII leaks in an LLM system
- Prompts sent to a model vendor.
- Logs and traces: the full prompt and response stored for debugging, often with weaker access control than the main database.
- Retrieved context: a RAG index built from tickets or emails contains other people's data.
- Outputs: the model repeats what it saw in context to a user who should not see it.
- Training data: fine-tuning on raw support chats can make a model memorise and later repeat a phone number.
- Derived stores: embeddings, caches and summaries created from personal data are still personal data.
A detection and redaction pipeline
1import re, uuid23def luhn_ok(digits):4 total, parity = 0, len(digits) % 25 for i, ch in enumerate(digits):6 d = int(ch)7 if i % 2 == parity:8 d = d * 2 - 9 if d * 2 > 9 else d * 29 total += d10 return total % 10 == 01112PATTERNS = {13 "CARD": re.compile(r"\b\d(?:[ -]?\d){12,18}\b"),14 "PAN": re.compile(r"\b[A-Z]{5}\d{4}[A-Z]\b"),15 "PHONE": re.compile(r"(?:\+91[ -]?)?\b[6-9]\d{9}\b"),16}1718def redact(text, vault):19 for label, rx in PATTERNS.items():20 def swap(m):21 raw = m.group()22 if label == "CARD" and not luhn_ok(re.sub(r"\D", "", raw)):23 return raw24 token = f"[{label}_{uuid.uuid4().hex[:6]}]"25 vault[token] = raw26 return token27 text = rx.sub(swap, text)28 return text2930vault = {}31print(redact("Card 4111 1111 1111 1111 charged twice. PAN ABCDE1234F, "32 "call 9876543210.", vault))33# Card [CARD_…] charged twice. PAN [PAN_…], call [PHONE_…].The Luhn checksum stops the redactor from removing every 16-digit reference number. The vault maps tokens back to real values; keep it in your own encrypted store so you can put the real values back into the final answer if needed. Regex handles structured IDs; for names, addresses and health details you add an NER model.
Presidio does both in one library: AnalyzerEngine().analyze(text=..., language="en") returns entity spans from pattern recognisers and a spaCy NER model, and AnonymizerEngine().anonymize(...) replaces, masks, hashes or encrypts them. You can add custom recognisers for local IDs. Test its recall on your own data, especially Indian names and addresses.
The rest of the controls
- Encryption in transit and at rest; role-based access to logs and traces.
- Retention limits with automatic deletion.
- A vendor contract (data processing agreement) with no training on your data and a known processing region.
- A deletion path that removes a person's data from the source, the vector index, derived summaries, caches and traces.
A real-life example
A bank's customer chatbot logs every conversation for quality review. An audit finds that 6% of chats contain full card numbers and 2% contain Aadhaar numbers, all stored in plain text in the tracing tool, which 40 engineers can read.
The team adds redaction at the API gateway, before both the model call and the trace write. They plant 500 synthetic PII records in a test set and measure recall per type: cards 99.8%, PAN 99%, phone numbers 97%, but person names only 81%, mostly on South Indian names with initials ("K. S. Raghavan"). They add a custom recogniser for initial-plus-surname patterns and retest at 93%. The trace tool now shows [CARD_…], and only two people in the fraud team can reverse a token.
Follow-up questions to expect
- "Redaction breaks the answer — the bot needs the account number. What now?" — Tokenise instead of deleting. The model reasons over
[ACCOUNT_7f2a]; your code swaps the real value back in after the model call, inside your boundary. - "Is hashing enough?" — Not for low-entropy data. A 10-digit phone number can be recovered from an unsalted hash by trying all numbers. Use keyed hashing or tokenisation with a vault.
- "Can you use a self-hosted model instead?" — Yes, that removes the vendor transfer risk, but logs, retrieval and outputs still need the same controls.