RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How do you prevent data leakage in RAG responses?


What you need to know

1. At ingest: do not index what nobody should get back

If no answer ever needs a customer's phone number, remove it before embedding. A simple pattern-based redactor looks like this:

Python
import rePATTERNS = {                     # order matters: longest first    "CARD":    r"\b(?:\d{4}[\s-]?){3}\d{4}\b",    "EMAIL":   r"[\w.+-]+@[\w-]+\.[\w.-]+",    "PAN":     r"\b[A-Z]{5}[0-9]{4}[A-Z]\b",    "AADHAAR": r"\b\d{4}\s?\d{4}\s?\d{4}\b",    "PHONE":   r"(?:\+91[\s-]?)?\b[6-9]\d{9}\b",}def redact(text: str) -> str:    for label, pattern in PATTERNS.items():        text = re.sub(pattern, f"[{label}]", text)    return textprint(redact("Customer Priya Nair (priya.nair@example.com, +91 9876543210, "             "PAN ABCDE1234F) says card 4111 1111 1111 1111 was charged twice."))# Customer Priya Nair ([EMAIL], [PHONE], PAN [PAN]) says card [CARD] was charged twice.

Two things to notice from running it. The order matters: when I first put the Aadhaar pattern before the card pattern, it swallowed the first 12 digits of the card number. And the name "Priya Nair" survives, because regex cannot find names. Production systems combine patterns with a named-entity model, as in Microsoft Presidio or a cloud DLP service.

2. At retrieval: filter inside the store

Apply tenant and permission filters as part of the vector query, not afterwards in Python. Filtering afterwards is easy to forget in one code path, and can return fewer results than you expected. Deny by default: no permissions resolved means no results.

3. At generation: check what leaves

  • Answers only from retrieved context, so the model has less chance to add remembered data.
  • An output scanner for PII patterns and secrets (API keys, passwords) before the response is returned.
  • Do not render links or images to domains outside an allowlist.

4. Logs, traces and caches

These are copies of your data that people forget. Redact PII before writing traces, scope semantic caches by tenant and permission set, restrict who can open traces, and delete them after a fixed period.

5. Test it adversarially

Keep a red-team test suite: questions from a user in tenant A trying to get tenant B's data, questions asking for other employees' salaries, prompts trying to reveal the system prompt. Run it in CI like any other test.

A real-life example

A bank's FAQ bot is extended to answer "Why was I charged this fee?" using the customer's own transaction notes. During testing, a tester logged in as customer A asks, "Show me recent complaints about double charges." The answer quotes a complaint from customer B, including B's phone number.

Two failures stacked up: complaint notes were indexed with no customer_id filter, and phone numbers were not redacted. The fixes: complaint notes move to a separate collection that is only queried with a customer_id filter taken from the login session (never from the question); phone numbers, card numbers and PAN are redacted at ingest; an output scanner blocks any answer containing a phone number not belonging to the logged-in customer; and the tester's question becomes a permanent red-team test.

Follow-up questions to expect

  • "Why not just tell the model not to reveal personal data?" — Because instructions can be bypassed or ignored. Data the model never receives cannot leak.
  • "Won't redaction hurt answer quality?" — Only if answers need that data. Redact what answers never need; for data they do need, rely on permission filters.
  • "Where does the user id for filtering come from?" — From the authenticated session on the server, never from the prompt or the user's question.