Course Content
RAG Systems
12 sections · 66 lessons
How do you ensure compliance when using RAG?
What you need to know
Compliance means proving to regulators, auditors and customers that data is handled according to laws and contracts. The laws most often named: the EU GDPR, India's Digital Personal Data Protection (DPDP) Act, 2023, the US HIPAA for health data, and sector rules from regulators such as the RBI for banks. The EU AI Act adds transparency and risk-management duties that phase in from 2025 onward. You do not need to recite them; you need to show the system can meet their common demands.
| Demand | What the RAG system needs |
|---|---|
| "Where is our data processed?" | Known regions for store, logs and model; contracts with each provider |
| "Delete everything about this person" | Chunk metadata with subject_id and source_id; tested delete across index, caches, traces, backups |
| "Why did it say that?" | Per-answer log of user, chunk ids, prompt version, model version |
| "Do you need this data?" | Minimisation at ingest, retention limits, redaction |
| "Is a human accountable?" | Human review for high-stakes answers, clear disclosure that answers are AI-generated |
The erasure path, step by step
- Find — query the index for every chunk with
subject_idorsource_idmatching the request. - Delete — remove those chunks, and the source documents or the relevant parts.
- Purge copies — semantic cache entries and traces that contain them.
- Record — log that the erasure happened, without storing the erased data.
- Verify — run a search for the person's identifiers and confirm nothing returns.
People forget step 3. The trace store is often the largest copy of personal data in a RAG system.
Reconstructing past answers
Store with every answer: timestamp, user, retrieved chunk ids with document versions, prompt template version, model name and version, and the answer. With versioned documents, you can show exactly what the system saw months later, which is what an incident review or a regulator will ask.
A real-life example
A bank's customer receives a wrong answer about a loan prepayment penalty and complains to the bank's grievance officer. The compliance team asks: what did the bot see, and why?
Because the bot logs versions, the team finds the answer in minutes: it cited version 3 of the home loan FAQ, which had been replaced by version 4 two days earlier, but the old chunks were still in the index due to a failed sync job. They can show the exact text shown, correct the answer for the customer, fix the sync, and add a freshness alert.
The same month, a customer asks for their data to be erased under the DPDP Act. The erasure job deletes 14 chunks from call-note summaries, 3 semantic cache entries and 22 traces, then confirms that a search for the customer's id returns nothing. The whole process takes one automated job instead of a week of manual searching.
Follow-up questions to expect
- "Does using a hosted LLM break GDPR or DPDP?" — Not by itself. You need a lawful basis, a processing agreement, the right region and retention terms, and to send only needed data.
- "How do you delete data from backups?" — Usually by keeping backups short-lived and re-applying erasure lists when restoring, following your legal team's policy.
- "How long should you keep traces?" — As short as debugging and audit needs allow, set with legal; often 30 to 90 days for full content, longer for metadata without content.