Course Content
RAG Systems
12 sections · 66 lessons
How does RAG help with private or sensitive data?
What you need to know
There are two ways to give a model private knowledge. Their privacy properties are very different.
Fine-tuning on private data
- Facts are mixed into the weights
- Cannot remove one person's data without retraining
- Everyone using the model can reach everything it learned
- Hard to prove where an answer came from
RAG over private data
- Facts stay in a store you control
- Delete the document and its chunks, and it is gone from answers
- Each query sees only what that user may see
- Every answer lists its sources
The four real benefits
- Revocation. Delete a document and its chunks, and future answers cannot use it. With fine-tuned weights, a model can still reproduce memorised text.
- Per-query access control. The retriever filters by the requesting user's permissions, so two employees asking the same question can get different, correct answers.
- Auditability. You can log exactly which chunks were shown for each answer.
- Data minimisation. Each request sends a few hundred to a few thousand tokens of relevant text, not your whole corpus.
Where the data still goes
- The model provider. If you call a hosted model, chunks leave your network in every prompt. You need a contract covering retention and training use, the right region, or a self-hosted model.
- Logs, traces and caches. Each of these holds copies of document text and questions.
- The vector store itself. Embeddings are not anonymous. Research such as Vec2Text (2023) showed that text can be largely reconstructed from some embeddings, so protect the vector store like the source documents.
- The wrong user. One missing permission filter, and a private chunk appears in someone else's answer.
A real-life example
A company wants its HR assistant to answer questions about pay bands, which are confidential by grade. Option A, fine-tuning a model on HR documents, is rejected: every employee could coax out other grades' bands, and when bands change each April the model would need retraining.
Option B is RAG. Each pay-band chunk carries allowed_groups: ["hr", "grade_L5"]. An L5 engineer asking "what is my pay band?" gets only the L5 band, with a citation. When the April 2026 bands are published, the old chunks are deleted and new ones indexed the same day. The company also chooses a model deployment in its own cloud region with zero data retention, and redacts employee names from traces.
Follow-up questions to expect
- "Is sending chunks to a hosted model safe?" — It depends on the contract and settings: data retention, training use, region, and certifications. For highly sensitive data, use a private deployment or a self-hosted model.
- "Are embeddings safe to share?" — No. Treat them as containing the source text.
- "Can the model leak data it saw in an earlier request?" — Not through the weights, since RAG does not train it. But shared caches, logs and conversation memory can, if they are not scoped per user.