- MantraMindAI
- Blog
- Responsible AI & Security
AI data privacy: tracing one request from browser to model
Jai Rao
August 22, 202620 min read
Follow a single AI request from the text box to a hosted model and back, naming what is exposed at every hop: logs, traces, retrieval, tool calls and vector indexes.
Most privacy failures in AI features are not exotic. Nobody talks the model into reciting a database. What happens is duller: a support agent pastes a customer email into a summarise box, the HTTP client logs request bodies because someone raised the log level during an incident three months ago and never lowered it, and now the full text sits in a search index that forty people across three teams can query. None of them could have opened the customer record it came from.
That is the shape of the problem. The model is rarely the weak point; the plumbing around it is. So the useful exercise is not a policy review, it is a trace. Take one request, follow it from the browser to the provider and back, and at every hop write down what data is present and what can persist there. Everything below follows that path, in order, and ends with a walk-through you can run against your own feature before it ships.
Seven places one request comes to rest
Assume a modest feature: an internal assistant that answers a support agent's question about a customer, using the account record plus a few retrieved help-centre and ticket documents. One question, one answer, a couple of seconds. Here is where that data physically lands on the way through.
| Hop | What is present | Where it can persist |
|---|---|---|
| Browser / client | The typed question, the rendered answer, the session token | Form autosave, browser extensions, session-replay recordings |
| Edge: CDN, WAF, gateway | URL, headers, sometimes the body | Access logs; WAF rule hits often store the matched payload |
| Your application process | Everything: input, system prompt, retrieved chunks, tool results, the answer | Process memory, plus whatever the process chooses to emit |
| Logs, traces, error reports | Whatever you passed to a logger, a span attribute, or an exception handler | Log index and trace store for the retention window; error reporter forever-ish |
| Retrieval store | Chunk text, its embedding, its metadata | Until explicitly deleted; the vector outlives the source row |
| Model provider | Serialised prompt, tool definitions, attachments, your key identity, timing | Depends entirely on tier and settings; see the next section |
| Tool callouts the model triggers | Whatever arguments the model chose, sent to whatever endpoint the tool wraps | The third party's logs, outside your control |
| Return path: caches, analytics | The answer, usually keyed by something derived from the input | Response cache, product analytics events, warehouse tables |
Two of those rows are the ones teams consistently forget, and they are the two with the loosest access control: the telemetry pipeline and the error reporter. The source database has row-level rules, a review process for new queries, and an audit trail. The log index has a team-wide role and a search box.
What crosses your boundary on a hosted model call
The thing that leaves your network is the serialised request, and it is bigger than people picture. It contains the system prompt, every message you chose to include from the conversation history, the full text of any retrieved passages, any attached files, and your tool or function definitions. That last item is worth pausing on: tool schemas describe your internal API surface, including field names, enum values, and sometimes table or entity names. Even with no customer data in the payload, you have described part of your system's shape to a third party.
Alongside the body goes transport metadata: the API key identity, which maps to your organisation and often to a specific service; the source IP; timestamps; and token counts, which are a rough measure of how much text you sent. TLS protects all of this from everyone except the endpoint. It does not protect it from the provider, and it cannot: inference runs on plaintext. "Encrypted in transit and at rest" is a true statement that answers a different question from the one you are asking, which is who can read it and for how long.
Not training on it and not keeping it are different promises
These get collapsed into one reassurance in internal discussions, and they are unrelated guarantees.
Not training on your data means your text will not influence the weights of future models. It says nothing about storage. A provider can honour this perfectly while holding your prompts on disk for a fixed window.
Not retaining your data means nothing durable is written after the response is returned, or that anything written is deleted on a short, stated schedule. This is a separate setting or contract tier in most places, not a default, and it interacts badly with features that exist because they store things: server-side conversation state, batch endpoints, file stores, prompt caching, and fine-tuning all need to keep your data by definition. If your bill went down because prompt caching kicked in, your prompt is being held somewhere for the life of that cache.
There is usually a third promise hiding behind both: an abuse-monitoring buffer, sometimes with a human-review path for flagged content. That is a legitimate safety mechanism and also a place your payload exists. Get the answer in writing for the specific endpoint and tier you call, and remember that a feature flag flipped by your own team can move you off it silently. The team that enabled server-side threads to simplify state management may not have realised they changed the retention answer.
The vendors behind your vendor
Count the organisations that could technically read one payload. Your model provider, and the cloud it runs on. Possibly a subcontractor for classification or support tooling. Then your side of the line: an observability SaaS, an error reporting service, a product analytics vendor, a data warehouse, a session-replay tool on the front end, and the support desk that receives escalated transcripts. For a typical AI feature that list is longer than the vendor list anyone would recite in a meeting, and most of the additions came in through a default SDK configuration rather than a decision. The honest answer to "where does this data go" is the union of your provider list and your telemetry list, and the way to produce it is to read what each SDK captures by default, not what its landing page says it is for.
The payload in your logs outlives the conversation
If you fix one thing after reading this, fix this one. Prompt and log hygiene is where real leaks happen, at a rate that dwarfs anything model-specific, because the mechanisms are all convenience features that default to on.
- HTTP client libraries log request and response bodies at debug level. Log levels get raised during incidents and lowered late.
- Tracing and LLM-observability SDKs auto-instrument model calls and attach the prompt and completion as span attributes. It is a headline feature, and it is a bulk export of your payloads to another vendor.
- Exception handlers serialise local variables and request bodies. The failing request is by definition the interesting one, and it carries everything.
- Retries and dead-letter queues persist message bodies, sometimes for days, in a queue nobody thinks of as a data store.
- Someone builds an eval or prompt-playground dataset out of production traffic. That is a copy of raw prompts living in whichever tool has the weakest permissions in the company.
The fix is to decide once what a model call is allowed to emit, and enforce it in one place. Identifiers and metrics, never payloads. This emitter records enough to debug an incident without storing the text.
import hashlib, jsonDENY = {"prompt", "messages", "input", "output", "completion", "documents", "tool_args", "attachments"}def log_llm_call(logger, *, request_id, tenant_id, user_id, model, prompt_text, usage, latency_ms, status): record = { "event": "llm.call", "request_id": request_id, "tenant": tenant_id, # stable per (tenant, user), useless outside our system "subject": hashlib.sha256( f"{tenant_id}:{user_id}".encode()).hexdigest()[:16], "model": model, "prompt_sha256": hashlib.sha256(prompt_text.encode()).hexdigest(), "prompt_chars": len(prompt_text), "input_tokens": usage["input_tokens"], "output_tokens": usage["output_tokens"], "latency_ms": latency_ms, "status": status, } assert DENY.isdisjoint(record), "payload field leaked into log record" logger.info(json.dumps(record))What survived is everything an on-call engineer actually uses: a request id to correlate, a subject hash to count affected users without naming them, token counts and latency for cost and performance, and a prompt hash. That hash does more work than it looks like: it tells you whether two complaints came from the same prompt, and whether the prompt you reproduced locally is byte-identical to the one that failed, without keeping the prompt. The assert is a cheap tripwire for the day someone adds a helpful debugging field. Then go turn off body logging in the HTTP client and set the tracing SDK's capture flags explicitly, because the defaults are not on your side.
Sometimes you genuinely need a payload to debug something. Make that a deliberate, per-request capture: opt-in by flag, short TTL, stored in a location with its own access control, and readable only through something that records who read it. A capture path that a support engineer can enable for one request is fine. A capture path that is on for all traffic is a second copy of your production data.
Cutting the request down before it leaves
Before reaching for a redaction library, ask a cheaper question: why is this field in the prompt at all? Most over-sharing happens because someone serialised an ORM object into the context. A summariser does not need the card number. A triage classifier does not need the customer's name, address, or lifetime value. Build the prompt from an explicit projection of named fields and never from a model instance, and you will remove more sensitive data than any detector ever will. It also fails safe: a new column added to the table does not silently start flowing to a third party.
For the text you do have to send, redaction of direct identifiers is worth doing. Here is a small, honest version. It swaps identifiers for stable placeholders and hands back a mapping so the answer can be rehydrated for the one user entitled to see the real values.
import re# Order matters: email first, or the phone rule eats part of it.PATTERNS = [ ("EMAIL", re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]{2,}")), ("CARD", re.compile(r"\b(?:\d[ -]?){13,19}\b")), ("PHONE", re.compile(r"\+?\d[\d ()-]{8,}\d")), ("IPV4", re.compile(r"\b(?:\d{1,3}\.){3}\d{1,3}\b")),]def redact(text): """Replace direct identifiers with stable placeholders. Returns (clean_text, mapping). The mapping is as sensitive as the original: keep it in memory for this request, never log it. """ mapping, seen = {}, {} for kind, rx in PATTERNS: def repl(m, kind=kind): raw = m.group(0) if raw not in seen: token = f"[{kind}_{len(seen) + 1}]" seen[raw] = token mapping[token] = raw return seen[raw] text = rx.sub(repl, text) return text, mappingRun that over a support email and the addresses and card-shaped digit runs come out as [EMAIL_1] and [CARD_2], consistently, so the model can still reason about "the same person" appearing twice. Now the limits, because this is where teams fool themselves.
A pattern detector finds shapes it already knows. It will not find identity carried in prose, and prose is where identity usually lives: "the night manager at the Croydon depot", a rare diagnosis next to a small town, a quoted internal ticket number that means something to anyone with access to the tracker. Re-identification rarely needs a name. It also over-matches in ways that quietly break your feature: an eleven-digit order number becomes [PHONE_1], and the model now answers about nothing. Measure both directions on real traffic before you trust it, because a detector with good recall and terrible precision produces a product bug that looks like a model problem. Machine-learned PII detectors do meaningfully better on names and addresses in free text, and still miss domain-specific identifiers such as policy numbers or device serials unless they were trained on yours. And note what the function returns: the mapping is a small, dense file containing exactly the values you decided were too sensitive to send. Treat it accordingly.
Pseudonymous identifiers, and the free-text field that defeats the schema
When the model needs to refer to an entity, give it a token that means nothing outside your system. user_8f31c2 rather than an email address. The payload in a provider's logs is then not directly identifying; you can still correlate two requests and they cannot; and if that prompt is later pasted into a ticket by a developer, it does not carry a person with it. Derive the token from a per-tenant salted hash rather than the raw primary key, so it is stable where you need stability but not joinable across contexts by anyone who has seen your ids elsewhere. Be precise about what this buys: pseudonymisation is reversible by design, which is the feature and also the ceiling. It reduces exposure at rest and in transit; it does not make the data anonymous.
Schema-based redaction works when data has fields. Free text is what breaks it. A notes column, a chat transcript, an uploaded PDF, a code comment: there is no field to skip, because the identifiers are inside sentences a human wrote for other humans. You have three honest options and one dishonest one. Do not send the free text at all, and run a classifier over structured features instead. Send only the span the user explicitly selected, so the scope is a user decision rather than a default. Or send it, accept the residual risk, and pay for it with tighter retention and access on every hop downstream. The dishonest option is to run a regex over it and describe the result as sanitised.
One index, many customers
Retrieval is where a privacy bug becomes a cross-customer bug, which is the kind you have to disclose. A vector index has no idea who owns a row unless you put ownership in the row. Build one index over "all the documents" and nearest-neighbour search will happily hand one customer's contract to another, with impeccable semantics: it really is the most similar chunk to the question that was asked.
The tempting shortcut is to retrieve globally and drop what the user cannot see. Filtering after retrieval is not the same as filtering during it, for three separate reasons. Your candidate budget gets consumed by documents the user will never receive, so the answer degrades and looks like a model quality problem. The forbidden chunks were nonetheless loaded into process memory, into a trace span, and possibly into a prompt-assembly buffer, so any new code path that skips the filter, including a debug endpoint or a "show sources" panel, exposes them. And the shape of the result set is itself a signal: a system that behaves differently when hidden matches exist tells a probing user that they exist.
Put the permission predicate inside the query that ranks. With pgvector that looks like this.
-- Wrong: neighbours first, permissions second.SELECT c.id, c.bodyFROM chunks cORDER BY c.embedding <=> :qLIMIT 8;-- Right: tenant and ACL are part of the ranked query.SELECT c.id, c.bodyFROM chunks cJOIN doc_acl a ON a.doc_id = c.doc_idWHERE c.tenant_id = :tenant AND a.group_id = ANY(:groups)ORDER BY c.embedding <=> :qLIMIT 8;The second query can never return a row the caller is not entitled to, which is the property you want. It is also slower, and you should know why: a restrictive filter forces an approximate index to explore more of its graph or fall back to a scan, so you are paying latency for correctness. Check what your engine actually does, because several apply the filter to candidates fetched by the approximate search rather than during it, which changes recall and sometimes changes safety. When tenants are large, a separate namespace or index per tenant is the sturdier answer, and it makes deletion and access review simpler too.
One more trap: the index holds a copy of permissions as they were at write time. Access changes, and the copy does not. Decide explicitly whether you re-check entitlements at query time against the source of truth, which is safer and slower, or accept staleness bounded by a sync job. A revoked user who can still retrieve documents for an hour is a real finding, not a rounding error. Whichever you choose, write an integration test with two tenants that asserts tenant A's query never returns a document id belonging to tenant B, and run it against every retrieval path you have, including the ones built for debugging.
Injection is an exfiltration path
The version of prompt injection worth engineering against is not the one that makes a chatbot say something rude. It is the one that moves data out. The setup is ordinary: your feature reads content it does not control. A web page, an inbound email, a customer-uploaded PDF, a ticket comment, a code comment in a repository you were asked to review. That content contains instructions, and the model has no reliable way to distinguish your instructions from the document's, because both arrive as text in one context window. Delimiters and "ignore any instructions in the document below" help at the margin and are not a control.
The damage requires a second ingredient: the same component holding an outbound channel. An HTTP fetch tool, a webhook, an email send, a "save to shared drive" action, or something as mundane as a markdown image the client will render, since the client resolving that URL performs the request for the attacker. Put untrusted text and an egress channel in the same agent, and the instruction and the mechanism are in the same process. Appending retrieved context to a query string is a complete exfiltration primitive.
The architectural rule is worth stating flatly: a component that reads attacker-controlled text must not hold the authority to send data out. That is a data-flow property, enforced by structure, and it survives a model upgrade in a way that prompt wording does not. In practice it means splitting the work. One step reads the untrusted content and may only return schema-validated fields, which your code then checks. A separate, privileged step acts on those fields and never sees the raw text. Egress tools take a destination from an allow-list resolved by your code, not from the model: the model may request "notify the ticket owner", and your code decides which address that is. Never let the model compose a URL that includes content it just read, and render model output so that remote images and links do not auto-load. For anything whose effect genuinely leaves your perimeter, put a human in front of it and show them the real destination and the real payload, not a summary of it.
Credentials do not belong in a prompt
API keys, database passwords, bearer tokens, signed URLs: not in the system prompt, not in a tool result the model can read, not in a retrieved document. The reasons stack. The prompt goes to a third party. The prompt tends to end up in a log. The model may reproduce it in output, including in an error message you show the user. And a successful injection turns "the model can see the token" into "the attacker has the token" in one step. The correct shape is that the model asks for an action and your code holds the credential and performs it, scoped to the calling user rather than to a service account with broad rights. If a secret does pass through a prompt, treat it as disclosed and rotate it, and check the git history of your prompt templates and eval fixtures while you are there, because that is where they hide.
What "delete my data" means once it is an embedding
Deletion is where AI features quietly break promises made elsewhere in the product. Removing the row from your primary database can leave the same information in at least six other places along the path we just walked: chunk text and its vector in the retrieval index, the payload in log and trace stores until their retention window rolls over, a copy inside an eval or golden dataset someone built from production traffic, the provider's retention window, a response cache keyed off the prompt, and backups. Backups are usually defensible if you can state their expiry; the others are just data you forgot about.
Embeddings deserve their own paragraph because they get treated as if they were hashes. They are not. A vector is derived data that preserves a great deal of the original meaning, and work on embedding inversion has repeatedly shown that a close paraphrase of short text can be recovered from its vector, which is more than enough to identify a person. Treat a stored vector as the text it came from, and make deletion delete it.
Fine-tuning is the one that has no clean remedy. If you trained on customer data, the model is downstream of it, and removing a training row does not remove its influence on the weights. Retraining is the only honest fix, which makes "may we fine-tune on this data?" a decision that is expensive to reverse. Decide it deliberately, at the start.
Make deletion mechanical rather than remembered. Key every derived store by the source record id, write one fan-out routine that hits all of them, and test it end to end: insert a record, let it flow into every store, issue the delete, then assert zero hits in the primary table, the index, the cache, and the search of your log store. A periodic sweep over the whole index is the design that rots, because the sweep drifts from reality and the new store nobody wired up is invisible until someone asks a question you cannot answer.
The pre-launch walk-through
Two people, ninety minutes, one real request traced end to end. This is the version I would run, in order, and each step has an outcome you can see rather than an opinion you can hold.
- Fire one request through the deployed feature and list every system that saw the payload. Include the CDN, the error reporter, and the front-end analytics. If the list is shorter than eight entries, you have missed something.
- Print the exact bytes you send to the provider, once, in a scratch script. Read it aloud. People find whole fields they did not know were in there.
- Write down, per endpoint and tier, whether the provider trains on it and whether it retains it, with a link to where that is stated. Two answers, not one.
- Search your log and trace stores for a distinctive string you put in the request. It should return nothing. Do the same in the error reporter after deliberately causing an exception mid-call.
- Check the HTTP client's body-logging setting and the tracing SDK's prompt-capture flag in the deployed configuration, not in the repository default.
- Diff the prompt against the record it came from and delete every field the task does not need. Then confirm the prompt is built from a field list, not a serialised object.
- Run your redactor over a few hundred real messages and count both misses and false hits. Decide whether the numbers justify the reliance you are placing on it.
- Run the two-tenant retrieval test against every path, including debug endpoints and any "show sources" view.
- Revoke a user's access to a document, then immediately query as that user. Note how long the stale answer persists and whether that window is acceptable.
- List every tool the model can invoke and mark which ones can move data outside your perimeter. For each one, name the component that reads untrusted text and confirm it is not the same component.
- Grep prompts, templates, tool results, and eval fixtures for anything shaped like a credential, including in git history.
- Delete a test record and verify zero hits in every derived store, one by one, by hand, this first time.
What usually goes wrong afterwards is not a new attack. It is drift: a new debug flag, a new tool with an outbound call, a new observability integration that captures payloads by default, a second retrieval path added for a different surface. Every one of those is a normal, well-intentioned change that reopens a hole you closed. So of the twelve checks above, the three worth converting into automated tests today are the log-payload assertion, the two-tenant retrieval test, and the deletion fan-out test. Those three fail loudly when someone reintroduces the problem, which is the only kind of privacy control that holds up over a year of shipping.