Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
You're building RAG for 500+ enterprise customers. A bug causes one tenant's chatbot to retrieve another's private data. How do you architect retrieval so this can never happen — even with a buggy query?
What you need to know
Bugs in query code are certain over time. The goal is a design where a buggy query cannot express "another tenant's data".
The three layers
- Partition — separate namespaces or collections per tenant, or one collection where the tenant key is mandatory and indexed. Crossing tenants requires naming a different partition.
- Derive the tenant from identity — the tenant id comes from the verified JWT claim at the edge, travels in a request context, and is injected by the data-access layer. No application function accepts a tenant id as a parameter.
- Assert after retrieval — every chunk stores
tenant_idin its payload; the retriever checks each result against the request's tenant and raises, alerts and returns nothing on a mismatch.
Layer 3 exists because layer 1 can be misconfigured. It turns a silent leak into a loud incident.
Choosing the partition model
| Model | Fits | Trade-off |
|---|---|---|
| Namespace per tenant (for example Pinecone namespaces) | Hundreds to many thousands of tenants | Strong logical separation inside one index |
| Shared collection with an indexed tenant key (for example Qdrant payload partitioning) | Many small tenants | Efficient; every query must carry the filter, so layers 2 and 3 matter more |
| Collection or cluster per tenant | A few large customers with contractual isolation | Most isolation, most operational cost; charge for it |
For 500 or more tenants, a separate collection each is usually too many objects to manage; vendors' multitenancy guides generally recommend namespaces or a tenant-keyed shared collection, with dedicated clusters for the few who pay for them.
The data-access layer
1class TenantRetriever:2 def __init__(self, auth: AuthContext): # built from the verified token3 self._tenant = auth.tenant_id45 def search(self, query_vec, k=10):6 hits = index.search(query_vec, k=k, namespace=self._tenant)7 for h in hits:8 if h.payload["tenant_id"] != self._tenant:9 alert("cross_tenant_hit", expected=self._tenant, got=h.payload["tenant_id"])10 raise CrossTenantError() # fail closed11 return hitsThe search method has no tenant parameter at all, so a caller cannot pass the wrong one.
Beyond retrieval
Caches must include the tenant in the cache key, or one tenant's cached answer can be served to another. Put the tenant id on every log line and trace so an audit can answer "who saw what". Use per-tenant encryption keys if contracts require it.
Prove it, continuously
Create two test tenants and give each a canary document containing a unique random string. A scheduled job asks tenant A's bot for tenant B's string every few minutes. If it ever comes back, page someone. That test turns a design claim into evidence you can show a customer's security team.
A real-life example
Scenario (illustrative numbers). An HR-software company serves 640 client companies from one vector index. A developer adds a "search across all my workspaces" feature and, in one code path, forgets the tenant filter. In staging, a test user at one client sees a salary-band document from another.
Because the post-retrieval assertion is in place, the request fails and an alert fires instead of leaking. The team then removes the tenant parameter from every retriever function, moves to namespaces per client, and adds a canary probe every 5 minutes across 20 pairs of test tenants. In the next year's security audit, the canary history showed zero cross-tenant hits across roughly 2 million probes.
Follow-up questions to expect
- "Isn't a metadata filter enough?" — It works until one code path forgets it. The point of the layers is that forgetting it is either impossible or caught.
- "How do you handle shared public documents?" — Put them in a separate shared partition and query it explicitly alongside the tenant's own, never by dropping the tenant filter.
- "What about the LLM provider seeing data?" — That is a separate control: data-processing agreements, regional endpoints or self-hosting, and no training on customer data.