Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

You're building RAG for 500+ enterprise customers. A bug causes one tenant's chatbot to retrieve another's private data. How do you architect retrieval so this can never happen — even with a buggy query?


Each layer can stop the leak on its ownTenant read from the verified tokenPartition: namespace or row policyCheck each result's tenant idCanary: A asks for B's secret string
The design goal is a query that cannot express another tenant's data, and a canary that proves it every few minutes.

What you need to know

Bugs in query code are certain over time. The goal is a design where a buggy query cannot express "another tenant's data".

The three layers

  1. Partition — separate namespaces or collections per tenant, or one collection where the tenant key is mandatory and indexed. Crossing tenants requires naming a different partition.
  2. Derive the tenant from identity — the tenant id comes from the verified JWT claim at the edge, travels in a request context, and is injected by the data-access layer. No application function accepts a tenant id as a parameter.
  3. Assert after retrieval — every chunk stores tenant_id in its payload; the retriever checks each result against the request's tenant and raises, alerts and returns nothing on a mismatch.

Layer 3 exists because layer 1 can be misconfigured. It turns a silent leak into a loud incident.

Choosing the partition model

ModelFitsTrade-off
Namespace per tenant (for example Pinecone namespaces)Hundreds to many thousands of tenantsStrong logical separation inside one index
Shared collection with an indexed tenant key (for example Qdrant payload partitioning)Many small tenantsEfficient; every query must carry the filter, so layers 2 and 3 matter more
Collection or cluster per tenantA few large customers with contractual isolationMost isolation, most operational cost; charge for it

For 500 or more tenants, a separate collection each is usually too many objects to manage; vendors' multitenancy guides generally recommend namespaces or a tenant-keyed shared collection, with dedicated clusters for the few who pay for them.

The data-access layer

Python
class TenantRetriever:    def __init__(self, auth: AuthContext):          # built from the verified token        self._tenant = auth.tenant_id    def search(self, query_vec, k=10):        hits = index.search(query_vec, k=k, namespace=self._tenant)        for h in hits:            if h.payload["tenant_id"] != self._tenant:                alert("cross_tenant_hit", expected=self._tenant, got=h.payload["tenant_id"])                raise CrossTenantError()           # fail closed        return hits

The search method has no tenant parameter at all, so a caller cannot pass the wrong one.

Beyond retrieval

Caches must include the tenant in the cache key, or one tenant's cached answer can be served to another. Put the tenant id on every log line and trace so an audit can answer "who saw what". Use per-tenant encryption keys if contracts require it.

Prove it, continuously

Create two test tenants and give each a canary document containing a unique random string. A scheduled job asks tenant A's bot for tenant B's string every few minutes. If it ever comes back, page someone. That test turns a design claim into evidence you can show a customer's security team.

A real-life example

Scenario (illustrative numbers). An HR-software company serves 640 client companies from one vector index. A developer adds a "search across all my workspaces" feature and, in one code path, forgets the tenant filter. In staging, a test user at one client sees a salary-band document from another.

Because the post-retrieval assertion is in place, the request fails and an alert fires instead of leaking. The team then removes the tenant parameter from every retriever function, moves to namespaces per client, and adds a canary probe every 5 minutes across 20 pairs of test tenants. In the next year's security audit, the canary history showed zero cross-tenant hits across roughly 2 million probes.

Follow-up questions to expect

  • "Isn't a metadata filter enough?" — It works until one code path forgets it. The point of the layers is that forgetting it is either impossible or caught.
  • "How do you handle shared public documents?" — Put them in a separate shared partition and query it explicitly alongside the tenant's own, never by dropping the tenant filter.
  • "What about the LLM provider seeing data?" — That is a separate control: data-processing agreements, regional endpoints or self-hosting, and no training on customer data.