Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
100 enterprise clients each upload their own documents. Client A must never see Client B's data, but you want one shared vector DB cluster to control costs. How do you enforce strict tenant isolation at the retrieval layer?
What you need to know
Filtering versus partitioning
There are two ways to keep tenants apart in one cluster.
Metadata filter only
- Every chunk has
tenant_id; every query adds a filter - One forgotten filter leaks everything
- Easy to add, cheap to run
- Isolation depends on every code path
Partition per tenant
- Each tenant's vectors live in a separate namespace or shard
- A query must name its tenant to run at all
- Deleting a tenant is one operation
- Isolation is enforced by the database
For 100 clients, partitioning is easy to operate. Common options in 2026:
| Store | Per-tenant unit |
|---|---|
| Pinecone | Namespace |
| Weaviate | Built-in multi-tenancy (one shard per tenant) |
| Milvus | Partition key, or a database/collection per tenant |
| Qdrant | Collection per tenant, or custom shard keys |
| Postgres + pgvector | Row-level security (RLS) policy on tenant_id |
With pgvector, RLS makes the database add the tenant condition to every query:
1ALTER TABLE chunks ENABLE ROW LEVEL SECURITY;2ALTER TABLE chunks FORCE ROW LEVEL SECURITY;3CREATE POLICY tenant_only ON chunks4 USING (tenant_id = current_setting('app.tenant_id')::uuid);5-- per request, inside a transaction: SET LOCAL app.tenant_id = '...';FORCE matters because a table owner otherwise skips RLS. Connect the app as a role that is not the owner.
Make the tenant id impossible to forge
The tenant id must come from the authenticated session — a claim in a verified JWT — never from a query parameter, a request body, or anything the LLM wrote. Put all vector access behind one repository module and block direct imports of the raw client with a lint rule.
1class TenantRetriever:2 def __init__(self, vdb, session):3 self.vdb, self.tenant = vdb, session.claims["tenant_id"] # verified token only45 def search(self, vector, k=8):6 hits = self.vdb.query(namespace=self.tenant, vector=vector, top_k=k)7 if any(h.metadata["tenant_id"] != self.tenant for h in hits):8 raise SecurityIncident("cross-tenant chunk returned") # should never fire9 return hitsThe caller cannot pass a tenant at all, so there is nothing to get wrong. The assertion is defence in depth: every chunk also stores tenant_id, and a mismatch is treated as a security incident, not a warning.
Prove it on every deploy
"Can never happen" needs a test, not a code review:
1def test_no_cross_tenant_leak(client):2 as_a = client.login("tenant-a-user")3 hits = as_a.search("ZEBRA-7731") # a string planted only in tenant B's docs4 assert hits == []Run it in CI on every deploy. Add per-tenant encryption keys if contracts ask for them, and per-tenant rate limits so one client cannot slow the cluster for the rest. Remember the other places data can leak: key any semantic cache and any logs by tenant too.
A real-life example
Scenario, numbers made up. An HR-tech SaaS serves 100 companies from one vector cluster holding 40M chunks. At first, isolation is a tenant_id filter added in the search function.
A developer adds an "admin search" endpoint for support staff and calls the vector client directly, without the filter. In staging, the new cross-tenant CI test fails: logged in as company A, a search returns a salary-policy chunk from company B. The team moves every tenant to its own namespace, makes the repository the only importable client, and derives the tenant from the token. The admin endpoint now has to name a tenant explicitly and is audit-logged. Offboarding a client drops from a day-long delete script to deleting one namespace.
Follow-up questions to expect
- "What changes at 100,000 tenants?" — A namespace each may hit store limits, so use a partition key with a filter enforced by the database, plus dedicated partitions for the few largest tenants.
- "Is the shared embedding model a risk?" — No. It is stateless: it turns text into vectors and keeps nothing. The risks are stored data, caches and logs.
- "How do you handle a user who belongs to two tenants?" — They pick one active tenant per session; the token carries that one id, and switching issues a new token.