Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your enterprise search assistant must answer questions from PDFs, Slack, Jira, emails, and Confluence together. How do you unify retrieval across heterogeneous data sources with different schemas and permissions?
What you need to know
One record shape
Every source becomes the same record at ingestion:
1from pydantic import BaseModel2from datetime import datetime34class Chunk(BaseModel):5 doc_id: str6 source: str # "slack" | "jira" | "confluence" | "mail" | "pdf"7 text: str8 title: str9 url: str10 author: str11 updated_at: datetime12 acl_principals: list[str] # group ids, e.g. ["grp:finance", "user:u123"]13 thread_id: str | None = None14 source_meta: dict = {}Downstream code — chunking, embedding, search, citations — only knows this shape. Source differences live in connectors.
Connectors are incremental
| Source | How to pick up changes |
|---|---|
| Slack | Events API for new and edited messages |
| Jira | Webhooks on issue create and update |
| Confluence | Query pages by last-modified time |
| Email (Microsoft 365) | Microsoft Graph delta queries |
| PDFs in drives | The drive's change feed |
Changes go to a queue that chunks and embeds only what changed. Full re-crawls are too slow and too expensive at company scale.
Permissions: pre-filter, groups, deny by default
- Mirror ACLs at sync — store group ids, not expanded user lists, so a person joining a team does not force re-indexing.
- Resolve at query time — get the caller's groups from the identity provider (Okta, Entra ID).
- Pre-filter —
acl_principalsmust overlap the caller's groups, inside the search. - Deny by default — a chunk with a missing or stale ACL is excluded.
- Recheck the final top-K against the source system before display, to catch access removed since the last sync.
Ranking across sources
Use hybrid search (BM25 plus dense vectors), then a cross-encoder reranker across sources. Add per-source quotas before reranking — for example up to 10 candidates per source — or long PDFs take every slot and Slack answers never appear.
Measure recall@k per source, sync lag p95, and permission leaks, which must be zero and are tested in CI with adversarial queries: log in as a user outside a group and search for text only that group can see.
A real-life example
Scenario, numbers made up. A 3,000-person fintech builds one assistant over Confluence, Jira, Slack and policy PDFs. In the pilot, two problems show up. Answers almost never cite Slack, where most real decisions live, because 200-page PDFs fill the top 20. And a tester finds a message from a private leadership channel in her results.
The leak comes from a connector that indexed private channels with an empty ACL, which the filter treated as "public". The team switches to deny-by-default and adds the leak suite: 50 planted secrets, each queried by users who should not see them. With per-source quotas, Slack citations rise from 3% to 24% of answers, and thumbs-up rises with them.
Follow-up questions to expect
- "Why group ids and not user ids?" — Group membership changes daily. With group ids, the index stays the same and only the query-time lookup changes.
- "What about email, which is personal?" — Usually index only the caller's own mailbox, or leave email out of shared search. The ACL model still applies.
- "What if someone loses access between syncs?" — The final top-K recheck against the source catches it, and a short sync interval keeps the window small.