Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your enterprise search assistant must answer questions from PDFs, Slack, Jira, emails, and Confluence together. How do you unify retrieval across heterogeneous data sources with different schemas and permissions?


What you need to know

One record shape

Every source becomes the same record at ingestion:

Python
from pydantic import BaseModelfrom datetime import datetimeclass Chunk(BaseModel):    doc_id: str    source: str                 # "slack" | "jira" | "confluence" | "mail" | "pdf"    text: str    title: str    url: str    author: str    updated_at: datetime    acl_principals: list[str]   # group ids, e.g. ["grp:finance", "user:u123"]    thread_id: str | None = None    source_meta: dict = {}

Downstream code — chunking, embedding, search, citations — only knows this shape. Source differences live in connectors.

Connectors are incremental

SourceHow to pick up changes
SlackEvents API for new and edited messages
JiraWebhooks on issue create and update
ConfluenceQuery pages by last-modified time
Email (Microsoft 365)Microsoft Graph delta queries
PDFs in drivesThe drive's change feed

Changes go to a queue that chunks and embeds only what changed. Full re-crawls are too slow and too expensive at company scale.

Permissions: pre-filter, groups, deny by default

  1. Mirror ACLs at sync — store group ids, not expanded user lists, so a person joining a team does not force re-indexing.
  2. Resolve at query time — get the caller's groups from the identity provider (Okta, Entra ID).
  3. Pre-filter — acl_principals must overlap the caller's groups, inside the search.
  4. Deny by default — a chunk with a missing or stale ACL is excluded.
  5. Recheck the final top-K against the source system before display, to catch access removed since the last sync.

Ranking across sources

Use hybrid search (BM25 plus dense vectors), then a cross-encoder reranker across sources. Add per-source quotas before reranking — for example up to 10 candidates per source — or long PDFs take every slot and Slack answers never appear.

Measure recall@k per source, sync lag p95, and permission leaks, which must be zero and are tested in CI with adversarial queries: log in as a user outside a group and search for text only that group can see.

A real-life example

Scenario, numbers made up. A 3,000-person fintech builds one assistant over Confluence, Jira, Slack and policy PDFs. In the pilot, two problems show up. Answers almost never cite Slack, where most real decisions live, because 200-page PDFs fill the top 20. And a tester finds a message from a private leadership channel in her results.

The leak comes from a connector that indexed private channels with an empty ACL, which the filter treated as "public". The team switches to deny-by-default and adds the leak suite: 50 planted secrets, each queried by users who should not see them. With per-source quotas, Slack citations rise from 3% to 24% of answers, and thumbs-up rises with them.

Follow-up questions to expect

  • "Why group ids and not user ids?" — Group membership changes daily. With group ids, the index stays the same and only the query-time lookup changes.
  • "What about email, which is personal?" — Usually index only the caller's own mailbox, or leave email out of shared search. The ACL model still applies.
  • "What if someone loses access between syncs?" — The final top-K recheck against the source catches it, and a short sync interval keeps the window small.