Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your organization wants every employee to have personalized AI assistants trained on their work style and documents. How do you architect secure, personalized LLM systems at enterprise scale?
What you need to know
Why not one fine-tuned model per person
Serving many small adapters is technically possible — engines such as vLLM can serve many LoRA adapters on one base model. The problems are elsewhere. Training data would include confidential documents, which the model can memorise and repeat. When an employee leaves or asks for deletion, removing their data from weights is not practical; you retrain. And thousands of adapters mean thousands of things to evaluate.
Personalise through context instead
| Layer | What it holds | Why it is safe |
|---|---|---|
| Identity-scoped retrieval | Documents the user can open in the source systems | Permissions mirrored from the source, enforced as a pre-filter |
| Structured profile | Role, team, projects, terms, preferred length | Visible, editable, versioned, deletable |
| Style examples | Two or three of the user's own documents, used as few-shot examples | Removed instantly if the user wants |
| Org-level adapter (optional) | Company jargon and formats | Trained on approved, non-personal data |
The classic leak: shared caches
A semantic cache stores answers to similar questions. If a finance director asks "What are Q3 bonus pools?" and the answer is cached globally, the next similar question from an intern returns it. Key every cache by the principal set:
1import hashlib23def cache_key(user, question_embedding_bucket: str, model_version: str) -> str:4 principals = sorted([user.id, *user.groups])5 scope = hashlib.sha256("|".join(principals).encode()).hexdigest()[:16]6 return f"sem:{model_version}:{scope}:{question_embedding_bucket}"Only users with exactly the same permissions share a cache entry. Include the model version so an upgrade does not serve old answers.
- Scope retrieval by identity — groups from the identity provider, enforced inside the search.
- Build the profile — structured, shown to the user, editable.
- Add style examples — the user's own documents, chosen by them.
- Scope caches and logs — by principal set; log what was retrieved for whom.
- Govern — data residency, deletion of profile and embeddings within an agreed time, leak tests in CI.
A real-life example
Scenario, numbers made up. A 12,000-person consulting firm plans per-consultant fine-tuning. The estimate: weeks of GPU time per refresh, no clean way to honour deletion, and client documents inside weights. They switch plans.
The assistant uses one model with permission-filtered retrieval over SharePoint and email, an editable profile, and style examples. In the pilot, 18% of users edit their profile in the first week — mostly correcting their current project — which shows why it must be visible. A CI leak test finds that the new semantic cache was keyed only by question; it is re-keyed by principal set before launch. Deletion requests are completed within a day, because nothing personal lives in the weights.
Follow-up questions to expect
- "Won't few-shot style examples cost tokens?" — Some, but two or three short samples in a cached prefix are cheap compared with training and maintaining per-user models.
- "How do you handle someone changing teams?" — Permissions update from the identity provider; the profile flags stale projects for the user to confirm.
- "How would you prove there are no cross-user leaks?" — Planted secrets per group and automated queries from users outside that group, run on every deploy.