Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your organization wants every employee to have personalized AI assistants trained on their work style and documents. How do you architect secure, personalized LLM systems at enterprise scale?


What you need to know

Why not one fine-tuned model per person

Serving many small adapters is technically possible — engines such as vLLM can serve many LoRA adapters on one base model. The problems are elsewhere. Training data would include confidential documents, which the model can memorise and repeat. When an employee leaves or asks for deletion, removing their data from weights is not practical; you retrain. And thousands of adapters mean thousands of things to evaluate.

Personalise through context instead

LayerWhat it holdsWhy it is safe
Identity-scoped retrievalDocuments the user can open in the source systemsPermissions mirrored from the source, enforced as a pre-filter
Structured profileRole, team, projects, terms, preferred lengthVisible, editable, versioned, deletable
Style examplesTwo or three of the user's own documents, used as few-shot examplesRemoved instantly if the user wants
Org-level adapter (optional)Company jargon and formatsTrained on approved, non-personal data

The classic leak: shared caches

A semantic cache stores answers to similar questions. If a finance director asks "What are Q3 bonus pools?" and the answer is cached globally, the next similar question from an intern returns it. Key every cache by the principal set:

Python
import hashlibdef cache_key(user, question_embedding_bucket: str, model_version: str) -> str:    principals = sorted([user.id, *user.groups])    scope = hashlib.sha256("|".join(principals).encode()).hexdigest()[:16]    return f"sem:{model_version}:{scope}:{question_embedding_bucket}"

Only users with exactly the same permissions share a cache entry. Include the model version so an upgrade does not serve old answers.

  1. Scope retrieval by identity — groups from the identity provider, enforced inside the search.
  2. Build the profile — structured, shown to the user, editable.
  3. Add style examples — the user's own documents, chosen by them.
  4. Scope caches and logs — by principal set; log what was retrieved for whom.
  5. Govern — data residency, deletion of profile and embeddings within an agreed time, leak tests in CI.

A real-life example

Scenario, numbers made up. A 12,000-person consulting firm plans per-consultant fine-tuning. The estimate: weeks of GPU time per refresh, no clean way to honour deletion, and client documents inside weights. They switch plans.

The assistant uses one model with permission-filtered retrieval over SharePoint and email, an editable profile, and style examples. In the pilot, 18% of users edit their profile in the first week — mostly correcting their current project — which shows why it must be visible. A CI leak test finds that the new semantic cache was keyed only by question; it is re-keyed by principal set before launch. Deletion requests are completed within a day, because nothing personal lives in the weights.

Follow-up questions to expect

  • "Won't few-shot style examples cost tokens?" — Some, but two or three short samples in a cached prefix are cheap compared with training and maintaining per-user models.
  • "How do you handle someone changing teams?" — Permissions update from the identity provider; the profile flags stale projects for the user to confirm.
  • "How would you prove there are no cross-user leaks?" — Planted secrets per group and automated queries from users outside that group, run on every deploy.