Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your AI assistant goes global. EU GDPR blocks sending European data to US models. India's DPDP adds more rules. How do you architect a multi-region LLM app that respects local data laws?
What you need to know
Three different obligations
| Requirement | Meaning | Example |
|---|---|---|
| Data residency | Personal data is stored and processed only in a region | An EU bank's contract requires EU-only processing |
| Transfer rules | Data may leave only with a legal basis | GDPR allows transfers to the US via the EU-US Data Privacy Framework or standard contractual clauses |
| Processing duties | Notice, consent, purpose limits, deletion | India's DPDP Act 2023, with its Rules phasing in from late 2025 |
Precision matters in an interview: GDPR does not flatly block US models, but transfers need a legal basis and paperwork, and many enterprise customers simply require residency. India's DPDP Act generally permits transfers abroad unless the government restricts a country, but sector rules can be stricter — RBI requires payment system data to be stored only in India.
Region as a hard boundary
- Tag each tenant — store a residency region at sign-up; it is data, not a guess.
- Route by the tag — a German customer travelling in Singapore still lands in the EU stack; never route by IP.
- Pin model endpoints — in-region cloud endpoints (EU data zones, AWS Bedrock in
eu-*orap-south-1), or self-hosted open models on in-region GPUs; confirm zero-retention and no-training terms in the contract. - Keep logs in region — traces contain full prompts and retrieved documents, so self-host tracing per region or log hashes and metrics only.
- Make deletion one operation — keep a map from user to every record, vector, cache entry, checkpoint and dataset that holds their data.
1REGIONS = {2 "eu": {"api": "https://eu.api.example.com", "llm": "bedrock:eu-central-1", "vdb": "vdb-eu"},3 "in": {"api": "https://in.api.example.com", "llm": "bedrock:ap-south-1", "vdb": "vdb-in"},4 "us": {"api": "https://us.api.example.com", "llm": "bedrock:us-east-1", "vdb": "vdb-us"},5}67def region_for(tenant_id: str) -> dict:8 region = control_plane.get_residency(tenant_id) # stored tag, not client IP9 if region not in REGIONS:10 raise PermissionError("no residency tag: refuse rather than guess")11 return REGIONS[region]Refusing when the tag is missing is deliberate: a default to the US region is exactly the bug an auditor finds.
The model catalogue differs by region
Not every model is offered in every region. Your golden eval set has to run per region, and each region may need its own model choice or a self-hosted fallback.
A real-life example
Scenario, numbers made up. A B2B support-assistant company signs a German insurer and an Indian NBFC in the same quarter. The insurer's contract requires EU-only processing; the NBFC's data includes payment information covered by RBI localisation.
The team splits into three cells: EU (Frankfurt), India (Mumbai) and US. Model calls go to in-region endpoints; the India cell uses a self-hosted open model for payment-related flows. Tracing is self-hosted in each region. A test deletes a synthetic user and verifies that 11 stores — database rows, vectors, caches, checkpoints, eval sets and backups on schedule — no longer hold the data. The German insurer's audit passes on the first round.
Follow-up questions to expect
- "Why not one global stack with encryption?" — Encryption protects data at rest, but the model still processes plaintext wherever it runs, which is what residency rules care about.
- "What about fine-tuning or evals across regions?" — Train and evaluate per region on that region's data, or use anonymised or synthetic data that is legally non-personal.
- "How do you handle backups?" — Keep them in region and include them in the deletion plan, usually through expiry within a documented period.