Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Sources, Freshness and Permissions


Meridian's second pilot was stopped by one line in a security review: "The assistant reads customer data through a single service account with access to all 2.4 million customers." Any staff member could ask about any customer, including their neighbour or a celebrity, and the assistant would answer. The core banking system's careful entitlements, built over twenty years, were bypassed by one new component.

The same pilot had a second, quieter problem. When a staff member asked about the late payment fee, it quoted the fee from a policy that had been withdrawn two months earlier. The old version was still in the shared folder the pilot had indexed.

Both failures are data architecture failures, not model failures. Before designing how context reaches the model, the architect answers three questions for every source: what is it, how fresh must it be, and who may see it. The answers form part A of MER-04, the data and context specification.

Who is the assistant when it reads a customer?Shared service account• One identity reads every customer• Assistantre-implements every entitlement• One bug exposes 2.4 million customers• Audit log says "assistant"On behalf of the staff member• The user's identitytravels with each call• Ledger and Atlas apply their own rules• A bug exposes only what that user sees• Audit log names the real person
The assistant should never see more than the person it serves, so identity passes through and restricted chunks are filtered before ranking.

The source inventory

Start with a complete list. Every source the assistant reads gets a row, with an owner who can answer questions about it. The trust column records whether the text in that source can contain words written by someone outside the bank. That matters for security in section 10.

SourceContentOwnerClassificationAccessTrust
PolicyHub260 policy documents, versionedPolicy teamInternalExport API, publish webhookTrusted
Ledger APIBalances, arrears, schedulesCore bankingCustomer confidentialREST via gatewayTrusted
Plan history viewHardship plans, last 5 yearsData platformCustomer confidentialSQL view, updated nightlyTrusted
Atlas case notesStaff notes and pasted customer emailsOperationsConfidential, may hold health dataRESTUntrusted
Vulnerability flagsRecorded vulnerability markersCustomer careSpecial categoryRESTTrusted
Hardship CalculatorEligibility and plan figuresCredit riskCustomer confidentialRESTTrusted
DocGen templatesApproved letter templatesComplianceInternalTemplate APITrusted

Case notes are marked untrusted even though staff write most of them. A single pasted email makes the whole field mixed authorship, and the pipeline cannot reliably tell which sentences a customer wrote. Treat the field by its worst content.

Freshness targets that follow harm

"As fresh as possible" is not a target. Each source gets a maximum staleness, chosen by asking: what harm does stale data cause here?

DataMaximum stalenessWhyMechanism
New or changed policy4 hours from publishNew rules usually have a start datePublish webhook triggers re-index
Withdrawn policy1 hour from withdrawalA withdrawn rule quoted with confidence misleads staffWithdrawal webhook deletes chunks first
Balance and arrears in a summary15 minutes, shown "as of"Payments can arrive during the dayRe-fetch on open if older
Figures in a letterZero: fetched at draft timeFigures go to the customerCalculator call per draft
Hardship plan history1 day, labelledPlans rarely change within a dayNightly view plus current plan from Ledger

Notice the asymmetry in the first two rows. Removing a withdrawn policy must be faster than adding a new one. A missing new policy produces "I could not find a policy on this", which is safe. A withdrawn policy still in the index produces a confident, wrong answer with a citation, which is the worst kind of failure. The ingestion pipeline therefore processes withdrawals before additions, and alerts if a withdrawal is not applied within the hour.

The plan history row shows another pattern: when data is allowed to be old, say so on screen. The summary prints "Plan history as of yesterday" and separately shows today's current plan from the live Ledger API. Staff can work with slightly old data if they know it is old.

The assistant sees what the user sees

The principle is short: at the moment it acts, the assistant must have no more access than the person it is acting for. Three mechanisms enforce it at Meridian.

Shared service account

  • One identity reads everything
  • The assistant must re-implement every entitlement rule
  • One bug exposes every customer
  • Audit logs show "assistant", not the person

On behalf of the user

  • The staff member's identity travels with each call
  • Ledger and Atlas apply their own existing rules
  • A bug exposes at most what that user could see
  • Audit logs show the real person

Identity pass-through. When a staff member asks a question, the orchestrator exchanges their sign-on token for a short-lived token that downstream APIs accept, using the standard OAuth token exchange pattern. Ledger and Atlas then apply the entitlements they already have.

Document permissions in the index. Twelve of the 260 policies are restricted, for example internal fraud procedures and credit risk thresholds. Each chunk carries the roles allowed to read it. Retrieval filters on the user's roles before ranking, so a restricted chunk never reaches the model for a user who cannot see it. Filtering after generation is too late; the model has already read it.

Case binding. The customer ID comes from the case open in Atlas, never from the text of a question. Staff can only open cases in their own queue, so the assistant inherits that rule.

One exception needs honest treatment. Summaries are precomputed when a hardship case opens, before any staff member is assigned. There is no user to act on behalf of. The design uses a separate background identity that can read only customers with an open hardship case, stores the summary against the case, and lets Atlas check the viewer's entitlement to the case before showing it. This is written up as an exception in the decision record, with its compensating controls, rather than hidden.

Minimise before the model

Some data should never reach a model at all. A scrubber runs on every record and note before it enters the context: patterns for card numbers, account numbers and national identity numbers, plus a small classifier for passwords and security answers, which staff sometimes type into notes by mistake. Its recall is measured on 500 labelled notes, like any other component.

Names, amounts and dates stay, because the summary and letter need them. Health mentions in notes stay in the model's input, because detecting possible vulnerability is one of the requirements, but they are masked in logs and traces. Every one of these choices appears in the data protection impact assessment, so the data protection officer can see what reaches the model and why.

With sources, freshness and permissions fixed, the next lesson designs how the chosen data is turned into the context a model actually reads.

Check your understanding

0 of 3 answered

1.Why must a withdrawn policy be removed from the index faster than a new policy is added?

2.Where should Meridian apply document permissions for the twelve restricted policies?

3.Summaries are precomputed before any staff member is assigned. How does the design handle permissions?