Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
The Complete Meridian Design
Meridian's first pilot went to its review with a 60-slide deck. It had a beautiful architecture diagram, a demo video and a slide titled "Responsible AI" with six icons. The reviewers asked for accuracy evidence, a threat model and a monitoring plan. None were in the deck, because none existed. The review ended without a decision, which in practice meant no.
This time, the dossier took four days to assemble, because it was not written from scratch. Twelve documents already existed, each reviewed by its main reader as it was built. The architect's last job was to put a front section on them that lets a busy board member understand the design in twenty minutes, and lets a specialist jump straight to the evidence they care about.
This lesson assembles MER-13, the design dossier. The next lesson defends it.
Two kinds of reader
A review board contains two kinds of reader, and the dossier must serve both. Generalists, such as the Chief Risk Officer, read the first few pages carefully and trust that specialists have checked the rest. Specialists, such as the head of model risk or the CISO's delegate, skip the front and go straight to their annex, looking for weaknesses.
So the dossier has a short front section and a long back section.
| Part | Length | Reader | Content |
|---|---|---|---|
| Executive summary | 1 page | Everyone | Decision requested, value, safety, residual risk |
| Architecture on a page | 1 page | Everyone | Components, flows, key decisions |
| Traceability matrix | 2 pages | Specialists, audit | Requirement to owner, end to end |
| Residual risks and owners | 1 page | Risk, business owner | From the register, with who accepts each |
| Release plan and gates | 1 page | Everyone | Stages, gates, dates, stop points |
| Open items and proposed conditions | 1 page | Board | What is not finished, and how it will be |
| Annexes MER-01 to MER-12 | As built | Specialists | The design record itself |
Seven pages of front matter. Everything in them links by ID to an annex, so nothing in the front makes a claim without evidence behind it.
The executive summary
The summary is the most important page in the dossier and the hardest to write, because it must be short, complete and honest at the same time. Here is Meridian's, in full.
# Loan servicing assistant: design approvalDecision requested: approve stage 1 (policy answers and account summaries)for build and release; note stage 2 (letter drafts) for separate approvalafter its pilot gate and a risk committee decision.What it does: helps 450 servicing staff find cited policy answers and seea customer's account history on one screen. It never decides eligibility,never sends communications and never writes to core banking.Value: about 170 staff hours a day saved in the pilot by stage 1. Hard benefit about$0.93m a year in avoided hiring (stage 1). Run cost about $0.43m a year.Three-year NPV +$0.26m. Adoption is the assumption that matters most.Safety, in three lines:- Figures, eligibility and customer identity are enforced by code, not by the model.- Quality is measured continuously; answers fall back to search if it drops.- Every output is traceable to its sources, versions and approver.Residual risk: staff will sometimes receive a wrong policy answer (R-01,medium). Pilot staff error rate with the assistant was 7%, against 18%without it. Accepted by the Head of Servicing.Letters: reduce letter errors (7% to a 2% target) but show a three-yearNPV of about -$0.1m. Justified by customer outcomes; referred to risk committee.Notice what the summary does. It states the decision in the first line. It separates stage 1 from stage 2, because they carry different risks and different returns. It gives numbers with their source. It names the weakest part, R-01, and the uncomfortable result, the letters NPV, on the first page, before anyone has to dig for them.
The architecture on a page
Atlas panel: answers with citations, summary, letter review screen |Orchestrator: policy_answer | account_summary | letter_draft (workflows, release bundles) | \AI gateway: auth, routing, quotas, metering Tool layer: policy-search (MCP), |-- provider A, in-region (primary) account-read (MCP), Hardship Calculator, |-- provider B (evaluated fallback) DocGen drafts (direct, idempotent) |-- guard and rewrite, 8B, self-hosted | | Systems of record via API gateway:Cross-cutting: tracing, audit store, Ledger, Atlas, PolicyHub, plan historyevaluation harness, kill switches Events: hardship.case.opened, policy.published, policy.withdrawnBeneath the diagram, three short paragraphs walk one request through each capability, because a reviewer understands a system faster by following a request than by reading boxes.
A policy answer starts in the Atlas panel. The self-hosted model checks the question and rewrites it into a search query; hybrid search finds 40 candidates the staff member is entitled to see, filtered by effective date; the reranker keeps six; the hosted model answers within a 6,000-token budget; code checks that every citation resolves to a current section; and the answer streams back with dates, in about 2.3 seconds to first text at p95.
An account summary starts before anyone asks. The hardship.case.opened event triggers a worker that fetches records in parallel, builds the must-include list, generates the narrative and checks it. When the specialist opens the case, live figures are fetched fresh beside the stored narrative.
A letter draft starts with a human choice of plan. The calculator produces every figure; the model writes the explanation around them; code checks figures, mandatory paragraphs and customer identity; a draft is created idempotently and older drafts are superseded; a named person reviews and approves; DocGen sends.
Under those walkthroughs sits a table of the decisions a reviewer is most likely to question, each with its record number: ADR-004 (never sends letters), ADR-007 (workflow, not agent, for summaries), ADR-010 (effective-date filter), ADR-011 (precompute summaries on events), ADR-012 (hybrid topology), ADR-013 (no model fallback for letters) and ADR-014 (customer ID bound in code). A reviewer who disagrees with any of them can read the alternatives that were rejected and why.
The traceability matrix
The traceability matrix follows every important requirement across the design record. Its main use is not for reviewers; it is for the architect, to find gaps before they do.
| Requirement | Decision | Control | Test | Objective or invariant | Owner |
|---|---|---|---|---|---|
| R-POL-01 correct, cited answers | ADR-009, ADR-010 | Citations with dates | Golden set, challenge set | 90% sampled correctness | Service owner |
| R-POL-02 declines when unsupported | Decline rule, score threshold | Search mode offered | 40 out-of-scope cases | 90% correct declines | Policy lead |
| R-SUM-01 every material event | Must-include checklist | Code check, fallback lines | Layer 1, 120 cases | 98% first generation | Eng lead |
| R-SUM-02 vulnerability surfaced | Excerpts, never set flags | Human decides | 80 labelled notes | Recall 95% | Vulnerable customers lead |
| R-LET-01 figures equal calculator | Calculator figures only | Invariant check, supersede, approval check | Layer 1, integration tests | Invariant, zero | Head of Servicing |
| R-LET-03 cannot send | No send tool | Tool matrix | Access test | Invariant | Security |
| R-SEC-01 entitlements enforced | Delegated identity | Role pre-filter, case binding | Access tests per role | Invariant | Security |
| R-NFR-01 answer latency | Six chunks, small rewrite model | Breaker, streaming | Load test | p95 2.5 s to first text | Eng lead |
Building the matrix found one real gap. R-SUM-02, surfacing vulnerability mentions, had a pre-launch test but no production measure: nothing would show if recall slipped after launch. The fix was a monthly sample of 40 new notes labelled by the vulnerable customers team and added to the quality objectives. That gap would have been an easy, and embarrassing, question in the review.
Release plan, open items and conditions
- Month 0 — board approval of stage 1.
- Months 1 to 6 — build, with layer 1 to 3 evaluation on every change; red team before pilot.
- Month 6 — pilot with 30 staff for four weeks; shadow and canary on the release process.
- Month 7 — release 1 to all 450 staff in waves of about 100 a week.
- Month 10 — gate 1: adoption of 70% weekly and the correctness objective met.
- Months 10 to 11 — letter pilot on 300 drafts; separate approval for stage 2.
Open items are listed plainly, each with an owner and a date: the final legal classification memo for letters, the signed contract with provider B, and an ISO/IEC 42001 gap assessment for the platform. The dossier also proposes the conditions the architect expects the board to want, such as a quarterly report on objectives, invariants, incidents, adoption and drill results. Proposing them first shows the board the team has thought about oversight after launch, and it tends to make the discussion shorter.
The dossier is ready. What remains is the forty minutes in the room, which the next lesson prepares.
Check your understanding
0 of 3 answered
1.Why does Meridian's executive summary mention the letters' negative three-year NPV on the first page?
2.What is the traceability matrix most useful for?
3.Why does the dossier propose oversight conditions to the board itself?