Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

The AI Platform and Operating Model


By the time Meridian's assistant reached pilot, eleven more AI ideas were waiting: complaints triage, a broker assistant, a summariser for fraud case notes, a meeting-notes tool for relationship managers, and seven others. If each team built its own gateway, guard, tracing, evaluation tooling and threat model, the bank would have twelve of everything, twelve different security stories and twelve arguments with model risk.

The opposite mistake is just as common. A central AI team that must approve every prompt change and build every feature becomes a queue. Ideas wait months, teams route around it, and shadow AI grows in spreadsheets and browser tabs.

This lesson designs what Meridian shares and what it does not, how decisions are made after go-live, and how it is paid for. It produces part B of MER-12, the operating model, and it is what turns one successful assistant into a capability the bank can repeat.

Platform owns how; the team owns whatShared platform• Gateway, keys, quotas and metering• Model catalogue and life cycle sheets• Evaluation harness and judge calibration• Guard, output filter, tracing, auditUse-case team• Workflows, prompts and release bundles• Golden sets and domain checks• User interface and adoption• Quality objectives and residual risk
Built once and assured once, the paved road took the second use case from five months to seven weeks without lowering the bar.

Platform or product: what goes where

The dividing line is simple to state: the platform owns how AI capabilities are provided safely; the product team owns what the capability does and how good it is for its use.

Shared platform

  • AI gateway: keys, routing, quotas, cost metering
  • Model catalogue with life cycle sheets
  • Evaluation harness, judge calibration tools, dataset storage
  • Guard classifier, output filter, scrubber
  • Tracing and audit store
  • Registry of MCP servers, and the shared ones

Use-case team

  • Workflows, prompts and release bundles
  • Golden sets and domain checks, such as figure checks
  • The user interface and adoption
  • Quality objectives and their error budgets
  • Business ownership and residual risk

Everything on the left is built once, secured once and assured once. Everything on the right depends on domain knowledge that only the use-case team has. The platform team never writes a use case's prompts; the use-case team never runs its own model keys.

What the platform costs, and when it pays

Sharing is not free. Someone has to build and run the shared components, and the bank pays for that team whether one use case uses it or twelve. Here is Meridian's platform budget for its first full year, at the same loaded rate MER-11 uses for engineers, $150,000 a year.

ComponentPlatform peoplePlatform cost per yearOne team building its own, per year
AI gateway: keys, routing, quotas, metering1.0$180,000$55,000
Model catalogue and life cycle sheets0.5$75,000$25,000
Evaluation harness, judge calibration, dataset store1.5$245,000$80,000
Guard, output filter, scrubber1.0$160,000$65,000
Tracing and audit store0.5$90,000$40,000
MCP registry and shared tool servers1.0$160,000$55,000
Head of AI platform1.0$150,000Not needed
Total6.5$1,060,000$320,000

The platform column is people plus $85,000 of base infrastructure. The right-hand column is what one use-case team would spend on its own copy: the build spread over three years, the same horizon as the business case, plus upkeep and the security and model risk reviews that every copy needs. Metered costs, such as tokens and GPU time, are left out of both columns, because a use case pays for them through chargeback either way. The $12,000 gateway share in MER-11 is one of those charges.

The platform also spends about $20,000 a year on each use case it serves: onboarding, quota set-up and help preparing evidence for the review board. So with N live use cases, the platform costs $1,060,000 + $20,000 × N, and building alone costs $320,000 × N. They are equal at N = 1,060,000 ÷ 300,000, about 3.5. The platform pays for itself from the fourth use case.

Live use casesShared platformEach team builds its ownDifference
1$1,080,000$320,000Platform costs $760,000 more
2$1,100,000$640,000Platform costs $460,000 more
4$1,140,000$1,280,000Platform saves $140,000
8$1,220,000$2,560,000Platform saves $1,340,000

Be honest about the first year. With only the servicing assistant and complaints triage live, the platform costs $460,000 more than letting both teams build alone. Meridian funded it anyway, and wrote down why: eleven ideas were already waiting, and the costs this table leaves out, twelve security stories and twelve arguments with model risk, are the ones that hurt a bank most. It also wrote a revisit condition, as it would for any design decision: if fewer than four use cases are live 18 months after the platform starts, shrink it to the gateway, tracing and evaluation harness.

The paved road

The payoff of a platform is the paved road: if a team builds on the approved components, it inherits their controls and the evidence already accepted for them. Security reviewed the gateway, the output filter and the audit store once; model risk accepted the evaluation harness and judge calibration method once. A new use case on the paved road only needs to show evidence for what is specific to it.

Not every idea needs Meridian's full thirteen-document design record. The review board uses three tiers, set at intake from impact and autonomy.

TierTypical useDesign record requiredTypical time to approval
3Internal, no customer data, L1Charter, data sheet, evaluation plan, threat checklist2 to 3 weeks
2Customer data, staff decide, L1 or L2Full design record, lighter business case6 to 10 weeks
1Affects customers directly, L3 or more, or possibly high-riskFull design record, independent validation before pilot3 to 6 months

Proportionality is what keeps people on the paved road. If a meeting-notes tool for twenty relationship managers needs the same paperwork as a credit decision system, teams will go around the process, and the bank will have less control, not more.

Decisions after go-live

The decision-rights table from section 1 covered building the system. Running it needs its own. The bank uses the familiar three lines of defence: the business and delivery teams own and manage risk, second-line risk and compliance functions (including model risk) oversee and challenge, and internal audit gives independent assurance.

DecisionService ownerBusiness ownerPlatform ownerModel riskSecurityInternal audit
Non-material prompt changeAIIII
Model version changeRACCC
New capabilityRACCCI
SEV1 incident responseAIRIRI
Quarterly monitoring reviewRACC
Annual revalidationCICAI
Platform component upgradeCIACC

The table says who decides each thing. Some decisions need several of those people at once: new use cases at tiers 1 and 2, material changes, and exceptions to the paved road. Those go to the AI review board. A board that is vague about what it needs and when it will answer becomes the queue this lesson warned about, so Meridian wrote its terms of reference as tightly as a service contract.

ElementMeridian's terms of reference
Voting membersCRO's delegate (chair), CISO's delegate, head of model risk, data protection officer, head of AI platform
Attending, not votingThe business owner, architect and finance partner of each item; a platform team secretary who keeps the decision log
QuorumFour of the five, always including model risk and security
CadenceMonthly, 90 minutes, at most three items: up to 40 minutes for a new use case, 20 for a change or exception; three members can sit out of cycle for urgent items
InputsIntake form with proposed tier; the design record documents the tier requires; evidence register; risk register entries with residual ratings signed by the business owner; platform components used and any exceptions; unit cost and business case summary
Paper deadlineSeven days before the meeting; an incomplete pack is not tabled
Decisions it can makeApprove; approve with conditions, each with an owner and a date; defer, naming the missing evidence; reject; pause a live capability
Decisions it cannot makeAccept residual risk for the business owner; release funding; redesign the system in the meeting

The last row matters as much as the one above it. Residual risk belongs to the business owner, through the risk register. Funding belongs to finance, through the stage gates in MER-11. Design belongs to the team. Keeping the board out of all three is what lets it finish in 90 minutes.

The board also publishes service levels, because teams only stay on the paved road if they can plan around it. Intake is triaged and a tier proposed within 5 working days. A tier 3 idea gets a written decision from the head of AI platform and the CISO's delegate within 10 working days, with no meeting. A tier 1 or 2 item is decided at the next meeting after its pack is complete, so never more than five weeks later. An urgent item, such as a model retirement notice, is decided out of cycle within 5 working days. Each quarter the board reports its own numbers: median days from complete pack to decision, and the share of items deferred for missing evidence. A rising deferral rate usually means the intake guidance is unclear, not that teams are careless.

Funding and chargeback

How the platform is paid for shapes how teams behave. Meridian chose three rules.

  • Base platform funded centrally. The first team to use a component should not pay for building it, or nobody will go first.
  • Variable costs charged back by metered use. The gateway already meters tokens and GPU time per use case, so each team sees its real consumption and cost per unit, and pays for it.
  • Evaluation runs are free to teams. If testing costs a team money, teams test less. Meridian wants the opposite, so evaluation traffic is billed to the platform budget.

People matter as much as components. All 450 servicing staff receive three hours of training before access: what the assistant does and does not do, how to read citations, how to review a draft, and how to report a wrong answer. Team leads receive a further half day on reading the review metrics. That programme is also how Meridian meets the EU AI Act's expectation that staff operating AI systems have sufficient AI literacy.

Part B of MER-12 records the operating model.

YAML
id: MER-12-Bplatform:  owner: head of AI platform  components: [ai_gateway, model_catalogue, eval_harness, guard, output_filter,               scrubber, tracing, audit_store, mcp_registry, policy_search, account_read]  people: 6.5  cost_per_year: 1060000      # people plus base infrastructure  per_use_case_support: 20000  break_even: 4 live use cases (building alone costs 320000 a year per team)  revisit: shrink to gateway, tracing and eval harness if fewer than 4 live at month 18  funding: central base; variable costs charged back by gateway metering  eval_runs: free to use-case teamsservicing_assistant:  service_owner: servicing assistant product lead  business_owner: Dana Whitfield, Head of Loan Servicing  tier: 2, letters with extra conditions  monitoring_review: quarterly, with model risk  revalidation: annual, and on material changereview_board:  chair: CRO delegate  members: [CRO delegate, CISO delegate, head of model risk, DPO, head of AI platform]  quorum: 4 of 5, including model risk and security  cadence: monthly, 90 minutes, at most 3 items; out of cycle with 3 members if urgent  papers_due: 7 days before the meeting  approves: [tier 1 and 2 use cases, material changes, paved-road exceptions]  outcomes: [approve, approve with conditions, defer, reject, pause]  service_levels: {intake: 5 working days, tier_3: 10 working days,                   tier_1_2: next meeting after a complete pack, urgent: 5 working days}exceptions:  - fraud case-notes summariser: vendor model outside the gateway; 3 conditions; review at 12 monthstraining:  all_staff: 3 hours before access  team_leads: extra half day on review metrics

Check your understanding

0 of 3 answered

1.Which of these belongs to the use-case team rather than the shared platform?

2.Why does Meridian use three tiers instead of requiring the full design record for every AI idea?

3.Why are evaluation runs free to use-case teams while tokens for production are charged back?