Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Hosted API, Self-Hosted or Hybrid


The first time the architect showed Meridian's security team a diagram with a hosted model on it, the reaction was immediate: "Customer data cannot leave the bank. We have to run the model ourselves." A week later, a finance partner looked at the per-token pricing and said the opposite for a different reason: "Why rent intelligence by the word when we could buy GPUs and own it?"

Both reactions are reasonable starting points. Both are also claims that need numbers. Where the model runs is one of the most expensive decisions in the design to reverse, and it touches security, cost, quality, resilience and regulation all at once.

This lesson compares the three topologies with Meridian's real traffic, shows how the bank made a hosted model acceptable for customer data, and records the result as part A of MER-05.

Weighted decision matrix, scores out of 54555355245142534.353.354.40HostedSelf-hostedHybridData protection, 25Quality, 25Operations, 15Cost at volume, 15Exit risk, 10WeightedLatency scored 4 for all three and is left out.
Hybrid wins by 0.05, so the argument that matters is about the weights, and the matrix puts that argument on the table.

Three topologies

Hosted API

  • Provider runs the model; you call it over a private connection
  • Strongest models available, no GPUs to run
  • Pay per token; costs rise with use
  • Data protection depends on contract and configuration

Self-hosted open weights

  • You run an open-weights model in your own cloud or data centre
  • Full control of data, versions and retirement
  • Pay for GPUs whether busy or idle, plus the team to run them
  • Quality depends on which open models fit your hardware

Hybrid

  • Each task goes where it fits best
  • Often a hosted model for hard generation and a small self-hosted model for frequent simple tasks
  • Two sets of operations to run
  • More design work, but no single dependency

"Hosted" at a bank rarely means calling a public endpoint over the internet. It usually means a model offered through the bank's existing cloud provider, in the bank's region, reached over private networking, under an enterprise agreement. That detail changes the security conversation completely.

What each option costs at Meridian's traffic

Start with the traffic, because the answer depends on it. Meridian's forecast is 3,000 policy answers, 600 summaries and 70 letters per working day. Section 4's token budgets turn that into about 23 million input tokens and 1.5 million output tokens a day.

At illustrative mid-tier hosted prices of $3 per million input tokens and $15 per million output tokens, that is about $70 plus $22, so roughly $100 a day, or about $2,200 a month. Your contract prices will differ; the method will not.

Now self-hosting. On the spike's golden set, the best open-weights model Meridian could serve on a single 8-GPU node scored 78%, against 83% for the hosted candidate on the same pipeline. To get close to the hosted quality, the bank would need a large model, and for resilience it would need at least two GPU nodes in separate zones.

ItemHosted APISelf-hosted, two GPU nodes
ComputeAbout $2,200 a month in tokens$30,000 to $60,000 a month in GPU rental
PeopleShare of the platform team1.5 to 2 extra engineers to run serving, scaling and patching
Quality on Meridian's set83% at spike78% at spike
Cost if traffic doublesAbout $4,400Same hardware, until it is full

At Meridian's volume, the API is more than ten times cheaper before counting people. Self-hosting wins on cost only at much higher, steadier volume, or when hardware is already owned and idle. Numbers like these change every year, so re-run the comparison at every annual review, but do not skip it.

The criteria a bank cares about

Cost is one criterion. A decision matrix makes the others visible and, more importantly, makes disagreements about their weight visible. Scores run from 1 (poor) to 5 (strong).

CriterionWeightHostedSelf-hostedHybrid
Data protection and residency25455
Quality on Meridian's evaluation25535
Operational burden15524
Cost at Meridian's volume15514
Exit and concentration risk10253
Latency10444
Weighted score, out of 54.353.354.40

Hybrid wins narrowly. The matrix is not a calculator of truth; if security had insisted on a weight of 50 for data protection, the result would move. That is the point: the argument is about the weights, and the matrix puts that argument on the table where it belongs, instead of hiding it inside adjectives.

Making a hosted model acceptable for customer data

Security's concern was real, and the answer was a set of specific controls, each one checkable.

  • In-region processing. The model endpoint runs in the bank's region; data does not leave it.
  • Private connectivity. Calls travel over the cloud provider's private network, not the public internet.
  • No retention, no training. The contract states that prompts and outputs are not stored beyond processing and never used for training. Some providers keep data for abuse monitoring by default, so the exemption must be agreed and written down, not assumed.
  • Encryption and keys. Data is encrypted in transit and at rest under the bank's standard.
  • Sub-processors and audit. The provider's list of sub-processors and its audit reports are reviewed by procurement.
  • Exit plan. The bank can move to another provider within a defined time.

The exit plan matters to regulators as well as to the bank. In the EU, the Digital Operational Resilience Act makes managing ICT third-party risk, including exit strategies, an explicit obligation for financial firms. Meridian's exit plan rests on three things: every model call goes through the bank's own gateway with a provider-neutral interface, a second provider is kept evaluated, and the evaluation suite itself is the migration tool.

The resulting decision is part A of MER-05.

YAML
id: MER-05-Adecision: hybridcomponents:  generation:           # answers, summaries, letter wording    where: hosted, provider A, in-region private endpoint    data_terms: no retention, no training (contract clause 7, signed)  guard_and_rewrite:    where: self-hosted, bank cloud tenancy, 1 GPU    model: open weights, about 8B parameters  embeddings_and_rerank:    where: hosted, provider A, same regiongateway: bank AI gateway, provider-neutral interface, all calls loggedsecond_provider: provider B, evaluated quarterly, contract in placeexit_time_target: 12 weeks to move generation to provider Breview: annually, or if daily volume exceeds 10x forecast

Check your understanding

0 of 3 answered

1.At Meridian's traffic, why does the hosted API cost far less than self-hosting?

2.Security says customer data cannot go to a hosted model. Which response best reflects the lesson?

3.Meridian's decision matrix gave hybrid 4.40 and hosted 4.35. What is the right way to use that result?