Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Hosted API, Self-Hosted or Hybrid
The first time the architect showed Meridian's security team a diagram with a hosted model on it, the reaction was immediate: "Customer data cannot leave the bank. We have to run the model ourselves." A week later, a finance partner looked at the per-token pricing and said the opposite for a different reason: "Why rent intelligence by the word when we could buy GPUs and own it?"
Both reactions are reasonable starting points. Both are also claims that need numbers. Where the model runs is one of the most expensive decisions in the design to reverse, and it touches security, cost, quality, resilience and regulation all at once.
This lesson compares the three topologies with Meridian's real traffic, shows how the bank made a hosted model acceptable for customer data, and records the result as part A of MER-05.
Three topologies
Hosted API
- Provider runs the model; you call it over a private connection
- Strongest models available, no GPUs to run
- Pay per token; costs rise with use
- Data protection depends on contract and configuration
Self-hosted open weights
- You run an open-weights model in your own cloud or data centre
- Full control of data, versions and retirement
- Pay for GPUs whether busy or idle, plus the team to run them
- Quality depends on which open models fit your hardware
Hybrid
- Each task goes where it fits best
- Often a hosted model for hard generation and a small self-hosted model for frequent simple tasks
- Two sets of operations to run
- More design work, but no single dependency
"Hosted" at a bank rarely means calling a public endpoint over the internet. It usually means a model offered through the bank's existing cloud provider, in the bank's region, reached over private networking, under an enterprise agreement. That detail changes the security conversation completely.
What each option costs at Meridian's traffic
Start with the traffic, because the answer depends on it. Meridian's forecast is 3,000 policy answers, 600 summaries and 70 letters per working day. Section 4's token budgets turn that into about 23 million input tokens and 1.5 million output tokens a day.
At illustrative mid-tier hosted prices of $3 per million input tokens and $15 per million output tokens, that is about $70 plus $22, so roughly $100 a day, or about $2,200 a month. Your contract prices will differ; the method will not.
Now self-hosting. On the spike's golden set, the best open-weights model Meridian could serve on a single 8-GPU node scored 78%, against 83% for the hosted candidate on the same pipeline. To get close to the hosted quality, the bank would need a large model, and for resilience it would need at least two GPU nodes in separate zones.
| Item | Hosted API | Self-hosted, two GPU nodes |
|---|---|---|
| Compute | About $2,200 a month in tokens | $30,000 to $60,000 a month in GPU rental |
| People | Share of the platform team | 1.5 to 2 extra engineers to run serving, scaling and patching |
| Quality on Meridian's set | 83% at spike | 78% at spike |
| Cost if traffic doubles | About $4,400 | Same hardware, until it is full |
At Meridian's volume, the API is more than ten times cheaper before counting people. Self-hosting wins on cost only at much higher, steadier volume, or when hardware is already owned and idle. Numbers like these change every year, so re-run the comparison at every annual review, but do not skip it.
The criteria a bank cares about
Cost is one criterion. A decision matrix makes the others visible and, more importantly, makes disagreements about their weight visible. Scores run from 1 (poor) to 5 (strong).
| Criterion | Weight | Hosted | Self-hosted | Hybrid |
|---|---|---|---|---|
| Data protection and residency | 25 | 4 | 5 | 5 |
| Quality on Meridian's evaluation | 25 | 5 | 3 | 5 |
| Operational burden | 15 | 5 | 2 | 4 |
| Cost at Meridian's volume | 15 | 5 | 1 | 4 |
| Exit and concentration risk | 10 | 2 | 5 | 3 |
| Latency | 10 | 4 | 4 | 4 |
| Weighted score, out of 5 | 4.35 | 3.35 | 4.40 |
Hybrid wins narrowly. The matrix is not a calculator of truth; if security had insisted on a weight of 50 for data protection, the result would move. That is the point: the argument is about the weights, and the matrix puts that argument on the table where it belongs, instead of hiding it inside adjectives.
Making a hosted model acceptable for customer data
Security's concern was real, and the answer was a set of specific controls, each one checkable.
- In-region processing. The model endpoint runs in the bank's region; data does not leave it.
- Private connectivity. Calls travel over the cloud provider's private network, not the public internet.
- No retention, no training. The contract states that prompts and outputs are not stored beyond processing and never used for training. Some providers keep data for abuse monitoring by default, so the exemption must be agreed and written down, not assumed.
- Encryption and keys. Data is encrypted in transit and at rest under the bank's standard.
- Sub-processors and audit. The provider's list of sub-processors and its audit reports are reviewed by procurement.
- Exit plan. The bank can move to another provider within a defined time.
The exit plan matters to regulators as well as to the bank. In the EU, the Digital Operational Resilience Act makes managing ICT third-party risk, including exit strategies, an explicit obligation for financial firms. Meridian's exit plan rests on three things: every model call goes through the bank's own gateway with a provider-neutral interface, a second provider is kept evaluated, and the evaluation suite itself is the migration tool.
The resulting decision is part A of MER-05.
1id: MER-05-A2decision: hybrid3components:4 generation: # answers, summaries, letter wording5 where: hosted, provider A, in-region private endpoint6 data_terms: no retention, no training (contract clause 7, signed)7 guard_and_rewrite:8 where: self-hosted, bank cloud tenancy, 1 GPU9 model: open weights, about 8B parameters10 embeddings_and_rerank:11 where: hosted, provider A, same region12gateway: bank AI gateway, provider-neutral interface, all calls logged13second_provider: provider B, evaluated quarterly, contract in place14exit_time_target: 12 weeks to move generation to provider B15review: annually, or if daily volume exceeds 10x forecastCheck your understanding
0 of 3 answered
1.At Meridian's traffic, why does the hosted API cost far less than self-hosting?
2.Security says customer data cannot go to a hosted model. Which response best reflects the lesson?
3.Meridian's decision matrix gave hybrid 4.40 and hosted 4.35. What is the right way to use that result?