Course Content
LLMOps & Deployment
6 sections · 40 lessons
What factors influence the choice between cloud and on-premise deployment?
What you need to know
What each side wins on
Managed API / cloud
- Start in days; no hardware
- Elastic for spikes and uncertain demand
- Frontier models you cannot self-host
- No GPU on-call; provider handles upgrades
Self-hosted / on-premise
- Data never leaves your network
- Lower cost per token at high, steady load
- Custom and fine-tuned open models
- Predictable latency; no shared quotas
The break-even calculation
Self-hosting is a fixed monthly cost (GPUs, including redundancy and idle time, plus people). An API is a variable cost (tokens × price). Break-even is where they meet.
1HOURS = 730 # hours in a month23def self_host_monthly(gpus, gpu_hourly, ops_monthly):4 return gpus * gpu_hourly * HOURS + ops_monthly56def api_monthly(tokens_m_in, tokens_m_out, price_in, price_out):7 return tokens_m_in * price_in + tokens_m_out * price_out89# Illustrative: 2 GPUs (one is redundancy) at $2.50/h, plus a share of an engineer10fixed = self_host_monthly(gpus=2, gpu_hourly=2.50, ops_monthly=4_000)1112# Workload: 3,000M input and 400M output tokens a month13for name, p_in, p_out in [("frontier API", 3.00, 15.00),14 ("small hosted model", 0.20, 0.80)]:15 api = api_monthly(3_000, 400, p_in, p_out)16 print(f"{name:20s} API ${api:>9,.0f} self-host ${fixed:,.0f}")Output: frontier API $15,000 vs self-host $7,650; small hosted model $920 vs self-host $7,650. All prices are illustrative. The lesson: if a small open model is good enough, compare with a hosted small model, which is often far cheaper than running your own GPUs. Self-hosting wins clearly only when volume is high and steady, or when something other than price forces it.
Also count: GPU utilisation (a GPU at 20% busy costs 2.5× more per token than at 50%), reserved versus on-demand pricing, and the engineering time for upgrades, security patches and on-call.
Constraints that override the maths
- Data that legally cannot leave a country, a network, or the building.
- Air-gapped or classified environments.
- Contract terms on retention and training that a provider cannot meet.
Before assuming these force self-hosting, check what providers already offer: regional endpoints (including Indian regions), zero-data-retention terms, private networking, and healthcare agreements. These often remove the constraint.
A real-life example
Two teams in the same group company reach opposite answers.
The hospital must keep patient data inside its own network — a rule, not a preference. It runs open models on four on-prem GPUs even though its volume is small and an API would cost less. The design goal becomes "cheapest on-prem setup that meets quality", which is why it uses 4-bit quantization and shares GPUs between workloads.
The code assistant for 2,000 engineers sends about 16 billion input tokens a month. On a frontier API that was $63,000 a month before optimisation and about $35,000 after prompt caching. Self-hosting an open coding model on eight H100s would cost about $14,600 a month in GPUs (illustrative reserved price) plus about $6,000 of engineering time. But the open model scores 6 points lower on refactoring. The team goes hybrid: the self-hosted model takes explain, document and small-edit requests (70% of traffic), and the frontier API keeps refactoring. The monthly bill lands near $30,000 with no quality loss on the hard tasks, and proprietary code for most requests never leaves the company network.
Follow-up questions to expect
- "What is the biggest hidden cost of self-hosting?" — Idle capacity and people: redundancy GPUs, headroom for peaks, and engineers for upgrades and on-call.
- "When does cloud GPU rental beat on-prem hardware?" — When demand is uncertain or spiky, or for the newest GPUs; buying hardware wins only with years of steady, high utilisation.
- "How do you avoid getting stuck either way?" — Keep a gateway and your own evals, so moving a route between API and self-hosted is a config change.