LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What factors influence the choice between cloud and on-premise deployment?


What you need to know

What each side wins on

Managed API / cloud

  • Start in days; no hardware
  • Elastic for spikes and uncertain demand
  • Frontier models you cannot self-host
  • No GPU on-call; provider handles upgrades

Self-hosted / on-premise

  • Data never leaves your network
  • Lower cost per token at high, steady load
  • Custom and fine-tuned open models
  • Predictable latency; no shared quotas

The break-even calculation

Self-hosting is a fixed monthly cost (GPUs, including redundancy and idle time, plus people). An API is a variable cost (tokens × price). Break-even is where they meet.

Python
HOURS = 730                                    # hours in a monthdef self_host_monthly(gpus, gpu_hourly, ops_monthly):    return gpus * gpu_hourly * HOURS + ops_monthlydef api_monthly(tokens_m_in, tokens_m_out, price_in, price_out):    return tokens_m_in * price_in + tokens_m_out * price_out# Illustrative: 2 GPUs (one is redundancy) at $2.50/h, plus a share of an engineerfixed = self_host_monthly(gpus=2, gpu_hourly=2.50, ops_monthly=4_000)# Workload: 3,000M input and 400M output tokens a monthfor name, p_in, p_out in [("frontier API", 3.00, 15.00),                          ("small hosted model", 0.20, 0.80)]:    api = api_monthly(3_000, 400, p_in, p_out)    print(f"{name:20s} API ${api:>9,.0f}   self-host ${fixed:,.0f}")

Output: frontier API $15,000 vs self-host $7,650; small hosted model $920 vs self-host $7,650. All prices are illustrative. The lesson: if a small open model is good enough, compare with a hosted small model, which is often far cheaper than running your own GPUs. Self-hosting wins clearly only when volume is high and steady, or when something other than price forces it.

Also count: GPU utilisation (a GPU at 20% busy costs 2.5× more per token than at 50%), reserved versus on-demand pricing, and the engineering time for upgrades, security patches and on-call.

Constraints that override the maths

  • Data that legally cannot leave a country, a network, or the building.
  • Air-gapped or classified environments.
  • Contract terms on retention and training that a provider cannot meet.

Before assuming these force self-hosting, check what providers already offer: regional endpoints (including Indian regions), zero-data-retention terms, private networking, and healthcare agreements. These often remove the constraint.

A real-life example

Two teams in the same group company reach opposite answers.

The hospital must keep patient data inside its own network — a rule, not a preference. It runs open models on four on-prem GPUs even though its volume is small and an API would cost less. The design goal becomes "cheapest on-prem setup that meets quality", which is why it uses 4-bit quantization and shares GPUs between workloads.

The code assistant for 2,000 engineers sends about 16 billion input tokens a month. On a frontier API that was $63,000 a month before optimisation and about $35,000 after prompt caching. Self-hosting an open coding model on eight H100s would cost about $14,600 a month in GPUs (illustrative reserved price) plus about $6,000 of engineering time. But the open model scores 6 points lower on refactoring. The team goes hybrid: the self-hosted model takes explain, document and small-edit requests (70% of traffic), and the frontier API keeps refactoring. The monthly bill lands near $30,000 with no quality loss on the hard tasks, and proprietary code for most requests never leaves the company network.

Follow-up questions to expect

  • "What is the biggest hidden cost of self-hosting?" — Idle capacity and people: redundancy GPUs, headroom for peaks, and engineers for upgrades and on-call.
  • "When does cloud GPU rental beat on-prem hardware?" — When demand is uncertain or spiky, or for the newest GPUs; buying hardware wins only with years of steady, high utilisation.
  • "How do you avoid getting stuck either way?" — Keep a gateway and your own evals, so moving a route between API and self-hosted is a config change.