Course Content
Token Economics and Cost Optimization
3 sections · 7 lessons
Local Models and Hybrid Architectures — Changing the Model Instead of the Prompt
The spreadsheet was compelling. A GPU instance costs 1.20 dollars an hour. That is 864 dollars a month. The API bill was 6,400 dollars a month. Self-hosting would therefore save about 5,500 dollars a month, and the decision took eleven minutes.
Six weeks later the real monthly cost was 4,928 dollars, not 864. Two instances were needed rather than one, because a single instance means an outage every time it restarts. An engineer was spending a day a week on the serving stack. And the model was handling only the simple 62 percent of traffic, so the API bill had not gone to zero — it had gone to 3,780. Total spend had not fallen at all: it had risen, from 6,400 to 8,708 dollars a month, and the team had acquired a GPU cluster to look after.
Nothing in that spreadsheet was a lie. It was just measuring the wrong thing. Hosted APIs and self-hosted models have fundamentally different cost shapes, and the decision between them is not a price comparison — it is a break-even calculation with a volume on one side and a fixed monthly commitment on the other.
Two cost shapes, not two prices
| Hosted API | Self-hosted model | |
|---|---|---|
| Cost structure | Purely variable, per token | Almost entirely fixed, per hour |
| Cost at zero traffic | Zero | Full |
| Cost of a 10× traffic spike | 10× the bill | Zero, until you exceed capacity — then a hard failure |
| Marginal cost of one more request | The token price | Effectively zero |
| Who absorbs idle time | The provider | You |
| Time to first request | Minutes | Days to weeks |
| Ongoing engineering | None | Continuous |
| Capability ceiling | Frontier | Whatever fits your hardware |
A hosted API charges you for what you use. A GPU charges you for what you reserved. Idle capacity is the entire hidden cost of self-hosting, and it is invisible in every price-per-hour comparison.
The honest fixed cost
Count everything, not just the instance.
| Element | Monthly (USD) | Why it is not optional |
|---|---|---|
| 2 GPU instances at 1.20/hour × 720 hours | 1,728 | One instance means downtime on every deploy or crash |
| Storage and egress | 120 | Model weights, logs, metrics |
| Monitoring and alerting | 80 | You now own an inference service |
| Engineering, 0.25 FTE fully loaded | 3,000 | Upgrades, OOM incidents, throughput tuning, weight updates |
| Total fixed cost | 4,928 |
The engineering line is the one teams leave out, and it is 61 percent of the total. It does not go away after launch: serving frameworks change, a new model release makes your current one obsolete, and someone has to be on call when the GPU node runs out of memory at 3am.
Computing the break-even volume
Define a standard request for your workload — say 800 input tokens and 200 output tokens. On a small hosted model at 1.00 input and 5.00 output per million:
API cost per request = 800 x 1.00/1e6 + 200 x 5.00/1e6 = 0.000800 + 0.001000 = 0.001800 USDbreak-even volume = fixed monthly cost / API cost per request = 4,928 / 0.0018 = 2,737,778 requests per monthThat single number replaces the eleven-minute spreadsheet. Below roughly 2.7 million requests a month the API is cheaper; above it the GPU is. Now look at how sharply it swings either side of that line.
| Requests / month | API cost (USD) | Self-hosted (USD) | Cheaper option |
|---|---|---|---|
| 100,000 | 180 | 4,928 | API, by 27× |
| 500,000 | 900 | 4,928 | API, by 5.5× |
| 1,000,000 | 1,800 | 4,928 | API, by 2.7× |
| 2,737,778 | 4,928 | 4,928 | Break-even |
| 5,000,000 | 9,000 | 4,928 | Local, by 1.8× |
| 20,000,000 | 36,000 | 4,928 | Local, by 7.3× |
| 50,000,000 | 90,000 | 6,656 (4 instances) | Local, by 13.5× |
Two things fall out of this table that most discussions of the topic miss. At low volume, self-hosting is not marginally worse — it is 27 times worse, and no amount of tuning closes that gap. At high volume, self-hosting is not marginally better; it is transformational, because the marginal cost of an extra request on hardware you already own is essentially zero. The mistake is treating this as a close call. It almost never is; the only hard part is knowing which side of the line you are on.
The capacity check nobody runs
A break-even volume is meaningless if the hardware cannot serve it. Suppose your serving stack sustains 8 requests per second per instance for this request shape. Two instances give you 8 × 2 × 2,592,000 seconds in a 30-day month, or about 41.5 million requests of capacity. That comfortably covers 20 million requests. It does not cover 50 million, which is why that row needs four instances.
Measure throughput under continuous batching, where the server packs many concurrent requests into each forward pass. Benchmarking with one request at a time understates real throughput by an order of magnitude and will make you buy four times the hardware you need.
Quantisation: fitting a large model on affordable hardware
Model weights are numbers. Storing each one in 16 bits is the training-time default, but inference tolerates far less precision than training does. Quantisation stores weights in fewer bits — 8, 4, sometimes fewer — which shrinks the memory footprint proportionally.
| Precision | Bytes per parameter | 70B model | 8B model | Typical benchmark loss | Hardware needed for 70B |
|---|---|---|---|---|---|
| FP16 / BF16 | 2 | 140 GB | 16 GB | Baseline | 2 × 80 GB |
| INT8 | 1 | 70 GB | 8 GB | Around 1% | 1 × 80 GB |
| INT4 (GPTQ, AWQ) | 0.5 | 35 GB | 4 GB | 2–4% | 1 × 48 GB |
| INT3 | 0.375 | 26 GB | 3 GB | 5–10% | 1 × 32 GB |
| INT2 | 0.25 | 17.5 GB | 2 GB | Severe | Rarely usable |
The economic effect is direct. Two 80 GB cards at 1.20 dollars an hour each is 2.40 an hour; one 48 GB card is around 0.80. Quantising from FP16 to INT4 cuts the hardware bill by a factor of three, in exchange for two to four points of benchmark accuracy. Whether that is a good trade depends entirely on your task — a 2-point drop on a reasoning benchmark may be a 0-point drop on entity extraction.
Quantisation also makes inference faster, which surprises people. Generation is bound by memory bandwidth, not arithmetic: every token requires reading the whole weight matrix. Halve the bytes and you roughly halve the read time.
The memory nobody budgets for: the KV cache
Weights are not the only thing in VRAM. Every token in every in-flight request keeps a key and value vector per layer, and that KV cache grows with concurrency.
per-token KV bytes = 2 (key and value) x layers x kv_heads x head_dim x bytes_per_valuefor a typical 8B model at FP16: 2 x 32 x 8 x 128 x 2 = 131,072 bytes = 128 KiB per token32 concurrent requests of 8,000 tokens each: 32 x 8,000 x 128 KiB = 32,768,000 KiB = 31.25 GiBA 4 GB quantised model needing over 30 GiB of KV cache is the normal situation, not an edge case. This is why concurrency, not model size, usually decides your GPU. It is also why long prompts are expensive on self-hosted infrastructure in a way that has nothing to do with token pricing — every extra 1,000 tokens of context costs about 125 MiB of VRAM per concurrent request, and when you run out you do not get a slow response, you get an out-of-memory crash.
Choosing a local model
| Parameters | VRAM at INT4 | Reliably good at | Reliably bad at |
|---|---|---|---|
| 1–3B | 1–2 GB | Classification, routing, PII redaction, language detection | Anything multi-step |
| 7–8B | 4–5 GB | Summarisation, structured extraction, rewriting, simple Q&A | Long-horizon reasoning, non-trivial code |
| 13–14B | 7–8 GB | The above, plus decent code and better instruction following | Frontier-level reasoning |
| 30–34B | 17–19 GB | Strong general-purpose work | Still a visible gap on hard tasks |
| 70B | 35–40 GB | Comparable to mid-tier hosted models on many tasks | The hardest reasoning and agentic work |
Do not choose from this table. Choose from your own evaluation set. Published benchmark scores measure performance on academic tasks that are almost certainly not your task, and the gap between "scores well on a leaderboard" and "handles our support tickets" is where most local-model projects die. Build 200 labelled examples of your real workload first; it is a day of work and it makes the entire decision empirical.
Beyond accuracy, four things determine whether a model is deployable at all: the licence (some forbid commercial use or impose revenue thresholds), the context window (check the length you can actually serve — the KV cache above, not the model card, usually sets the practical limit), tool-calling support (if your application needs function calls, a model without reliable structured output is unusable), and measured throughput at your batch size.
Hybrid architectures: route, do not choose
The framing "cloud or local" is a false binary. Real workloads are a mixture of easy and hard requests, and the winning architecture sends each one to the cheapest thing that can handle it.
Complexity-based routing
Classify the request first, then dispatch. The classifier itself should be cheap — a heuristic, or a 1B local model, never a frontier call.
1from dataclasses import dataclass23@dataclass(frozen=True)4class Route:5 backend: str # "local" | "hosted_small" | "hosted_frontier"6 model: str78ROUTES = {9 "classify": Route("local", "llama-3.1-8b-instruct-awq"),10 "extract": Route("local", "llama-3.1-8b-instruct-awq"),11 "summarise": Route("hosted_small", "claude-haiku-4-5"),12 "chat": Route("hosted_small", "claude-haiku-4-5"),13 "analyse": Route("hosted_frontier", "claude-opus-5-5"),14 "code": Route("hosted_frontier", "claude-opus-5-5"),15}1617FALLBACK = Route("hosted_frontier", "claude-opus-5-5")1819def route_for(task: str) -> Route:20 r = ROUTES.get(task)21 if r is None:22 # Loud, not silent — an unrouted task is a cost incident.23 log.error("unrouted task %r, falling back to frontier", task)24 metrics.increment("router.fallback", tags={"task": task})25 return r or FALLBACKThe bug that quietly deletes your savings
Model identifiers contain hyphens and dots: llama-3.1-8b-instruct-awq. Python identifiers cannot. Every so often someone writes a router that treats model names as attributes, and it fails in the worst possible way — silently.
1# WRONG — will not even parse2class Model(Enum):3 llama-3.1-8b = "llama-3.1-8b-instruct" # SyntaxError45# WRONG — parses, runs, and is always wrong6def pick(config, model_name):7 # model_name is "llama-3.1-8b", which is never a valid attribute name,8 # so getattr always misses and every request takes the fallback.9 return getattr(config, model_name, config.default_model)1011# RIGHT — identifiers for code, strings for data12class Model(str, Enum):13 LLAMA_31_8B = "llama-3.1-8b-instruct-awq"14 HAIKU_45 = "claude-haiku-4-5"15 OPUS_55 = "claude-opus-5-5"1617def pick(config: dict, model_name: str) -> str:18 if model_name not in config:19 raise KeyError(f"no backend configured for model {model_name!r}")20 return config[model_name]The middle version is the expensive one. getattr with a default never raises, so every request falls through to default_model, which is the frontier model, and your carefully designed hybrid router becomes a very slow way of sending everything to the most expensive backend. On the workload below, that mistake turns an 8,708-dollar month into a 27,000-dollar one, and nothing in the logs says so.
Any routing fallback that is silent is a cost incident waiting to happen. Make the fallback path emit an error-level log and increment a counter, and alert when that counter is non-zero.
Confidence-based escalation
The other pattern is to try the cheap model first and escalate only when the answer fails a check — a schema validation, a self-reported confidence score, a retrieval-grounding test.
The economics have a clean threshold. If the cheap path costs c and the expensive path costs C, and a fraction e of requests escalate, the average cost is c plus e times C. Escalation saves money whenever e is below 1 minus c divided by C:
c = 0.0018 (hosted small)C = 0.0090 (frontier)break-even escalation rate = 1 - (0.0018 / 0.0090) = 0.80at e = 0.15: 0.0018 + 0.15 x 0.0090 = 0.00315 USD versus 0.0090 always-frontier -> 65% savingAn 80 percent break-even sounds like enormous headroom, and on cost it is. Latency is the real constraint: an escalated request pays both calls sequentially, so your p95 latency becomes the sum of the two. At a 15 percent escalation rate, roughly one request in seven is twice as slow. For a background job that is irrelevant; for a typeahead suggestion it is fatal.
A worked hybrid stack
Workload: 3,000,000 requests a month, 800 input and 200 output tokens each. Measured task mix: 62 percent simple extraction and classification, 30 percent moderate summarisation and chat, 8 percent genuinely hard analysis.
| Tier | Backend | Requests | Unit cost | Monthly cost (USD) |
|---|---|---|---|---|
| Simple | Local 8B, INT4 | 1,860,000 | Fixed | 4,928 |
| Moderate | Hosted small (1.00 / 5.00) | 900,000 | 0.0018 | 1,620 |
| Hard | Hosted frontier (5.00 / 25.00) | 240,000 | 0.0090 | 2,160 |
| Hybrid total | 3,000,000 | 8,708 | ||
| Baseline: everything on frontier | 3,000,000 | 0.0090 | 27,000 |
A saving of 18,292 dollars a month, or 67.7 percent. But a cost number alone is not a result, because routing makes mistakes — about 2 percent of hard requests get misclassified as simple and answered badly. Measured end-to-end accuracy is 94.1 percent for the all-frontier baseline and 93.4 percent for the hybrid. Divide to get the number that actually decides it:
all frontier : 27,000 / (3,000,000 x 0.941) = 0.009564 USD per correct answerhybrid : 8,708 / (3,000,000 x 0.934) = 0.003108 USD per correct answerhybrid is 3.08x better per correct answerSeven tenths of an accuracy point bought a three-fold improvement in cost per correct answer. Whether you take that trade depends on what a wrong answer costs — trivial for a tag suggestion, unacceptable for a medication interaction check. State that cost explicitly before you route anything.
What people get wrong
| Belief | Reality | Consequence |
|---|---|---|
| "Local models are free" | They are fixed-cost, and fixed cost is charged at zero traffic | Paying 4,928 a month to serve 100,000 requests |
| "The instance price is the cost" | Engineering is usually the largest line item | Break-even volume understated by 3–5× |
| "One instance is enough" | Deploys, crashes and upgrades all mean downtime | An outage on the first weight update |
| "Quantisation is basically lossless" | INT4 typically costs 2–4 benchmark points | A quality regression discovered by customers |
| "A bigger local model closes the gap" | Only your evaluation set can answer that | 40 GB of VRAM bought for no measured gain |
| "Benchmark with one request at a time" | Continuous batching changes throughput by 10× | Four times the hardware you needed |
| "Route on keywords" | Brittle; misroutes exactly the unusual requests that need the strong model | Bad answers concentrated on hard cases |
| "Escalation always saves money" | Only below the break-even escalation rate, and it doubles tail latency | A slow product that costs the same |
| "Self-hosting solves data governance" | It removes third-party transfer, not your retention, access-control or audit obligations | A compliance gap assumed to be closed |
How to actually make this decision
Work in this order, and stop as soon as the answer is obvious.
Compute your break-even volume before anything else. Fixed monthly cost divided by API cost per request. If your traffic is under half that number, self-hosting is not a close call and you should stop here — the answer is the API, and the engineering time is better spent on prompt compression and caching, which cost nothing to run.
If you are near or above break-even, split the workload by difficulty before you buy hardware. Sample 500 real requests and label them simple, moderate or hard. The proportion of simple requests is what determines the size of the prize, and it is frequently much higher than teams expect — 62 percent in the example above, and it is not unusual to find 80 percent.
Validate the cheap tier on your own evaluation set, at the quantisation you intend to ship. Not the FP16 version, not the leaderboard score. The 4-bit quantised model, on your 200 labelled examples, with your prompt.
Ship the routing before the GPU. Route the simple tier to a small hosted model first. That captures most of the saving in an afternoon with no infrastructure, and it tells you whether the routing works before you commit to hardware. In the worked example, routing 62 percent of traffic to a hosted small model instead of a local one costs 1,860,000 times 0.0018 equals 3,348 dollars a month against the local option's 4,928 — the hosted version is cheaper at that volume, and needs no cluster at all.
That last point deserves emphasis, because it inverts the usual instinct. Self-hosting only wins on the tier that is both high-volume and simple enough for a small model. If you cannot fill the GPU, the API is not a compromise. It is the correct answer.