Token Economics and Cost Optimization

Local Models and Hybrid Architectures — Changing the Model Instead of the Prompt


The spreadsheet was compelling. A GPU instance costs 1.20 dollars an hour. That is 864 dollars a month. The API bill was 6,400 dollars a month. Self-hosting would therefore save about 5,500 dollars a month, and the decision took eleven minutes.

Six weeks later the real monthly cost was 4,928 dollars, not 864. Two instances were needed rather than one, because a single instance means an outage every time it restarts. An engineer was spending a day a week on the serving stack. And the model was handling only the simple 62 percent of traffic, so the API bill had not gone to zero — it had gone to 3,780. Total spend had not fallen at all: it had risen, from 6,400 to 8,708 dollars a month, and the team had acquired a GPU cluster to look after.

Nothing in that spreadsheet was a lie. It was just measuring the wrong thing. Hosted APIs and self-hosted models have fundamentally different cost shapes, and the decision between them is not a price comparison — it is a break-even calculation with a volume on one side and a fixed monthly commitment on the other.

The crossover the eleven-minute spreadsheet missed1,600 USD5,630 USDAPI6,400 USD5,630 USDSelf-host,barely25,600 USD8,220 USDSelf-host64,000 USD13,400 USDSelf-hostAPI billSelf-host, honestCheaper0.5 M req/month2 M req/month8 M req/month20 M req/monthThe honest fixed cost is not one GPU at 864 USD — it is redundancy, monitoring and a name on the pager.
Self-hosting is a fixed cost with a capacity ceiling; the API is a slope, and only volume decides which wins.

Two cost shapes, not two prices

Hosted APISelf-hosted model
Cost structurePurely variable, per tokenAlmost entirely fixed, per hour
Cost at zero trafficZeroFull
Cost of a 10× traffic spike10× the billZero, until you exceed capacity — then a hard failure
Marginal cost of one more requestThe token priceEffectively zero
Who absorbs idle timeThe providerYou
Time to first requestMinutesDays to weeks
Ongoing engineeringNoneContinuous
Capability ceilingFrontierWhatever fits your hardware

A hosted API charges you for what you use. A GPU charges you for what you reserved. Idle capacity is the entire hidden cost of self-hosting, and it is invisible in every price-per-hour comparison.

The honest fixed cost

Count everything, not just the instance.

ElementMonthly (USD)Why it is not optional
2 GPU instances at 1.20/hour × 720 hours1,728One instance means downtime on every deploy or crash
Storage and egress120Model weights, logs, metrics
Monitoring and alerting80You now own an inference service
Engineering, 0.25 FTE fully loaded3,000Upgrades, OOM incidents, throughput tuning, weight updates
Total fixed cost4,928

The engineering line is the one teams leave out, and it is 61 percent of the total. It does not go away after launch: serving frameworks change, a new model release makes your current one obsolete, and someone has to be on call when the GPU node runs out of memory at 3am.

Computing the break-even volume

Define a standard request for your workload — say 800 input tokens and 200 output tokens. On a small hosted model at 1.00 input and 5.00 output per million:

Text
API cost per request = 800 x 1.00/1e6 + 200 x 5.00/1e6                     = 0.000800 + 0.001000                     = 0.001800 USDbreak-even volume    = fixed monthly cost / API cost per request                     = 4,928 / 0.0018                     = 2,737,778 requests per month

That single number replaces the eleven-minute spreadsheet. Below roughly 2.7 million requests a month the API is cheaper; above it the GPU is. Now look at how sharply it swings either side of that line.

Requests / monthAPI cost (USD)Self-hosted (USD)Cheaper option
100,0001804,928API, by 27×
500,0009004,928API, by 5.5×
1,000,0001,8004,928API, by 2.7×
2,737,7784,9284,928Break-even
5,000,0009,0004,928Local, by 1.8×
20,000,00036,0004,928Local, by 7.3×
50,000,00090,0006,656 (4 instances)Local, by 13.5×

Two things fall out of this table that most discussions of the topic miss. At low volume, self-hosting is not marginally worse — it is 27 times worse, and no amount of tuning closes that gap. At high volume, self-hosting is not marginally better; it is transformational, because the marginal cost of an extra request on hardware you already own is essentially zero. The mistake is treating this as a close call. It almost never is; the only hard part is knowing which side of the line you are on.

The capacity check nobody runs

A break-even volume is meaningless if the hardware cannot serve it. Suppose your serving stack sustains 8 requests per second per instance for this request shape. Two instances give you 8 × 2 × 2,592,000 seconds in a 30-day month, or about 41.5 million requests of capacity. That comfortably covers 20 million requests. It does not cover 50 million, which is why that row needs four instances.

Measure throughput under continuous batching, where the server packs many concurrent requests into each forward pass. Benchmarking with one request at a time understates real throughput by an order of magnitude and will make you buy four times the hardware you need.

Quantisation: fitting a large model on affordable hardware

Model weights are numbers. Storing each one in 16 bits is the training-time default, but inference tolerates far less precision than training does. Quantisation stores weights in fewer bits — 8, 4, sometimes fewer — which shrinks the memory footprint proportionally.

PrecisionBytes per parameter70B model8B modelTypical benchmark lossHardware needed for 70B
FP16 / BF162140 GB16 GBBaseline2 × 80 GB
INT8170 GB8 GBAround 1%1 × 80 GB
INT4 (GPTQ, AWQ)0.535 GB4 GB2–4%1 × 48 GB
INT30.37526 GB3 GB5–10%1 × 32 GB
INT20.2517.5 GB2 GBSevereRarely usable

The economic effect is direct. Two 80 GB cards at 1.20 dollars an hour each is 2.40 an hour; one 48 GB card is around 0.80. Quantising from FP16 to INT4 cuts the hardware bill by a factor of three, in exchange for two to four points of benchmark accuracy. Whether that is a good trade depends entirely on your task — a 2-point drop on a reasoning benchmark may be a 0-point drop on entity extraction.

Quantisation also makes inference faster, which surprises people. Generation is bound by memory bandwidth, not arithmetic: every token requires reading the whole weight matrix. Halve the bytes and you roughly halve the read time.

The memory nobody budgets for: the KV cache

Weights are not the only thing in VRAM. Every token in every in-flight request keeps a key and value vector per layer, and that KV cache grows with concurrency.

Text
per-token KV bytes = 2 (key and value)                   x layers                   x kv_heads                   x head_dim                   x bytes_per_valuefor a typical 8B model at FP16:  2 x 32 x 8 x 128 x 2 = 131,072 bytes = 128 KiB per token32 concurrent requests of 8,000 tokens each:  32 x 8,000 x 128 KiB = 32,768,000 KiB = 31.25 GiB

A 4 GB quantised model needing over 30 GiB of KV cache is the normal situation, not an edge case. This is why concurrency, not model size, usually decides your GPU. It is also why long prompts are expensive on self-hosted infrastructure in a way that has nothing to do with token pricing — every extra 1,000 tokens of context costs about 125 MiB of VRAM per concurrent request, and when you run out you do not get a slow response, you get an out-of-memory crash.

Choosing a local model

ParametersVRAM at INT4Reliably good atReliably bad at
1–3B1–2 GBClassification, routing, PII redaction, language detectionAnything multi-step
7–8B4–5 GBSummarisation, structured extraction, rewriting, simple Q&ALong-horizon reasoning, non-trivial code
13–14B7–8 GBThe above, plus decent code and better instruction followingFrontier-level reasoning
30–34B17–19 GBStrong general-purpose workStill a visible gap on hard tasks
70B35–40 GBComparable to mid-tier hosted models on many tasksThe hardest reasoning and agentic work

Do not choose from this table. Choose from your own evaluation set. Published benchmark scores measure performance on academic tasks that are almost certainly not your task, and the gap between "scores well on a leaderboard" and "handles our support tickets" is where most local-model projects die. Build 200 labelled examples of your real workload first; it is a day of work and it makes the entire decision empirical.

Beyond accuracy, four things determine whether a model is deployable at all: the licence (some forbid commercial use or impose revenue thresholds), the context window (check the length you can actually serve — the KV cache above, not the model card, usually sets the practical limit), tool-calling support (if your application needs function calls, a model without reliable structured output is unusable), and measured throughput at your batch size.

Hybrid architectures: route, do not choose

The framing "cloud or local" is a false binary. Real workloads are a mixture of easy and hard requests, and the winning architecture sends each one to the cheapest thing that can handle it.

Complexity-based routing

Classify the request first, then dispatch. The classifier itself should be cheap — a heuristic, or a 1B local model, never a frontier call.

Python
from dataclasses import dataclass@dataclass(frozen=True)class Route:    backend: str          # "local" | "hosted_small" | "hosted_frontier"    model: strROUTES = {    "classify":  Route("local",           "llama-3.1-8b-instruct-awq"),    "extract":   Route("local",           "llama-3.1-8b-instruct-awq"),    "summarise": Route("hosted_small",    "claude-haiku-4-5"),    "chat":      Route("hosted_small",    "claude-haiku-4-5"),    "analyse":   Route("hosted_frontier", "claude-opus-5-5"),    "code":      Route("hosted_frontier", "claude-opus-5-5"),}FALLBACK = Route("hosted_frontier", "claude-opus-5-5")def route_for(task: str) -> Route:    r = ROUTES.get(task)    if r is None:        # Loud, not silent — an unrouted task is a cost incident.        log.error("unrouted task %r, falling back to frontier", task)        metrics.increment("router.fallback", tags={"task": task})    return r or FALLBACK

The bug that quietly deletes your savings

Model identifiers contain hyphens and dots: llama-3.1-8b-instruct-awq. Python identifiers cannot. Every so often someone writes a router that treats model names as attributes, and it fails in the worst possible way — silently.

Python
# WRONG — will not even parseclass Model(Enum):    llama-3.1-8b = "llama-3.1-8b-instruct"      # SyntaxError# WRONG — parses, runs, and is always wrongdef pick(config, model_name):    # model_name is "llama-3.1-8b", which is never a valid attribute name,    # so getattr always misses and every request takes the fallback.    return getattr(config, model_name, config.default_model)# RIGHT — identifiers for code, strings for dataclass Model(str, Enum):    LLAMA_31_8B  = "llama-3.1-8b-instruct-awq"    HAIKU_45     = "claude-haiku-4-5"    OPUS_55      = "claude-opus-5-5"def pick(config: dict, model_name: str) -> str:    if model_name not in config:        raise KeyError(f"no backend configured for model {model_name!r}")    return config[model_name]

The middle version is the expensive one. getattr with a default never raises, so every request falls through to default_model, which is the frontier model, and your carefully designed hybrid router becomes a very slow way of sending everything to the most expensive backend. On the workload below, that mistake turns an 8,708-dollar month into a 27,000-dollar one, and nothing in the logs says so.

Any routing fallback that is silent is a cost incident waiting to happen. Make the fallback path emit an error-level log and increment a counter, and alert when that counter is non-zero.

Confidence-based escalation

The other pattern is to try the cheap model first and escalate only when the answer fails a check — a schema validation, a self-reported confidence score, a retrieval-grounding test.

The economics have a clean threshold. If the cheap path costs c and the expensive path costs C, and a fraction e of requests escalate, the average cost is c plus e times C. Escalation saves money whenever e is below 1 minus c divided by C:

Text
c = 0.0018   (hosted small)C = 0.0090   (frontier)break-even escalation rate = 1 - (0.0018 / 0.0090) = 0.80at e = 0.15:  0.0018 + 0.15 x 0.0090 = 0.00315 USD              versus 0.0090 always-frontier  ->  65% saving

An 80 percent break-even sounds like enormous headroom, and on cost it is. Latency is the real constraint: an escalated request pays both calls sequentially, so your p95 latency becomes the sum of the two. At a 15 percent escalation rate, roughly one request in seven is twice as slow. For a background job that is irrelevant; for a typeahead suggestion it is fatal.

A worked hybrid stack

Workload: 3,000,000 requests a month, 800 input and 200 output tokens each. Measured task mix: 62 percent simple extraction and classification, 30 percent moderate summarisation and chat, 8 percent genuinely hard analysis.

TierBackendRequestsUnit costMonthly cost (USD)
SimpleLocal 8B, INT41,860,000Fixed4,928
ModerateHosted small (1.00 / 5.00)900,0000.00181,620
HardHosted frontier (5.00 / 25.00)240,0000.00902,160
Hybrid total3,000,0008,708
Baseline: everything on frontier3,000,0000.009027,000

A saving of 18,292 dollars a month, or 67.7 percent. But a cost number alone is not a result, because routing makes mistakes — about 2 percent of hard requests get misclassified as simple and answered badly. Measured end-to-end accuracy is 94.1 percent for the all-frontier baseline and 93.4 percent for the hybrid. Divide to get the number that actually decides it:

Text
all frontier : 27,000 / (3,000,000 x 0.941) = 0.009564 USD per correct answerhybrid       :  8,708 / (3,000,000 x 0.934) = 0.003108 USD per correct answerhybrid is 3.08x better per correct answer

Seven tenths of an accuracy point bought a three-fold improvement in cost per correct answer. Whether you take that trade depends on what a wrong answer costs — trivial for a tag suggestion, unacceptable for a medication interaction check. State that cost explicitly before you route anything.

What people get wrong

BeliefRealityConsequence
"Local models are free"They are fixed-cost, and fixed cost is charged at zero trafficPaying 4,928 a month to serve 100,000 requests
"The instance price is the cost"Engineering is usually the largest line itemBreak-even volume understated by 3–5×
"One instance is enough"Deploys, crashes and upgrades all mean downtimeAn outage on the first weight update
"Quantisation is basically lossless"INT4 typically costs 2–4 benchmark pointsA quality regression discovered by customers
"A bigger local model closes the gap"Only your evaluation set can answer that40 GB of VRAM bought for no measured gain
"Benchmark with one request at a time"Continuous batching changes throughput by 10×Four times the hardware you needed
"Route on keywords"Brittle; misroutes exactly the unusual requests that need the strong modelBad answers concentrated on hard cases
"Escalation always saves money"Only below the break-even escalation rate, and it doubles tail latencyA slow product that costs the same
"Self-hosting solves data governance"It removes third-party transfer, not your retention, access-control or audit obligationsA compliance gap assumed to be closed

How to actually make this decision

Work in this order, and stop as soon as the answer is obvious.

Compute your break-even volume before anything else. Fixed monthly cost divided by API cost per request. If your traffic is under half that number, self-hosting is not a close call and you should stop here — the answer is the API, and the engineering time is better spent on prompt compression and caching, which cost nothing to run.

If you are near or above break-even, split the workload by difficulty before you buy hardware. Sample 500 real requests and label them simple, moderate or hard. The proportion of simple requests is what determines the size of the prize, and it is frequently much higher than teams expect — 62 percent in the example above, and it is not unusual to find 80 percent.

Validate the cheap tier on your own evaluation set, at the quantisation you intend to ship. Not the FP16 version, not the leaderboard score. The 4-bit quantised model, on your 200 labelled examples, with your prompt.

Ship the routing before the GPU. Route the simple tier to a small hosted model first. That captures most of the saving in an afternoon with no infrastructure, and it tells you whether the routing works before you commit to hardware. In the worked example, routing 62 percent of traffic to a hosted small model instead of a local one costs 1,860,000 times 0.0018 equals 3,348 dollars a month against the local option's 4,928 — the hosted version is cheaper at that volume, and needs no cluster at all.

That last point deserves emphasis, because it inverts the usual instinct. Self-hosting only wins on the tier that is both high-volume and simple enough for a small model. If you cannot fill the GPU, the API is not a compromise. It is the correct answer.