Token Economics and Cost Optimization

Mini Project: Build an LLM Cost Optimisation Dashboard


Finance forwards last month's invoice for the support assistant: 11,379.60 dollars across three million requests. They ask two questions. Where did it go, and what happens when traffic triples?

The team already has a dashboard. It shows one number — mean cost per request, 0.0038 dollars — sitting flat and healthy on a green line. That number is true and it is useless, because 70 percent of requests cost 0.0010 dollars, about a quarter of the mean, and 3 percent of requests cost sixty times it. The average describes no request that the system has ever served.

The same dashboard reports a cache hit rate of 43.4 percent, which everyone treats as a triumph. That cache is avoiding 85 dollars a month of a 5,600-dollar bill, and the semantic-similarity version somebody has proposed would cost 585 dollars a month to run. It would be the most expensive optimisation in the system.

Both failures come from the same place: a dashboard that aggregates before it attributes. This project builds one that does not. You will instrument a workload down to the individual model call, aggregate it without lying, display the handful of metrics that change decisions, and produce a before-and-after report you could defend in the meeting where the budget is set.

Turning an invoice into an answerable questionEmit a usageevent per callCost function,per model rateRoll up byfeatureand tenantPanel thep95 cost,not the meanCompareagainstthe budget11,379.60 USD over three million requests averages 0.0038 USD, and the average is the misleading part.
Per-request attribution is what turns "the bill went up" into "this feature's retrieval step doubled its context".

What you are building, and how you know it works

#CriterionTarget
1Attribution granularityEvery model call carries request id, conversation id, route, model, prompt/cached/completion tokens and cost
2Attribution completenessAuxiliary calls (routing, embedding, reranking, safety) and retries roll up to their parent request; failed calls are recorded, not dropped
3Pricing fidelityCost function handles at least three token classes: fresh input, cache-read input, output
4Distribution reportingDashboard shows p50, p90, p99 and mean — never the mean alone
5Cache honestyReports cost-weighted hit rate and net cache saving after lookup cost
6Drill-downTop 50 most expensive requests are clickable through to their full call tree
7Before/afterBaseline and optimised pipelines run over the identical workload, savings attributed per technique
8Budget verdictProjection against a stated monthly budget, plus one sensitivity scenario
9No credentials requiredWhole harness runs against a mock client with no API key and no network

The workload and the constraints

Text
Product        AI customer-support assistantVolume         3,000,000 requests/month               faq       2,100,000  (70%)  short factual answers               support     810,000  (27%)  order and billing problems               analysis      90,000   (3%)  retrieval-heavy, long answersBudget         4,000 USD/monthQuality bar    faq       may use a small model and may be served from cache               support   must not be silently downgraded without an eval               analysis  always the strongest model availablePricing (USD per million tokens; illustrative figures for this exercise)               premium   3.00 input   0.30 cached input   15.00 output               small     0.25 input   0.025 cached input   1.25 output               embedding-small 0.02    embedding-large 0.65

Architecture

Text
  request in      |      +--> tracer.start(request_id, conversation_id, route)      |      v  [ router call ] [ embedding ] [ rerank ] [ main completion ] [ safety check ]      |               |             |              |                  |      +---------------+-------------+--------------+------------------+                                    |                                    v                            CostLedger.record(UsageEvent)   one row per CALL                                    |                                    v                    raw events table  (14 days, full fidelity)                                    |                        +-----------+-----------+                        v                       v              hourly roll-ups (90d)      daily roll-ups (forever)                        |                       |                        +-----------+-----------+                                    v                            dashboard queries                     percentiles / group-by route / top-N                     cost-weighted cache rate / burn-down

Per-request cost attribution

The unit problem, which comes first

One user message does not equal one model call. In this assistant a single message triggers a routing classification, an embedding of the query, a rerank of retrieved chunks, the main completion, and a safety check — five calls. A ledger that records one row per call and then reports "mean cost per request" is reporting mean cost per call, which in this system is 0.00076 dollars: exactly one fifth of the truth, and a number no invoice will ever agree with.

Record at the granularity of the call, report at the granularity of the thing a user asked for. Getting those two backwards is how a dashboard ends up disagreeing with the bill by a factor of five.

That requires a correlation id created at the entry point and threaded through every downstream call, including retries. A retry is not a new request; it is the same request costing twice, and a ledger that treats it as new will show a healthy per-request cost while the bill climbs. Failed calls belong in the ledger too — a completion that times out after generating 300 tokens is billed for those tokens.

The event schema and the cost function

Python
from dataclasses import dataclass, fieldPRICES = {  # USD per million tokens    "premium": {"input": 3.00, "cached_input": 0.30, "output": 15.00},    "small":   {"input": 0.25, "cached_input": 0.025, "output": 1.25},    "embed-s": {"input": 0.02, "cached_input": 0.02, "output": 0.00},}@dataclassclass UsageEvent:    request_id: str          # one user message    conversation_id: str     # a session; the unit finance cares about    turn: int    call_kind: str           # router | embed | rerank | completion | safety    route: str               # faq | support | analysis    model: str    prompt_tokens: int       # fresh input, billed at full rate    cached_prompt_tokens: int = 0    completion_tokens: int = 0    cache_hit: bool = False  # served entirely from the response cache    attempt: int = 1         # 2+ means this is a retry of the same work    ok: bool = True          # failures still cost money    latency_ms: int = 0def cost(e: UsageEvent) -> float:    if e.cache_hit:        return 0.0    p = PRICES[e.model]    return (e.prompt_tokens * p["input"]            + e.cached_prompt_tokens * p["cached_input"]            + e.completion_tokens * p["output"]) / 1_000_000

The cached_prompt_tokens field is the one most implementations omit, and omitting it makes the dashboard wrong in the direction that flatters you. Provider prompt caching typically bills re-read input at around a tenth of the normal rate and charges a premium to write the cache entry. A cost function with only input and output rates will price 8,000 cached tokens as 8,000 fresh ones and report a saving that never happened — or, if you subtract them entirely, report a saving that is ten times too large.

The mock client makes the whole harness runnable with no key. Deterministic output length keyed off a hash of the prompt means two runs of the same workload produce identical numbers, which is what makes a before/after comparison meaningful rather than a sampling artefact.

Python
class MockClient:    """Same surface as a real SDK: .chat() returns usage counts."""    def __init__(self, model): self.model = model    def chat(self, system, user, max_tokens=300, cached_prefix_tokens=0):        prompt = system + "\n" + user        n_in = count_tokens(prompt)        rng = random.Random(int(hashlib.sha256(prompt.encode()).hexdigest(), 16) % 2**32)        n_out = min(rng.randint(40, 2200), max_tokens)        return Usage(model=self.model,                     prompt_tokens=n_in - cached_prefix_tokens,                     cached_prompt_tokens=cached_prefix_tokens,                     completion_tokens=n_out)

Aggregation that does not lie

Three rules, each of which exists because breaking it produces a plausible wrong number.

Keep raw events; aggregate at query time. The temptation is to push cost into a metrics system as labelled counters. Do not label them per user: 400,000 users becomes 400,000 time series, and the cardinality will take the metrics backend down long before it answers a question. Raw rows in a columnar table, rolled up hourly after 14 days and daily after 90, cost almost nothing and answer questions you had not thought of yet.

Percentiles do not average. The mean of 24 hourly p99 values is not the daily p99, and it is usually well below it. If you pre-aggregate, store histograms or t-digests, not summary statistics.

Roll up calls to requests before computing any per-request statistic. Otherwise every percentile you report describes calls.

Python
def request_costs(events):    """Sum every call - including retries, failures and auxiliaries -    into the user-visible request that caused it."""    per_request = defaultdict(float)    for e in events:        per_request[e.request_id] += cost(e)    return per_requestdef summary(events):    v = sorted(request_costs(events).values())    q = lambda p: v[min(len(v) - 1, int(p * len(v)))]    return {"requests": len(v), "total": sum(v), "mean": sum(v) / len(v),            "p50": q(.50), "p90": q(.90), "p99": q(.99), "p999": q(.999),            "max": v[-1]}def cost_weighted_cache_rate(events):    """The only cache metric worth putting on a dashboard."""    avoided = sum(counterfactual_cost(e) for e in events if e.cache_hit)    paid    = sum(cost(e) for e in events) + sum(lookup_cost(e) for e in events)    return avoided / (avoided + paid) if (avoided + paid) else 0.0

The panels worth building

PanelQuestion it answersFailure it prevents
Cost per request: p50, p90, p99, meanIs the body or the tail driving spend?Optimising the 70 percent that is already cheap
Spend by route, stacked, month to dateWhere is the money?One flat number that hides everything
Top 50 most expensive requests, drillable to the call treeWhat does an expensive request actually look like?Percentiles with no examples behind them
Cost per conversationWhat does a user session cost?Per-request cost that ignores 9-turn sessions
Cost per resolved ticketUnit economicsCost per request falling while resolution rate falls faster
Cost-weighted cache hit rate and net savingIs the cache making money?Celebrating a hit rate (see below)
Input : output token ratio by routeWhich lever even applies here?Compressing prompts on an output-bound route
Cost per 1,000 requests, 30-day trendIs efficiency improving?Traffic growth masking a per-request regression
Budget burn-down with end-of-month projectionWill we make budget?Finding out from the invoice

Trap one: the average hides the tail

Here is the baseline workload — everything on the premium model, a 99-token system prompt on every call, max_tokens left at 500 — priced out per route.

RouteRequestsInput tokensOutput tokensCost/request (USD)Monthly (USD)Share of spend
faq2,100,000113450.0010142,129.4018.7%
support810,0002792100.0039873,229.4728.4%
analysis90,00011,9992,0600.0668976,020.7352.9%
Total3,000,000mean 0.00379311,379.60100%

Work one row: an analysis request sends 11,999 input tokens at 3.00 per million, which is 0.035997, and returns 2,060 output tokens at 15.00 per million, which is 0.030900 — 0.066897 per request, and 6,020.73 across 90,000 of them.

Now the distribution behind that mean:

StatisticCost (USD)Reading
p500.001014The typical request is an faq
p900.003987Still ordinary support traffic
p990.06216× the mean
p99.90.14839× the mean
mean0.0037933.7× the median; describes nothing

Three percent of requests are 52.9 percent of the bill. That single fact reorders the entire optimisation backlog. Shaving 20 percent off the faq route — the one everybody attacks first, because it is 70 percent of the traffic and feels like where the volume is — saves 426 dollars a month. Shaving 20 percent off analysis saves 1,204 dollars from three percent of the traffic.

Traffic volume tells you where the requests are. Only the cost distribution tells you where the money is, and in every LLM application worth optimising those are different places.

Two arithmetic checks worth internalising. A mean is not a description of a heavy-tailed distribution, it is a total divided by a count — useful for forecasting the bill, useless for deciding what to fix. And a per-request cost target ("get under 0.002 dollars") is unachievable in a system where the honest answer is that 97 percent of requests already are and the other 3 percent never will be.

The optimised pipeline, and where the saving really came from

Five changes, each independently measurable — which matters, because "we applied some optimisations" is not a report and "compression was worth 693 dollars, routing 1,507" is.

Python
SYSTEM = {"verbose": VERBOSE_99_TOKENS, "compressed": "You are a polite, accurate support assistant."}MAX_OUT  = {"faq": 60, "support": 180, "analysis": 900}ROUTE_TO = {"faq": "small", "support": "premium", "analysis": "premium"}TOP_K    = {"analysis": 10}          # was 20; recall unchanged on the eval setdef run(workload, optimised: bool):    ledger, cache = CostLedger(), ExactMatchCache()    for req in workload:        model  = ROUTE_TO[req.route] if optimised else "premium"        system = SYSTEM["compressed" if optimised else "verbose"]        ctx    = retrieve(req, k=TOP_K.get(req.route, 20)) if optimised else retrieve(req, k=20)        key    = cache.key(model, system, req.text)   # model MUST be in the key        if optimised and (hit := cache.get(key)):            ledger.record(req, hit, cache_hit=True)            continue        r = clients[model].chat(system, ctx + req.text,                                max_tokens=MAX_OUT[req.route] if optimised else 500)        cache.put(key, r)        ledger.record(req, r)    return ledger
RouteBefore (USD)After (USD)Saved (USD)Share of total saving
faq2,129.4054.592,074.8135.7%
support3,229.472,677.86551.619.5%
analysis6,020.732,840.943,179.7954.8%
Total11,379.605,573.395,806.2151.0%

Now decompose that by technique rather than by route, which is the table that actually changes what you work on next.

TechniqueSaving (USD/month)Share
Trim retrieved context on analysis, 11,900 → 6,000 tokens1,593.0027.4%
Cap and instruct output length on analysis, 2,060 → 9001,566.0027.0%
Route faq to the small model1,507.2826.0%
Compress the system prompt, 99 → 22 tokens693.0011.9%
Cap output on support, 210 → 180364.506.3%
Exact-match response cache at a 43.4% hit rate84.961.5%
Cache lookup cost−2.52−0.04%

Deleting 77 tokens from one string is worth 693 dollars a month because it is paid on all three million requests: 77 × 3.00 / 1,000,000 × 3,000,000. Halving the retrieved context on 90,000 requests is worth more than twice that. And the cache — the component teams reliably build first, and the one this dashboard was proudest of — is worth 1.5 percent.

Trap two: a cache hit rate that costs money

The optimised faq request costs 0.00006525 dollars. A cache hit avoids exactly that and nothing more, so 1,302,000 hits avoid 84.96 dollars. Against a total of 5,573.39, the cost-weighted hit rate is 1.5 percent, while the raw hit rate on the same dashboard reads 43.4 percent. Both numbers are correct. Only one of them is about money.

Worse, a cache is not free to consult. Every lookup costs something, and there is a threshold below which the cache is a net expense:

Text
break-even hit rate  =  cost per lookup / cost of the call it replaces
Cache designCost per lookup (USD)Replaces a call costingBreak-even hit rateVerdict
Exact match, in-process hash~00.00006525 (faq)~0%Keeps 85 a month
Semantic, small embedding, 60-token key0.00000120.000065251.8%Fine; small extra win
Semantic, large embedding, 300-token key0.0001950.00006525299%Impossible; loses 500 a month
Semantic, large embedding, on analysis0.0001950.0315660.62%A 12% hit rate saves about 323 a month after lookups

Row three is the proposal that was on the table at the start of this lesson. Three million lookups at 0.000195 is 585 dollars a month to avoid 85 — a net loss of 500 dollars, sold as a cost optimisation. No hit rate can rescue it, because the lookup costs three times the call it is trying to replace.

Cache where the money is, not where the repetition is. Repetition concentrates in cheap traffic; cost concentrates in expensive traffic; a cache placed by hit rate will reliably land on the wrong one.

Two correctness notes that belong in the same review. The cache key must include the model name and every parameter that changes the output, or a small-model answer will be served for a premium-model request — corrupting the response and mis-attributing the cost. And a cache has no idea when your returns policy changes; put a TTL on anything policy-shaped and an invalidation hook on anything you can detect changing.

The verdict against a budget

Savings percentages do not pass or fail. Budgets do. The optimised pipeline lands at 5,573.39 against a 4,000-dollar budget: 39 percent over, after a 51 percent saving. A report that stops at "we cut costs by half" has not answered the question that was asked.

One lever remains, and it is deliberately the last one: route support to the small model. That takes the tier from 2,677.86 to 223.16 and the total to 3,118.69, 22 percent under budget. It is also the only change in this project that can degrade what a customer receives, so it is gated on evidence rather than arithmetic — 300 labelled support tickets answered by both models, blind-scored, shipped only if the small model stays within an agreed margin. That evaluation costs about two dollars of tokens and decides the fate of 2,454 dollars a month.

Then state what would break the plan. Suppose analysis grows from 3 percent of traffic to 6 percent as one retrieval-heavy feature gets popular. Nothing about your cost engineering has changed, and the bill becomes 5,959.63 — 49 percent over budget again. That is what a spend distribution concentrated in 3 percent of requests means: the budget is one product decision away from breaking, and the dashboard should have an alert on route mix, not only on total spend.

Finally, the option everyone raises in this meeting. Self-hosting the faq tier on two small GPU instances, fully loaded with a share of engineering time, is roughly 1,800 dollars a month in fixed cost. The faq tier currently costs 54.59. Break-even is 1,800 / 0.00006525 = 27,586,207 requests a month, thirteen times the current faq volume. Self-hosting is not a marginal call here; it is 33 times more expensive, and the reason is that a fixed cost is charged in full at any traffic level while an optimised API tier is charged at nearly nothing.

What goes wrong, and what to do about it

SymptomCauseFix
Dashboard total is a fifth of the invoicePer-call rows reported as per-request metricsRoll calls up by request_id before any statistic
Per-request cost flat, invoice climbingRetries counted as new requestsCarry the correlation id into retries; track attempt
Prompt-caching savings look implausibly largeCached input priced at zero, or at the full rateAdd a cached_input rate; it is roughly a tenth, not nothing
Mean cost per request looks healthy, budget missedHeavy tail; the mean is a total divided by a countReport p50/p90/p99 and spend share by route
Optimised the biggest route, saved almost nothingBiggest by volume is not biggest by spendRank routes by spend share before choosing work
Cache hit rate 43 percent, bill unchangedHits land on the cheapest trafficReport cost-weighted hit rate; cache the expensive route
Adding a semantic cache increased spendLookup costs more than the call it replacesCompute break-even hit rate before building it
Cached answer quotes the wrong policyNo TTL, no invalidationTTL on policy-shaped content, hook on known changes
Premium-quality request answered by the small modelModel name missing from the cache keyKey on model plus every parameter affecting output
Metrics backend falls overPer-user labels; cardinality explosionRaw rows in a columnar table, aggregate at query time
Daily p99 lower than several hourly p99sPercentiles were averagedStore histograms or t-digests, or recompute from raw
Answers truncate mid-sentence after a cost pushmax_tokens treated as a quality dialAsk for brevity in the prompt; the cap is a backstop

What this means when you build

Instrument before you optimise, and instrument the tail specifically. The single highest-value panel in this whole project is the drillable list of the 50 most expensive requests, because it converts a percentile into an artefact you can read. Nobody looks at "p99 = 0.062" and knows what to do. Everybody looks at a request that pulled 20 chunks of context to answer a question about shipping and knows exactly what to do.

Attribute by cost, not by count, whenever you decide what to work on. Volume-ranked backlogs send teams to optimise the cheap majority; spend-ranked backlogs send them to the 3 percent that is half the bill. In this workload the two orderings are exact opposites.

Pair every saving with the quality evidence it depends on, and refuse to ship the ones that have none. Compression, output caps, context trimming and caching are all defensible from measurement alone. Downgrading a model is not, and it is the one lever whose failure your dashboard will never show you — cost falls, the graph turns green, and the damage appears weeks later in escalation rates that nobody connected to the change.

And keep the dashboard pointed at the decision. A cost dashboard exists to answer three questions: what will the bill be, which change would move it most, and are we still inside the budget. Any panel that does not serve one of those is decoration, and decoration is how a team ends up with a green dashboard and an invoice nobody can explain.