Token Economics and Cost Optimization

Measuring Token Usage in Applications


A team built a cost dashboard. It was not a toy: it logged every call, multiplied tokens by the published rate, and rolled the result up by day. At the end of the month it read 1,740 dollars. The invoice read 2,398 dollars — 38 percent higher.

The dashboard was not buggy in the ordinary sense. Every line it recorded was correct. The problem was the lines it never recorded at all.

Their assistant handled 200,000 user messages a month, each averaging 1,800 input and 220 output tokens on a model priced at 3.00 and 15.00 USD per million. That is 0.0054 plus 0.0033, or 0.0087 dollars a message — 1,740 dollars across 200,000 messages, exactly as the dashboard said. But 30 percent of those messages triggered a tool call, and a tool call is not one API request, it is two: the first returns a request to run the tool, the second resends the whole conversation plus the assistant's tool-use block plus the tool result, and generates the real answer. That second call carried about 2,370 input and 180 output tokens — 0.00981 dollars — and there were 60,000 of them, for 588.60 dollars that the dashboard never saw. Another 4 percent of calls came back with malformed JSON and were silently re-asked by a retry wrapper: 8,000 extra calls at 0.0087, or 69.60 dollars. Add it up: 1,740 plus 588.60 plus 69.60 equals 2,398.20.

Nobody had lied and nothing had leaked. The measurement simply counted user messages when the meter counts API calls. This lesson is about closing that gap: reading what each provider actually reports, normalising it into one shape, and turning it into a number you can forecast and alert on.

Why the dashboard read 1,740 and the invoice 2,398What the tracker counted• One usage record per returned call• The published per-token list rate• Tokens the SDK handed back• Rolled up by calendar dayWhat the provider also billed• Retries after a client-side timeout• Streams that droppedbefore the last chunk• Cached and system tokens priced apart• Failed calls that still consumed input
The 38 percent gap was every call the code never saw return — you cannot log usage from a response you abandoned.

Estimating and measuring are two different jobs

People conflate these constantly, and then wonder why their pre-flight token estimate does not match the invoice. They are not the same activity and they do not have the same purpose.

EstimatingMeasuring
When it happensBefore the requestAfter the response
HowLocal tokenizer or a counting endpointRead the usage block on the response
What it knowsInput onlyInput, output, cache reads, cache writes, reasoning
AccuracyGood on input; blind to outputExact — it is the billing record
What it is forGuardrails: reject or trim before you payAccounting, forecasting, anomaly detection
Failure if you skip itA 400,000-token pasted log slips through and costs 1.20 USD of input at 3.00 per million — on every retryYou find out what you spent from the invoice

You need both. Estimation is a gate on the way in. Measurement is the ledger on the way out. Using an estimate as your ledger is the specific mistake that makes dashboards drift from invoices, because an estimate cannot know how many tokens the model chose to generate.

An estimate answers "should I send this?" A measurement answers "what did I spend?" A system that only does one of them is either uncontrolled or unaccounted.

Every response tells you what it cost — if you read it

Each provider attaches a usage object to the response. The field names differ, and one difference between them will silently corrupt your numbers if you miss it.

Anthropic

JSON
{  "usage": {    "input_tokens": 320,    "output_tokens": 214,    "cache_creation_input_tokens": 0,    "cache_read_input_tokens": 4180  }}

Critical detail: input_tokens here counts only the tokens that were not served from cache. The 4,180 cached tokens are reported separately and are not included in the 320. To get the total tokens the model saw, you add all three input-side fields together. To get the cost, you multiply each of them by its own rate.

OpenAI

JSON
{  "usage": {    "prompt_tokens": 4500,    "completion_tokens": 214,    "total_tokens": 4714,    "prompt_tokens_details": { "cached_tokens": 4180 },    "completion_tokens_details": { "reasoning_tokens": 0 }  }}

Here prompt_tokens is the total, cached tokens included. The 4,180 cached tokens are already inside the 4,500. To price it correctly you must subtract: 4,500 minus 4,180 equals 320 full-price tokens, plus 4,180 at the cached rate. On GPT-5.6 and later models OpenAI also bills cache writes, at 1.25 times the input rate, and reports them as prompt_tokens_details.cache_write_tokens (input_tokens_details.cache_write_tokens in the Responses API). Those are inside the total too, so subtract them as well.

Anthropic excludes cached tokens from the input count; OpenAI includes them. A cost function written for one and pointed at the other will be wrong by the entire size of your cache — which, on a cache-heavy workload, is most of the bill.

Work the numbers to see how bad it gets. Suppose the cached prefix is 4,180 tokens and the fresh input is 320, on a model at 3.00 input and 0.30 cached read per million:

Text
correct : 320 x 3.00/1e6  +  4,180 x 0.30/1e6        = 0.000960        +  0.001254          = 0.002214 USDwrong (treat all 4,500 as full-price input):        = 4,500 x 3.00/1e6                     = 0.013500 USD

That is a 6.1-times overstatement on a single call. Run that dashboard for a month and you will "prove" that caching made things worse.

Google Gemini

JSON
{  "usageMetadata": {    "promptTokenCount": 4500,    "candidatesTokenCount": 214,    "cachedContentTokenCount": 4180,    "thoughtsTokenCount": 0,    "totalTokenCount": 4714  }}

The pattern underneath

ConceptAnthropicOpenAIGoogleGotcha
Fresh inputinput_tokensprompt_tokens minus cachedpromptTokenCount minus cachedOnly Anthropic gives it directly
Cache readscache_read_input_tokensprompt_tokens_details.cached_tokenscachedContentTokenCountBilled at ~10% of input
Cache writescache_creation_input_tokensprompt_tokens_details.cache_write_tokens (GPT-5.6 and later)Storage billed per hourBilled at ~125% of input
Generated textoutput_tokenscompletion_tokenscandidatesTokenCountIncludes reasoning on some models
Reasoning tokensInside output_tokens; broken out in output_tokens_details.thinking_tokenscompletion_tokens_details.reasoning_tokensthoughtsTokenCountInvisible in the text, fully billed

That last row deserves its own warning. With extended or adaptive thinking enabled, a model can spend 3,000 tokens reasoning and then emit a 200-token answer. At an output rate of 15.00 USD per million, the visible answer costs 0.003 dollars and the invisible reasoning costs 0.045 — fifteen times more. If your tracker measures the length of the text you displayed, you are measuring one sixteenth of the output bill.

Streaming: where the usage number hides

In a streaming response there is no single JSON body to read usage from. The counts arrive as events: input usage in the opening event, and a running output count in the delta events, with the final tally in the last one. If your code loops over text chunks and returns, you get the text and no usage at all.

Python
def stream_answer(client, messages, ledger, tenant):    usage = {"input": 0, "output": 0, "cache_read": 0, "cache_write": 0}    try:        with client.messages.stream(            model="claude-sonnet-5", max_tokens=1024, messages=messages,        ) as stream:            for event in stream:                if event.type == "message_start":                    u = event.message.usage                    usage["input"] = u.input_tokens                    usage["cache_read"] = u.cache_read_input_tokens or 0                    usage["cache_write"] = u.cache_creation_input_tokens or 0                elif event.type == "message_delta":                    # cumulative, so overwrite rather than add                    usage["output"] = event.usage.output_tokens                elif (event.type == "content_block_delta"                      and event.delta.type == "text_delta"):                    yield event.delta.text      # skip thinking and tool-input deltas    finally:        # runs even if the client disconnects mid-stream        ledger.record(model="claude-sonnet-5", tenant=tenant, **usage)

Two things make this correct. The message_delta output count is cumulative, so you assign it rather than accumulate it — adding produces a triangular over-count that grows with response length. And the ledger write sits in a finally block, so when a user closes the tab after 400 of a planned 900 tokens, you still record the 400 you were billed for. Aborting a stream stops the display; it does not always stop the generation, and it never refunds what was already produced.

A cost tracker you can actually reuse

The design goal is one normalised shape that every provider adapter feeds into, so that pricing, aggregation, and alerting are written once.

Python
from dataclasses import dataclassfrom datetime import datetime, timezone@dataclass(frozen=True)class Rate:    inp: float; out: float; cache_read: float; cache_write: floatRATES = {                                    # USD per 1M, as of September 2026    "claude-sonnet-5":  Rate(2.00, 10.00, 0.20, 2.50),    "claude-haiku-4-5": Rate(1.00,  5.00, 0.10, 1.25),}@dataclassclass Usage:    input: int = 0    output: int = 0    cache_read: int = 0    cache_write: int = 0def from_anthropic(u):    return Usage(        input=u.input_tokens,                          # already excludes cache        output=u.output_tokens,        cache_read=getattr(u, "cache_read_input_tokens", 0) or 0,        cache_write=getattr(u, "cache_creation_input_tokens", 0) or 0,    )def from_openai(u):    d = u.prompt_tokens_details                        # an object, not a dict    cached = (d.cached_tokens or 0) if d else 0    written = (d.cache_write_tokens or 0) if d else 0    return Usage(        input=u.prompt_tokens - cached - written,      # subtract, or double-count        output=u.completion_tokens,        cache_read=cached,        cache_write=written,    )class CostLedger:    def __init__(self):        self.rows = []    def record(self, model, usage, feature, tenant, request_id, batch=False):        r = RATES[model]        cost = (usage.input * r.inp                + usage.output * r.out                + usage.cache_read * r.cache_read                + usage.cache_write * r.cache_write) / 1_000_000        if batch:            cost *= 0.5                                # async batch discount        self.rows.append({            "ts": datetime.now(timezone.utc),            "model": model, "feature": feature, "tenant": tenant,            "request_id": request_id, "cost_usd": cost,            "input": usage.input, "output": usage.output,            "cache_read": usage.cache_read, "cache_write": usage.cache_write,        })        return cost

Three design decisions in there are worth stating explicitly, because each one corresponds to a real failure.

  • Record per API call, not per user action. The request_id is the provider's, so a tool-use round trip produces two rows. That single choice would have closed 588 of the 658 dollars in the opening story.
  • Carry feature and tenant on every row. A total is nearly useless for decisions. "Which feature and which customer" is what lets you act.
  • Keep the raw token counts alongside the cost. When rates change, you can re-price history. If you only stored dollars, your historical trend becomes a mixture of two rate cards and you can never compare across the change.

From ledger to forecast

A ledger becomes a forecast when you divide it by something meaningful. Take the corrected figure from the opening — 2,398.20 dollars a month — for a product with 3,200 active accounts on a 25-dollar plan, so 80,000 dollars of monthly revenue.

Text
LLM spend as % of revenue : 2,398.20 / 80,000        = 3.0 %LLM cost per account      : 2,398.20 / 3,200         = 0.7494 USD

Three percent of revenue is comfortable. But averages hide the thing that will hurt you. Break the same ledger down by tenant and the usual shape appears: the heaviest 5 percent of accounts — 160 of them — generate 41 percent of the tokens.

Text
spend from the top 160     : 0.41 x 2,398.20         = 983.26 USDper heavy account          : 983.26 / 160            = 6.15 USDas % of that account's plan: 6.15 / 25               = 24.6 %

Those 160 accounts consume a quarter of their subscription price in inference. They are still profitable, but the margin curve is steep, and if usage on that tail doubles, they stop being profitable while the average still looks like a healthy 3 percent. This is why per-tenant attribution is not a nice-to-have. It is the difference between a metric that reassures you and a metric that warns you.

Aggregate cost tells you what happened. Cost per tenant, per feature, and per successful outcome tells you what to do about it.

The most useful denominator is usually not a request at all. For a support product it is cost per resolved ticket; for a coding tool, cost per accepted change. A change that raises cost per request by 20 percent but raises resolution rate by 50 percent is a good change, and only the outcome-denominated metric can see that.

Catching anomalies before the invoice does

Cost incidents are fast and quiet. Consider a real shape of bug: an agent loop appends the tool result to the message list twice per iteration. Nothing errors. Sessions still complete. But context grows twice as fast, and average session cost goes from 0.09 to 2.40 dollars.

At 300 sessions an hour, an alert that fires after 40 minutes costs you 200 sessions times 2.31 dollars of excess, or 462 dollars. The same bug discovered on the invoice a week later costs 300 times 24 times 7 equals 50,400 sessions times 2.31, or 116,424 dollars. The bug is identical. The detection latency is the entire difference.

Alert on rates and distributions, never on running totals — a monthly total only crosses its threshold once the money is gone.

SignalSensible triggerUsually means
Hourly spend rateAbove 3× the trailing 7-day hourly medianTraffic spike, retry storm, or a loop bug
Mean input tokens per callAbove 2× baseline for 15 minutesContext accumulating that should be trimmed
p99 input tokensAny single call above a hard ceilingAn unbounded document or a pasted log file
Cache read ratioDrops below 50% of its usual shareSomething volatile crept into the cached prefix
Calls per user actionAbove 2× the expected tool-loop depthThe agent is looping without converging
Spend for one tenantAbove a per-tenant daily capAbuse, a runaway script, or a genuinely huge customer

Pair the alerts with a hard circuit breaker. A per-tenant daily ceiling that returns a clear "daily limit reached" beats an unbounded bill every time, and it is about fifteen lines of code sitting in front of your client.

Framework-level tracking, and why it is not enough

Orchestration frameworks offer built-in usage callbacks. In LangChain, a callback handler receives every model response and can pull the usage metadata out of it:

Python
from langchain_core.callbacks import BaseCallbackHandlerclass LedgerCallback(BaseCallbackHandler):    def __init__(self, ledger, feature, tenant):        self.ledger, self.feature, self.tenant = ledger, feature, tenant    def on_llm_end(self, response, **kwargs):        for generation in response.generations:            for gen in generation:                msg = getattr(gen, "message", None)                if msg is None:                    continue                meta = msg.usage_metadata or {}                d = meta.get("input_token_details") or {}                read = d.get("cache_read") or 0                write = sum(d.get(k) or 0 for k in ("cache_creation",                            "ephemeral_5m_input_tokens", "ephemeral_1h_input_tokens"))                self.ledger.record(                    model=msg.response_metadata.get("model_name", "claude-sonnet-5"),                    usage=Usage(                        input=meta.get("input_tokens", 0) - read - write,                        output=meta.get("output_tokens", 0),                        cache_read=read, cache_write=write,                    ),                    feature=self.feature, tenant=self.tenant,                    request_id=msg.id,                )# chain.invoke(payload, config={"callbacks": [LedgerCallback(ledger, "summarise", "acme")]})

This is genuinely useful — a chain that makes five internal model calls produces five ledger rows without you touching the chain. Note the subtraction: LangChain's usage_metadata["input_tokens"] is the total input, cache reads and writes included, for every provider — the opposite of Anthropic's raw field — so the handler takes the cached parts out and prices them separately. But be clear about what framework tracking does not give you. Built-in cost helpers ship with a hard-coded price table that lags real rates and usually has no notion of cache read rates or the batch discount. They also cannot attribute spend to your tenant or feature unless you thread that context through yourself. Use the framework for capture; keep pricing and attribution in your own ledger.

The costs that break naive measurement

Hidden costWhy it is missedTypical sizeFix
Tool-use round tripsOne user action, several API calls+30–100% on agentic workloadsRecord per provider request ID
Reasoning / thinking tokensNever appear in the visible answerUp to 15× the visible outputRead the reasoning field; price it as output
Application-level retriesWrapped in a decorator nobody instruments+2–10%Record inside the retry, not around it
Aborted streamsNo final response object to readProportional to abandonment rateRecord from stream events in a finally
Cache writesFeel like part of caching's saving1.25× input rate on the first callPrice cache writes as their own line
Embeddings for retrievalDifferent endpoint, different meterSmall per call, huge at re-index timeRoute embeddings through the same ledger
Images and PDFsSent as files, billed as tokens~1,000–1,600 tokens per page or imageRead the usage block; do not assume
Batch requests priced at full rateTracker does not know the request was batchedOverstates by 2×Carry a batch flag into pricing
Evaluation and test trafficRuns in CI, hits the same keySometimes larger than productionSeparate API keys or a feature tag

Instrument first, optimise second

Every optimisation covered anywhere in this field — shorter prompts, tighter output limits, caching, cheaper models, routing — is a hypothesis about where the money goes. Without measurement you cannot test the hypothesis, and worse, you cannot tell an improvement from a regression that happened to coincide with a quiet week.

The order that works is narrow and boring. Put a ledger in front of every model call and make it record per API request, tagged with feature and tenant. Wire the provider's usage object into it through an adapter, keeping cached and fresh input separate. Add a rate-based alert and a per-tenant circuit breaker before you add anything clever. Then, and only then, look at the split between input, output, cached reads and reasoning, and go after whichever one is largest.

The team in the opening story did the optimisation work first. They spent two weeks shortening prompts and shaved 9 percent off a number that was, at the time, missing 38 percent of the actual spend. The measurement was not the boring prerequisite to the interesting work. It was the work — everything after it took half the effort and produced results they could prove.