Course Content
Token Economics and Cost Optimization
3 sections · 7 lessons
Measuring Token Usage in Applications
A team built a cost dashboard. It was not a toy: it logged every call, multiplied tokens by the published rate, and rolled the result up by day. At the end of the month it read 1,740 dollars. The invoice read 2,398 dollars — 38 percent higher.
The dashboard was not buggy in the ordinary sense. Every line it recorded was correct. The problem was the lines it never recorded at all.
Their assistant handled 200,000 user messages a month, each averaging 1,800 input and 220 output tokens on a model priced at 3.00 and 15.00 USD per million. That is 0.0054 plus 0.0033, or 0.0087 dollars a message — 1,740 dollars across 200,000 messages, exactly as the dashboard said. But 30 percent of those messages triggered a tool call, and a tool call is not one API request, it is two: the first returns a request to run the tool, the second resends the whole conversation plus the assistant's tool-use block plus the tool result, and generates the real answer. That second call carried about 2,370 input and 180 output tokens — 0.00981 dollars — and there were 60,000 of them, for 588.60 dollars that the dashboard never saw. Another 4 percent of calls came back with malformed JSON and were silently re-asked by a retry wrapper: 8,000 extra calls at 0.0087, or 69.60 dollars. Add it up: 1,740 plus 588.60 plus 69.60 equals 2,398.20.
Nobody had lied and nothing had leaked. The measurement simply counted user messages when the meter counts API calls. This lesson is about closing that gap: reading what each provider actually reports, normalising it into one shape, and turning it into a number you can forecast and alert on.
Estimating and measuring are two different jobs
People conflate these constantly, and then wonder why their pre-flight token estimate does not match the invoice. They are not the same activity and they do not have the same purpose.
| Estimating | Measuring | |
|---|---|---|
| When it happens | Before the request | After the response |
| How | Local tokenizer or a counting endpoint | Read the usage block on the response |
| What it knows | Input only | Input, output, cache reads, cache writes, reasoning |
| Accuracy | Good on input; blind to output | Exact — it is the billing record |
| What it is for | Guardrails: reject or trim before you pay | Accounting, forecasting, anomaly detection |
| Failure if you skip it | A 400,000-token pasted log slips through and costs 1.20 USD of input at 3.00 per million — on every retry | You find out what you spent from the invoice |
You need both. Estimation is a gate on the way in. Measurement is the ledger on the way out. Using an estimate as your ledger is the specific mistake that makes dashboards drift from invoices, because an estimate cannot know how many tokens the model chose to generate.
An estimate answers "should I send this?" A measurement answers "what did I spend?" A system that only does one of them is either uncontrolled or unaccounted.
Every response tells you what it cost — if you read it
Each provider attaches a usage object to the response. The field names differ, and one difference between them will silently corrupt your numbers if you miss it.
Anthropic
1{2 "usage": {3 "input_tokens": 320,4 "output_tokens": 214,5 "cache_creation_input_tokens": 0,6 "cache_read_input_tokens": 41807 }8}Critical detail: input_tokens here counts only the tokens that were not served from cache. The 4,180 cached tokens are reported separately and are not included in the 320. To get the total tokens the model saw, you add all three input-side fields together. To get the cost, you multiply each of them by its own rate.
OpenAI
1{2 "usage": {3 "prompt_tokens": 4500,4 "completion_tokens": 214,5 "total_tokens": 4714,6 "prompt_tokens_details": { "cached_tokens": 4180 },7 "completion_tokens_details": { "reasoning_tokens": 0 }8 }9}Here prompt_tokens is the total, cached tokens included. The 4,180 cached tokens are already inside the 4,500. To price it correctly you must subtract: 4,500 minus 4,180 equals 320 full-price tokens, plus 4,180 at the cached rate. On GPT-5.6 and later models OpenAI also bills cache writes, at 1.25 times the input rate, and reports them as prompt_tokens_details.cache_write_tokens (input_tokens_details.cache_write_tokens in the Responses API). Those are inside the total too, so subtract them as well.
Anthropic excludes cached tokens from the input count; OpenAI includes them. A cost function written for one and pointed at the other will be wrong by the entire size of your cache — which, on a cache-heavy workload, is most of the bill.
Work the numbers to see how bad it gets. Suppose the cached prefix is 4,180 tokens and the fresh input is 320, on a model at 3.00 input and 0.30 cached read per million:
correct : 320 x 3.00/1e6 + 4,180 x 0.30/1e6 = 0.000960 + 0.001254 = 0.002214 USDwrong (treat all 4,500 as full-price input): = 4,500 x 3.00/1e6 = 0.013500 USDThat is a 6.1-times overstatement on a single call. Run that dashboard for a month and you will "prove" that caching made things worse.
Google Gemini
1{2 "usageMetadata": {3 "promptTokenCount": 4500,4 "candidatesTokenCount": 214,5 "cachedContentTokenCount": 4180,6 "thoughtsTokenCount": 0,7 "totalTokenCount": 47148 }9}The pattern underneath
| Concept | Anthropic | OpenAI | Gotcha | |
|---|---|---|---|---|
| Fresh input | input_tokens | prompt_tokens minus cached | promptTokenCount minus cached | Only Anthropic gives it directly |
| Cache reads | cache_read_input_tokens | prompt_tokens_details.cached_tokens | cachedContentTokenCount | Billed at ~10% of input |
| Cache writes | cache_creation_input_tokens | prompt_tokens_details.cache_write_tokens (GPT-5.6 and later) | Storage billed per hour | Billed at ~125% of input |
| Generated text | output_tokens | completion_tokens | candidatesTokenCount | Includes reasoning on some models |
| Reasoning tokens | Inside output_tokens; broken out in output_tokens_details.thinking_tokens | completion_tokens_details.reasoning_tokens | thoughtsTokenCount | Invisible in the text, fully billed |
That last row deserves its own warning. With extended or adaptive thinking enabled, a model can spend 3,000 tokens reasoning and then emit a 200-token answer. At an output rate of 15.00 USD per million, the visible answer costs 0.003 dollars and the invisible reasoning costs 0.045 — fifteen times more. If your tracker measures the length of the text you displayed, you are measuring one sixteenth of the output bill.
Streaming: where the usage number hides
In a streaming response there is no single JSON body to read usage from. The counts arrive as events: input usage in the opening event, and a running output count in the delta events, with the final tally in the last one. If your code loops over text chunks and returns, you get the text and no usage at all.
1def stream_answer(client, messages, ledger, tenant):2 usage = {"input": 0, "output": 0, "cache_read": 0, "cache_write": 0}3 try:4 with client.messages.stream(5 model="claude-sonnet-5", max_tokens=1024, messages=messages,6 ) as stream:7 for event in stream:8 if event.type == "message_start":9 u = event.message.usage10 usage["input"] = u.input_tokens11 usage["cache_read"] = u.cache_read_input_tokens or 012 usage["cache_write"] = u.cache_creation_input_tokens or 013 elif event.type == "message_delta":14 # cumulative, so overwrite rather than add15 usage["output"] = event.usage.output_tokens16 elif (event.type == "content_block_delta"17 and event.delta.type == "text_delta"):18 yield event.delta.text # skip thinking and tool-input deltas19 finally:20 # runs even if the client disconnects mid-stream21 ledger.record(model="claude-sonnet-5", tenant=tenant, **usage)Two things make this correct. The message_delta output count is cumulative, so you assign it rather than accumulate it — adding produces a triangular over-count that grows with response length. And the ledger write sits in a finally block, so when a user closes the tab after 400 of a planned 900 tokens, you still record the 400 you were billed for. Aborting a stream stops the display; it does not always stop the generation, and it never refunds what was already produced.
A cost tracker you can actually reuse
The design goal is one normalised shape that every provider adapter feeds into, so that pricing, aggregation, and alerting are written once.
1from dataclasses import dataclass2from datetime import datetime, timezone34@dataclass(frozen=True)5class Rate:6 inp: float; out: float; cache_read: float; cache_write: float78RATES = { # USD per 1M, as of September 20269 "claude-sonnet-5": Rate(2.00, 10.00, 0.20, 2.50),10 "claude-haiku-4-5": Rate(1.00, 5.00, 0.10, 1.25),11}1213@dataclass14class Usage:15 input: int = 016 output: int = 017 cache_read: int = 018 cache_write: int = 01920def from_anthropic(u):21 return Usage(22 input=u.input_tokens, # already excludes cache23 output=u.output_tokens,24 cache_read=getattr(u, "cache_read_input_tokens", 0) or 0,25 cache_write=getattr(u, "cache_creation_input_tokens", 0) or 0,26 )2728def from_openai(u):29 d = u.prompt_tokens_details # an object, not a dict30 cached = (d.cached_tokens or 0) if d else 031 written = (d.cache_write_tokens or 0) if d else 032 return Usage(33 input=u.prompt_tokens - cached - written, # subtract, or double-count34 output=u.completion_tokens,35 cache_read=cached,36 cache_write=written,37 )3839class CostLedger:40 def __init__(self):41 self.rows = []4243 def record(self, model, usage, feature, tenant, request_id, batch=False):44 r = RATES[model]45 cost = (usage.input * r.inp46 + usage.output * r.out47 + usage.cache_read * r.cache_read48 + usage.cache_write * r.cache_write) / 1_000_00049 if batch:50 cost *= 0.5 # async batch discount51 self.rows.append({52 "ts": datetime.now(timezone.utc),53 "model": model, "feature": feature, "tenant": tenant,54 "request_id": request_id, "cost_usd": cost,55 "input": usage.input, "output": usage.output,56 "cache_read": usage.cache_read, "cache_write": usage.cache_write,57 })58 return costThree design decisions in there are worth stating explicitly, because each one corresponds to a real failure.
- Record per API call, not per user action. The
request_idis the provider's, so a tool-use round trip produces two rows. That single choice would have closed 588 of the 658 dollars in the opening story. - Carry
featureandtenanton every row. A total is nearly useless for decisions. "Which feature and which customer" is what lets you act. - Keep the raw token counts alongside the cost. When rates change, you can re-price history. If you only stored dollars, your historical trend becomes a mixture of two rate cards and you can never compare across the change.
From ledger to forecast
A ledger becomes a forecast when you divide it by something meaningful. Take the corrected figure from the opening — 2,398.20 dollars a month — for a product with 3,200 active accounts on a 25-dollar plan, so 80,000 dollars of monthly revenue.
LLM spend as % of revenue : 2,398.20 / 80,000 = 3.0 %LLM cost per account : 2,398.20 / 3,200 = 0.7494 USDThree percent of revenue is comfortable. But averages hide the thing that will hurt you. Break the same ledger down by tenant and the usual shape appears: the heaviest 5 percent of accounts — 160 of them — generate 41 percent of the tokens.
spend from the top 160 : 0.41 x 2,398.20 = 983.26 USDper heavy account : 983.26 / 160 = 6.15 USDas % of that account's plan: 6.15 / 25 = 24.6 %Those 160 accounts consume a quarter of their subscription price in inference. They are still profitable, but the margin curve is steep, and if usage on that tail doubles, they stop being profitable while the average still looks like a healthy 3 percent. This is why per-tenant attribution is not a nice-to-have. It is the difference between a metric that reassures you and a metric that warns you.
Aggregate cost tells you what happened. Cost per tenant, per feature, and per successful outcome tells you what to do about it.
The most useful denominator is usually not a request at all. For a support product it is cost per resolved ticket; for a coding tool, cost per accepted change. A change that raises cost per request by 20 percent but raises resolution rate by 50 percent is a good change, and only the outcome-denominated metric can see that.
Catching anomalies before the invoice does
Cost incidents are fast and quiet. Consider a real shape of bug: an agent loop appends the tool result to the message list twice per iteration. Nothing errors. Sessions still complete. But context grows twice as fast, and average session cost goes from 0.09 to 2.40 dollars.
At 300 sessions an hour, an alert that fires after 40 minutes costs you 200 sessions times 2.31 dollars of excess, or 462 dollars. The same bug discovered on the invoice a week later costs 300 times 24 times 7 equals 50,400 sessions times 2.31, or 116,424 dollars. The bug is identical. The detection latency is the entire difference.
Alert on rates and distributions, never on running totals — a monthly total only crosses its threshold once the money is gone.
| Signal | Sensible trigger | Usually means |
|---|---|---|
| Hourly spend rate | Above 3× the trailing 7-day hourly median | Traffic spike, retry storm, or a loop bug |
| Mean input tokens per call | Above 2× baseline for 15 minutes | Context accumulating that should be trimmed |
| p99 input tokens | Any single call above a hard ceiling | An unbounded document or a pasted log file |
| Cache read ratio | Drops below 50% of its usual share | Something volatile crept into the cached prefix |
| Calls per user action | Above 2× the expected tool-loop depth | The agent is looping without converging |
| Spend for one tenant | Above a per-tenant daily cap | Abuse, a runaway script, or a genuinely huge customer |
Pair the alerts with a hard circuit breaker. A per-tenant daily ceiling that returns a clear "daily limit reached" beats an unbounded bill every time, and it is about fifteen lines of code sitting in front of your client.
Framework-level tracking, and why it is not enough
Orchestration frameworks offer built-in usage callbacks. In LangChain, a callback handler receives every model response and can pull the usage metadata out of it:
1from langchain_core.callbacks import BaseCallbackHandler23class LedgerCallback(BaseCallbackHandler):4 def __init__(self, ledger, feature, tenant):5 self.ledger, self.feature, self.tenant = ledger, feature, tenant67 def on_llm_end(self, response, **kwargs):8 for generation in response.generations:9 for gen in generation:10 msg = getattr(gen, "message", None)11 if msg is None:12 continue13 meta = msg.usage_metadata or {}14 d = meta.get("input_token_details") or {}15 read = d.get("cache_read") or 016 write = sum(d.get(k) or 0 for k in ("cache_creation",17 "ephemeral_5m_input_tokens", "ephemeral_1h_input_tokens"))18 self.ledger.record(19 model=msg.response_metadata.get("model_name", "claude-sonnet-5"),20 usage=Usage(21 input=meta.get("input_tokens", 0) - read - write,22 output=meta.get("output_tokens", 0),23 cache_read=read, cache_write=write,24 ),25 feature=self.feature, tenant=self.tenant,26 request_id=msg.id,27 )2829# chain.invoke(payload, config={"callbacks": [LedgerCallback(ledger, "summarise", "acme")]})This is genuinely useful — a chain that makes five internal model calls produces five ledger rows without you touching the chain. Note the subtraction: LangChain's usage_metadata["input_tokens"] is the total input, cache reads and writes included, for every provider — the opposite of Anthropic's raw field — so the handler takes the cached parts out and prices them separately. But be clear about what framework tracking does not give you. Built-in cost helpers ship with a hard-coded price table that lags real rates and usually has no notion of cache read rates or the batch discount. They also cannot attribute spend to your tenant or feature unless you thread that context through yourself. Use the framework for capture; keep pricing and attribution in your own ledger.
The costs that break naive measurement
| Hidden cost | Why it is missed | Typical size | Fix |
|---|---|---|---|
| Tool-use round trips | One user action, several API calls | +30–100% on agentic workloads | Record per provider request ID |
| Reasoning / thinking tokens | Never appear in the visible answer | Up to 15× the visible output | Read the reasoning field; price it as output |
| Application-level retries | Wrapped in a decorator nobody instruments | +2–10% | Record inside the retry, not around it |
| Aborted streams | No final response object to read | Proportional to abandonment rate | Record from stream events in a finally |
| Cache writes | Feel like part of caching's saving | 1.25× input rate on the first call | Price cache writes as their own line |
| Embeddings for retrieval | Different endpoint, different meter | Small per call, huge at re-index time | Route embeddings through the same ledger |
| Images and PDFs | Sent as files, billed as tokens | ~1,000–1,600 tokens per page or image | Read the usage block; do not assume |
| Batch requests priced at full rate | Tracker does not know the request was batched | Overstates by 2× | Carry a batch flag into pricing |
| Evaluation and test traffic | Runs in CI, hits the same key | Sometimes larger than production | Separate API keys or a feature tag |
Instrument first, optimise second
Every optimisation covered anywhere in this field — shorter prompts, tighter output limits, caching, cheaper models, routing — is a hypothesis about where the money goes. Without measurement you cannot test the hypothesis, and worse, you cannot tell an improvement from a regression that happened to coincide with a quiet week.
The order that works is narrow and boring. Put a ledger in front of every model call and make it record per API request, tagged with feature and tenant. Wire the provider's usage object into it through an adapter, keeping cached and fresh input separate. Add a rate-based alert and a per-tenant circuit breaker before you add anything clever. Then, and only then, look at the split between input, output, cached reads and reasoning, and go after whichever one is largest.
The team in the opening story did the optimisation work first. They spent two weeks shortening prompts and shaved 9 percent off a number that was, at the time, missing 38 percent of the actual spend. The measurement was not the boring prerequisite to the interesting work. It was the work — everything after it took half the effort and produced results they could prove.