Course Content
Token Economics and Cost Optimization
3 sections · 7 lessons
Token Breakdown Across LLM Providers
In the first week of a launch, a four-person team put a number in front of their finance lead: about 22,500 dollars a month for the feature that ranks search results and writes a one-line explanation for each. They got there honestly. Every request sends a JSON block of roughly 100,000 characters describing the candidate products. They had read the common rule of thumb that one token is about four characters, so they wrote down 25,000 input tokens per request, multiplied by 300,000 requests a month, multiplied by the model's input rate of 3.00 USD per million tokens, and shipped.
The first full month's invoice was 37,500 dollars.
Nothing had broken. There was no retry storm, no runaway agent loop, no bug. Traffic was exactly 300,000 requests, as forecast. The only wrong number in the whole calculation was 25,000. Their payload was not English prose — it was JSON: braces, quotes, colons, product SKUs, UUIDs, ISO timestamps. Real tokenizers grind through that at roughly 2.4 characters per token, not 4. Each request actually carried about 41,667 input tokens, which is 1.67 times the estimate, and the invoice came out 1.67 times as large.
That gap is the subject of this lesson. A token is the unit you are billed in, and you cannot reliably infer it from characters, words, or bytes. You have to measure it, per provider, on the payload you will actually send.
What a token actually is
A language model does not see letters. Before any text reaches the model, a tokenizer — a fixed lookup table plus a merging algorithm — cuts the string into pieces called tokens and replaces each piece with an integer ID. The model only ever sees those integers.
The table is built once, before training, by an algorithm called byte pair encoding (BPE). It starts with individual bytes and repeatedly merges the most frequent adjacent pair into a new single entry, until the table reaches a target size — typically somewhere between 100,000 and 200,000 entries for a modern model. The consequence is the thing you need to internalise:
Common sequences get one token; rare sequences get chopped into many. The tokenizer is a compression scheme tuned to the text it was built from, so text that looks unlike that corpus costs more.
Two details trip people up constantly. First, the leading space is part of the token. In most BPE vocabularies " token" (with a space) and "token" (without) are different entries. That is why joining fragments with different spacing can change a token count without changing a single visible character. Second, a token is not a word. "Cat" is one token. "Antidisestablishmentarianism" is six or seven. A UUID such as 9f8b2c1e-4d3a-4f7b-9c2d-1e5a6b7c8d9e is over twenty, because random hex has no frequent sequences for BPE to have merged.
You are billed per token in both directions: tokens you send in (the input or prompt side) and tokens the model generates (the output or completion side). Those two have different prices, usually by a factor of four to six, which matters enormously and which we will come back to.
Why the same string costs different amounts on different providers
Every provider trains its own tokenizer on its own corpus. There is no shared standard. The same sentence is 10 tokens on one model and 12 on another, and the divergence widens sharply for anything that is not ordinary English.
The tokenizer can also change within one provider. Anthropic's pricing page says that Claude models from 4.7 onward use a newer tokenizer that produces roughly 30 percent more tokens for the same text than Claude Sonnet 4.6 and earlier models (as of September 2026). A model upgrade at the same per-token price can therefore still raise the bill, so re-measure your real payload whenever you change model, not only when you change provider.
The table below is a representative measurement, not a guarantee — the exact number depends on the tokenizer version you happen to be calling. What matters is the shape: the characters-per-token ratio, which is the real determinant of your bill.
| Content type | Sample | Characters | Approx. tokens | Chars per token |
|---|---|---|---|---|
| Plain English prose | "The quick brown fox jumps over the lazy dog." | 44 | 10 | 4.4 |
| Technical English | "Idempotent retries prevent duplicate charges." | 45 | 12 | 3.8 |
| Markdown with structure | Headings, bullets, bold markers | 48 | 18 | 2.7 |
| JSON payload | {"user_id": 12345, "status": "active"} | 38 | 16 | 2.4 |
| Python source | for i, (k, v) in enumerate(d.items()): | 38 | 16 | 2.4 |
| UUIDs / hashes | 9f8b2c1e-4d3a-4f7b-9c2d-1e5a6b7c8d9e | 36 | 22 | 1.6 |
| CJK text | Japanese or Chinese sentences | 11 | 9 | 1.2 |
Read the last column again. The four-characters-per-token rule is roughly right for the first row and roughly wrong for everything else. Applied to JSON it undercounts by about 1.7 times. Applied to a table of UUIDs it undercounts by about 2.5 times. Applied to CJK text it undercounts by more than three times — which is exactly why teams serving Japanese or Chinese users often discover their per-user cost is triple what they modelled on English test data.
The practical consequence for prompt design
If you have a choice about the wire format of the data you send, you have a choice about your bill. Sending 40 product records as JSON with fully-spelled-out keys costs meaningfully more than sending the same 40 records as a compact table with a header row, because every repeated {"product_name": is paying for punctuation the model does not need repeated. That is a real optimisation, and it is available to you at design time before you write a single line of caching logic.
Reading the true count: what each provider gives you
Guessing is the mistake. Every major provider ships a way to get the exact count, and they work differently enough that it is worth knowing which is which.
OpenAI — tiktoken, offline
OpenAI publishes its tokenizers as a local library, so counting costs you nothing and involves no network call. You load the encoding for a specific model family and encode the string.
1import tiktoken23enc = tiktoken.encoding_for_model("gpt-4o") # resolves to "o200k_base"45payload = open("catalogue.json").read()6print(len(enc.encode(payload))) # token count of this string onlyencoding_for_model only knows the model names that were mapped when your tiktoken version was released. In tiktoken 0.14, for example, encoding_for_model("gpt-6-sol") raises an error, and you must pick the encoding yourself with tiktoken.get_encoding(...) — which means you are trusting that you picked the right one. When you need the billed number for a current model, OpenAI also offers a counting endpoint, client.responses.input_tokens.count(model=..., input=..., tools=...), which counts the whole request the way the Anthropic and Google endpoints below do.
The catch with any local count: len(enc.encode(text)) counts that string. A real chat request also carries per-message framing tokens, a system prompt, and — if you use tools — the serialised JSON schema of every tool definition. Tool schemas are the single most commonly forgotten cost in the whole field. A six-tool agent can carry 1,500 tokens of schema on every request before the user has typed anything.
Anthropic — the token counting endpoint
Anthropic does not publish an offline tokenizer. Instead there is a free API endpoint that takes exactly the request body you were going to send and returns the number you will be billed for. Because it takes the whole request, it includes the system prompt, message framing, and tool schemas automatically.
1import anthropic23client = anthropic.Anthropic()45resp = client.messages.count_tokens(6 model="claude-sonnet-5",7 system="You rank search results and explain each ranking in one line.",8 tools=[PRODUCT_LOOKUP_TOOL], # schema tokens counted for you9 messages=[{"role": "user", "content": payload}],10)11print(resp.input_tokens) # the exact number the invoice will useThe trade-off versus a local tokenizer is latency and a network dependency, so you would not call this on the hot path for every single request. You call it in tests, in a cost-estimation script, and in a pre-flight guard for unusually large inputs.
Google — the countTokens endpoint
Gemini follows the same shape as Anthropic: a free network call that mirrors the generation request.
1from google import genai23client = genai.Client()4result = client.models.count_tokens(5 model="gemini-3.5-flash",6 contents=payload,7)8print(result.total_tokens)Which one to reach for
| Offline? | Costs money? | Counts tool schemas? | Best used for | |
|---|---|---|---|---|
| tiktoken (OpenAI) | Yes | No | Not automatically | Per-request guards, batch pre-sizing, unit tests |
| input_tokens.count (OpenAI) | No | No, but adds latency | Yes, when tools are passed | Exact estimates for current models |
| count_tokens (Anthropic) | No | No, but adds latency | Yes | Exact estimates, pre-flight size checks, CI budget tests |
| countTokens (Google) | No | No, but adds latency | Yes, when tools are passed | Same as above |
| Any "4 chars ≈ 1 token" heuristic | Yes | No | No | A rough order-of-magnitude sanity check, nothing more |
Use a heuristic to decide whether something is worth measuring. Never use a heuristic to decide what to tell your finance team.
The rate card, and the four numbers that matter
Pricing is quoted per million tokens. For any given model you need four numbers, not two:
- Input rate — what you pay per million tokens you send.
- Output rate — what you pay per million tokens the model generates. Almost always four to eight times the input rate.
- Cache read rate — what you pay for input tokens served from a prompt cache instead of being reprocessed. Typically about one tenth of the input rate, sometimes less.
- Cache write rate — the surcharge for putting a prefix into the cache in the first place. Typically about 1.25 times the input rate for a short-lived cache entry.
Here is the rate card used for the worked examples in this lesson. These are list prices from each provider's pricing page as of September 2026. Prices change several times a year, so treat them as working figures for the arithmetic, not as a quotation, and check the pricing pages before you forecast anything real.
| Model (as of September 2026) | Input (USD / 1M) | Output (USD / 1M) | Cache read | Cache write | Output ÷ input |
|---|---|---|---|---|---|
| Claude Opus 5.5 | 4.00 | 20.00 | 0.20 | 5.00 | 5.0× |
| Claude Sonnet 5 | 2.00 | 10.00 | 0.20 | 2.50 | 5.0× |
| Claude Haiku 4.5 | 1.00 | 5.00 | 0.10 | 1.25 | 5.0× |
| GPT-6 Sol | 2.00 | 10.00 | 0.20 | 2.50 | 5.0× |
| GPT-6 Luna | 0.10 | 0.50 | 0.01 | 0.125 | 5.0× |
| Gemini 3.1 Pro (prompts ≤200K) | 2.00 | 12.00 | 0.20 | storage, per hour | 6.0× |
| Gemini 3.1 Flash-Lite | 0.25 | 1.50 | 0.025 | storage, per hour | 6.0× |
The Claude cache-write column is the 5-minute cache; a 1-hour cache entry costs twice the input rate to write. Google bills an explicit context cache by storage time rather than by a write surcharge.
Why is output more expensive? Because input is processed in one parallel pass over the whole sequence, while output is generated one token at a time, each step requiring a full forward pass through the model. Generating 1,000 tokens is 1,000 sequential model invocations. The price reflects the hardware time.
Two modifiers are worth knowing at this stage. Batch processing — submitting a set of requests to run asynchronously within a 24-hour window — is billed at 50 percent of the standard rate by Anthropic, OpenAI and Google (as of September 2026), which is the single largest discount available and costs you nothing but latency. And some providers apply tiered long-context pricing: once a single request's input crosses a threshold (200,000 tokens on Gemini 3.1 Pro, for example), the entire request is billed at a higher rate, not just the excess. Crossing 200,000 tokens there doubles the input rate and raises the output rate by half for the whole call. OpenAI's GPT-6 models also list separate, higher long-context rates. Anthropic's current models do not: their 1M-token window is billed at the same per-token rate throughout. Check this per model, because it decides whether splitting a large request saves money.
Three worked cost estimates
The formula is always the same:
cost = (input_tokens x input_rate_per_million / 1,000,000) + (output_tokens x output_rate_per_million / 1,000,000)Example 1 — a single classification call
A support ticket router. System prompt of 60 tokens, ticket text of 40 tokens, output is one label of 5 tokens. On Claude Haiku 4.5 (1.00 in, 5.00 out):
input : 100 tokens x 1.00 / 1,000,000 = 0.000100 USDoutput : 5 tokens x 5.00 / 1,000,000 = 0.000025 USDtotal : 0.000125 USD per requestAt 2,000,000 classifications a month that is 0.000125 x 2,000,000 = 250 USD per month. Run the identical workload on Claude Opus 5.5 (4.00 in, 20.00 out) and each request costs 0.000500 USD — four times as much — for 1,000 USD per month. The extra 750 dollars a month buys you frontier reasoning on a task that is picking one of eight labels. That is the trade you are making, stated in money.
Example 2 — document summarisation
A 12,000-token document plus a 150-token instruction, producing a 400-token summary. On Claude Sonnet 5 (2.00 in, 10.00 out):
input : 12,150 tokens x 2.00 / 1,000,000 = 0.024300 USDoutput : 400 tokens x 10.00 / 1,000,000 = 0.004000 USDtotal : 0.028300 USD per documentAt 5,000 documents a day, that is 150,000 documents a month: 0.0283 x 150,000 = 4,245 USD per month.
Now look at the split. Input is 0.0243 of the 0.0283, which is 85.9 percent of the cost. If you spend a week making the summaries shorter — halving output to 200 tokens — you save 0.002 per document, or 300 dollars a month, a 7 percent cut. If instead you cut the document down to the 6,000 tokens that actually contain the relevant sections, you save 0.012 per document, or 1,800 dollars a month, a 42 percent cut. Same week of engineering, six times the return, and the only thing that told you which to do was the split.
Before optimising anything, compute what fraction of the bill is input and what fraction is output. Almost every wasted optimisation effort comes from skipping that one division.
Switching this workload to Claude Haiku 4.5 gives 12,150 x 1.00 / 1e6 = 0.01215 plus 400 x 5.00 / 1e6 = 0.002, for 0.01415 per document, or 2,122.50 USD a month — exactly half the Sonnet bill, a saving of 2,122.50 dollars a month. Whether the summaries are good enough is a quality question, but you now know precisely what the quality is costing you.
Example 3 — a ten-turn support conversation
This is where almost everyone's estimate collapses, so work through it slowly. Setup: a 500-token system prompt, each user turn is 80 tokens, each assistant reply is 150 tokens, ten turns.
The intuitive estimate is "10 turns of 80 in and 150 out, plus a 500-token system prompt" — about 1,300 input tokens and 1,500 output tokens. That estimate is wrong by a factor of more than twelve on the input side, because the API is stateless. It has no memory of your conversation. On every turn you resend the entire history.
turn 1 input = 500 + 80 = 580turn 2 input = 500 + (80+150) + 80 = 810turn 3 input = 500 + (80+150)x2 + 80 = 1,040...turn 10 input = 500 + (80+150)x9 + 80 = 2,650total input = 10 x 580 + 230 x (0+1+...+9) = 5,800 + 230 x 45 = 5,800 + 10,350 = 16,150 tokenstotal output = 10 x 150 = 1,500 tokens16,150 input tokens, not 1,300. The history is resent nine extra times, and the cost of a conversation grows with the square of its length, not linearly. On Claude Sonnet 5:
input : 16,150 x 2.00 / 1,000,000 = 0.032300 USDoutput : 1,500 x 10.00 / 1,000,000 = 0.015000 USDtotal : 0.047300 USD per conversationAt 40,000 conversations a month: 0.0473 x 40,000 = 1,892 USD per month. The naive estimate would have said 1,300 x 2 / 1e6 + 1,500 x 10 / 1e6 = 0.0176 per conversation, or 704 USD a month. The forecast would have been 2.7 times too low — and it would have got worse as user engagement improved, because longer conversations cost superlinearly. A product success metric quietly wired to a cost explosion is one of the nastiest failure modes in this whole area.
Turning the formula into code
Put the rate card in exactly one place, in code, under version control, with the date you copied it. Rates change; hard-coded magic numbers scattered across a codebase are how a price change silently invalidates every dashboard you own.
1from dataclasses import dataclass23@dataclass(frozen=True)4class Rate:5 """All figures are USD per 1,000,000 tokens."""6 inp: float7 out: float8 cache_read: float9 cache_write: float1011RATES_AS_OF = "2026-09" # list prices; re-check the pricing pages12RATES = {13 "claude-opus-5-5": Rate(4.00, 20.00, 0.20, 5.00),14 "claude-sonnet-5": Rate(2.00, 10.00, 0.20, 2.50),15 "claude-haiku-4-5": Rate(1.00, 5.00, 0.10, 1.25),16}1718MILLION = 1_000_0001920def cost_usd(model, input_tokens=0, output_tokens=0,21 cache_read_tokens=0, cache_write_tokens=0):22 r = RATES[model]23 return (24 input_tokens * r.inp +25 output_tokens * r.out +26 cache_read_tokens * r.cache_read +27 cache_write_tokens * r.cache_write28 ) / MILLION2930def monthly(model, per_request_in, per_request_out, requests_per_month):31 unit = cost_usd(model, per_request_in, per_request_out)32 return unit, unit * requests_per_month3334unit, month = monthly("claude-sonnet-5", 12_150, 400, 150_000)35print(f"{unit:.6f} per doc, {month:,.2f} per month")36# 0.028300 per doc, 4,245.00 per monthNote the shape of cost_usd: cached reads and cache writes are separate arguments, not folded into input_tokens. Providers report them as separate fields precisely because they are billed at separate rates, and a cost function that sums them together will overstate a cache-heavy workload by roughly ten times on the cached portion.
Where estimates go wrong
| Mistake | Why it happens | What it does to the number | Fix |
|---|---|---|---|
| Four characters per token, applied to JSON or code | The heuristic is quoted without its caveat | Undercounts by ~1.7× | Measure the real payload with the provider's counter |
| Forgetting tool schemas | They are invisible in application code | Adds 200–2,000 tokens to every request | Count the full request body, tools included |
| Counting only the newest turn | Chat UIs hide the resent history | Undercounts a 10-turn thread by ~12× | Model conversation cost as a running sum |
| Treating output as the same price as input | Only one headline price gets remembered | Undercounts output spend by 4–8× | Always carry two rates per model |
| Assuming tokenizers match across providers | Token counts look interchangeable | 10–30% drift on prose, more on other text | Re-measure when you switch provider |
| Ignoring aborted streams and retries | The user closed the tab, so it feels free | Adds silent spend proportional to your error rate | Bill on tokens generated, not responses delivered |
| Testing on English, shipping to a CJK market | Test fixtures are written in the team's language | Per-user cost up to 3× the model | Include representative locales in cost fixtures |
What to do before you write the first line
Three habits separate teams that get their forecast right from teams that get an invoice surprise.
Measure on a real payload, not a sample sentence. Take an actual production request — the largest realistic one, with the tool schemas and the system prompt attached — and run it through the provider's counter. Write that number down next to your traffic forecast. It takes ten minutes and it is the difference between a 22,500-dollar estimate and a 37,500-dollar invoice.
Compute the input/output split before you optimise. Two numbers, one division. If input is 86 percent of the bill, work on the prompt; if output is 70 percent of it, work on response length and caching. Optimising the wrong side is the most common way to spend a week and save nothing.
Model conversation cost as an accumulator, not a per-request constant. If your product involves multi-turn interaction, the honest unit of cost is the conversation, not the request, and its cost is quadratic in turn count. Put that in the forecast from day one, because that curve does not announce itself until you are already successful.
None of this is about being frugal for its own sake. It is about being able to answer the question "what does this feature cost per user, and what changes if usage triples?" with a number you can defend — before someone else asks it with an invoice in their hand.