Course Content
Token Economics and Cost Optimization
3 sections · 7 lessons
Context Length vs. Cost Tradeoffs
An engineer spent three days building a map-reduce pipeline so that contracts would never be sent to the model in one piece. The reasoning sounded airtight: long context is expensive, so split the document into chunks, summarise each chunk, then summarise the summaries. It shipped, it worked, and it made the bill go up by 71 percent.
Here is the arithmetic that nobody ran first. An 8,000-token contract with a 200-token instruction, answered in about 600 tokens, on a model priced at 3.00 USD per million input and 15.00 per million output, costs 0.0336 dollars as a single call. Split into four 2,000-token chunks, each chunk still needs the 200-token instruction attached, each produces its own 300-token intermediate summary, and then a fifth call has to read all four summaries and produce the final answer. That is 0.0576 dollars — 1.71 times the price of just sending the whole thing.
The chunking did not save context. It multiplied the instruction, added four sets of output tokens at five times the input rate, and threw away every relationship that crossed a chunk boundary. Three days of engineering bought a worse answer for more money.
The mistake is a confusion between two things that sound the same and are not: the size of a model's context window, and the number of tokens you actually spend.
A context window is a capacity, not a price
The context window is the maximum number of tokens a model can hold in a single request — your system prompt, tool schemas, the entire conversation history, any documents you attach, and the tokens it generates in reply, all counted together. It is an architectural ceiling. Exceed it and you get a hard error, not a degraded answer.
Two consequences follow immediately, and both are counterintuitive.
The window costs nothing to have. A model with a 1,000,000-token window does not charge you for 1,000,000 tokens. If you send it 500 tokens, you pay for 500 tokens. Capacity is free; consumption is billed. A 1M-context model at 3.00 USD per million input is cheaper per token than a 200K-context model at 5.00 — the bigger window is not a surcharge.
Output shares the window. If a model has a 200,000-token window and you send 199,000 tokens of input, there is room for roughly 1,000 tokens of output before the request fails. This is a genuinely common production bug: an input that grew slowly over months finally squeezes the reply into nothing, and the model starts truncating answers mid-sentence for reasons that have nothing to do with max_tokens.
You are never billed for the size of the window. You are billed for the tokens you put in it and the tokens that come out of it. Those are completely different quantities and confusing them is the most expensive mistake in this entire topic.
What long context does cost you
Price scales linearly with input tokens. Latency does not. The attention mechanism the model uses to relate every token to every other token has a cost that grows roughly with the square of the sequence length, and although providers charge you linearly, you feel the quadratic part in time to first token.
| Input size | Typical time to first token | Input cost at 3.00 / 1M |
|---|---|---|
| 2,000 tokens | Well under a second | 0.006 USD |
| 8,000 tokens | Around one second | 0.024 USD |
| 50,000 tokens | Several seconds | 0.150 USD |
| 200,000 tokens | Ten seconds or more | 0.600 USD |
The latency figures are indicative and vary by provider and load, but the shape is reliable: cost rises tenfold from 2,000 to 20,000 tokens, and so does the wait. In an interactive product, the wait is often the binding constraint long before the money is.
Worked comparison: three ways to analyse an 8,000-token document
Take the contract-analysis task and price all three architectures honestly. Rates: 3.00 input, 15.00 output, per million tokens. Instruction is 200 tokens; final answer is 600 tokens.
Approach A — one call with the whole document
input : 8,000 + 200 = 8,200 tokens x 3.00/1e6 = 0.024600output : 600 tokens x 15.00/1e6 = 0.009000total 0.033600 USDApproach B — map-reduce over four chunks
MAP (x4): each chunk 2,000 + 200 instruction = 2,200 in, 300 out per chunk : 2,200 x 3.00/1e6 = 0.006600 + 300 x 15.00/1e6 = 0.004500 -> 0.011100 four chunks 0.044400REDUCE : 4 x 300 summaries + 200 instruction = 1,400 in, 600 out 1,400 x 3.00/1e6 = 0.004200 + 600 x 15.00/1e6 = 0.009000 -> 0.013200total 0.057600 USDApproach C — retrieve only the relevant sections
Embed the document once, index it, and at query time pull the three passages that actually bear on the question — about 1,500 tokens.
input : 1,500 + 200 = 1,700 tokens x 3.00/1e6 = 0.005100output : 600 tokens x 15.00/1e6 = 0.009000total 0.014100 USD(plus a one-off embedding cost to build the index)The comparison
| Approach | Per document | Relative to A | At 20,000 docs/month | API round trips | Quality risk |
|---|---|---|---|---|---|
| A — full document, one call | 0.0336 USD | 1.00× | 672 USD | 1 | None; the model sees everything |
| B — map-reduce chunking | 0.0576 USD | 1.71× | 1,152 USD | 5 | High; cross-chunk relations are lost |
| C — retrieval | 0.0141 USD | 0.42× | 282 USD | 1 + a vector lookup | Medium; retrieval can miss a passage |
Chunking to "save context" costs 71 percent more than not bothering, and 4.1 times more than retrieval. The reason is structural and worth stating plainly: chunking duplicates your fixed costs and adds output tokens, which are the expensive kind. The 200-token instruction is paid five times instead of once. Four intermediate summaries are generated at the output rate purely so they can be thrown away after the reduce step.
Splitting a document into chunks does not reduce the tokens you send. It increases them, and converts cheap input tokens into expensive output tokens along the way.
Retrieval wins because it is the only approach that genuinely reduces the token count — it sends 1,700 tokens instead of 8,200 because it worked out that 6,500 of them were irrelevant. That is the actual lever: relevance, not chunking.
When chunking is still right
Chunking earns its place in exactly two situations. First, when the document does not fit in the window at all — a 2,000,000-token archive against a 1,000,000-token model leaves you no choice. Second, when the task is genuinely independent per chunk, such as "extract every monetary amount from each page". In that case there is no reduce step, no cross-chunk reasoning to lose, and the chunks can be sent through an asynchronous batch endpoint at half price. What does not earn its place is chunking a document that fits, for a task that requires seeing the whole thing.
Tiered pricing, and the cliff at the threshold
Some providers charge a higher rate once a single request's input crosses a threshold — and the higher rate applies to the entire request, not just the tokens above the line. Gemini 3.1 Pro is a real example (list prices as of September 2026):
| Request input size | Input rate (USD / 1M) | Output rate (USD / 1M) |
|---|---|---|
| Up to 200,000 tokens | 2.00 | 12.00 |
| Above 200,000 tokens | 4.00 | 18.00 |
Not every provider does this. Anthropic's current models bill their whole 1M-token window at one rate, so on those the cliff below does not exist. Read the pricing page of the exact model you call.
Now price two requests that differ by 2,000 tokens — one percent of their size:
Request A: 199,000 in / 1,000 out input : 199,000 x 2.00/1e6 = 0.398000 output : 1,000 x 12.00/1e6 = 0.012000 total 0.410000 USDRequest B: 201,000 in / 1,000 out input : 201,000 x 4.00/1e6 = 0.804000 output : 1,000 x 18.00/1e6 = 0.018000 total 0.822000 USDratio B/A = 2.00One percent more input, 100 percent more cost. At 50,000 such requests a month that is 20,500 dollars versus 41,100 — a difference of 20,600 dollars a month caused by 2,000 tokens. Any amount of engineering effort under that figure is justified purely by staying below the line.
And here the earlier advice inverts. Splitting request B into two roughly 100,500-token halves puts both under the threshold:
two halves: 2 x (100,500 x 2.00/1e6) = 0.402000outputs : 2 x ( 1,000 x 12.00/1e6) = 0.024000total 0.426000 USD (vs 0.822000)That is a 48 percent saving from the very technique that lost money in the previous section. The difference is not the technique, it is the pricing structure it is applied against. Within a single tier, splitting adds duplicated overhead. Across a tier boundary, splitting avoids a step change that dwarfs the overhead.
Read your provider's tier thresholds before you design your batching strategy. A rule that is correct on flat pricing can be exactly backwards on tiered pricing.
When quality decides instead of cost
There is a failure mode in very long contexts that has nothing to do with money. Models attend less reliably to material buried in the middle of a very long input than to material at the beginning or the end — a pattern usually described as lost in the middle. Fill 120,000 tokens with an entire knowledge base and ask a question whose answer sits at token 60,000, and accuracy drops measurably compared with sending the 8,000 tokens that actually matter.
Suppose a real evaluation gives you these two options:
| Stuff the whole knowledge base | Retrieve the relevant sections | |
|---|---|---|
| Input tokens | 120,000 | 8,200 |
| Cost per call (3.00 / 15.00, 600 out) | 0.369 USD | 0.0336 USD |
| Answer accuracy | 71% | 89% |
| Cost per correct answer | 0.5197 USD | 0.0378 USD |
Dividing cost by accuracy gives the number that actually matters, and it is 13.8 times better for the retrieval approach. Note what happened: the cheaper option was also the more accurate one, so there was no trade-off to agonise over. That is the common case, and it is why "just put everything in the context window" is rarely the right default even for teams with generous budgets.
The genuine trade-off appears when the task needs global reasoning. "Does any clause in this 300-page agreement contradict any other clause?" cannot be answered from three retrieved passages, because the contradiction is precisely the relationship between passages that retrieval did not co-select. For that task, pay for the long context. You are buying a capability, not being wasteful.
A decision framework you can apply in five minutes
| Question | If yes | Why |
|---|---|---|
| Does the answer depend on relationships across the whole input? | Send it whole, in one call | Splitting destroys the relationships you need |
| Is only a small, identifiable fraction of the input relevant per query? | Retrieve, then send | The only technique that truly reduces token count |
| Is the task independent per unit of input? | Many small calls, via the batch endpoint | No reduce step to pay for; 50% async discount applies |
| Does the input exceed the window? | Chunk — you have no choice | Accept the overhead and design the reduce step carefully |
| Are you within about 10% of a pricing tier boundary? | Trim, or split into sub-threshold requests | The whole request reprices, not just the excess |
| Is this an interactive, user-facing path? | Cap input well below the window | Time to first token grows faster than cost does |
| Is the same large prefix reused across requests? | Cache the prefix | Cached reads cost roughly a tenth of fresh input |
Routing by measured size, in code
Once the framework is clear, the implementation is a routing function: measure first, then choose. The important detail is that the measurement is real — an actual token count, not a character-length guess — because every threshold in the table above is denominated in tokens.
1TIER_THRESHOLD = 200_000 # repricing boundary of the model you call;2 # None for flat-priced models (current Claude)3SAFETY_MARGIN = 5_000 # stay clear of the cliff4INTERACTIVE_CAP = 25_000 # latency budget for user-facing paths56def plan_request(input_tokens, task_kind, interactive):7 """Return (model, strategy) for a request of a measured size."""89 if task_kind == "classify" and input_tokens < 2_000:10 # tiny, high volume, no reasoning needed11 return ("claude-haiku-4-5", "single_call")1213 if task_kind == "global_reasoning":14 # must see everything; only question is which tier15 if TIER_THRESHOLD and input_tokens > TIER_THRESHOLD - SAFETY_MARGIN:16 return ("claude-sonnet-5", "split_below_tier")17 return ("claude-sonnet-5", "single_call")1819 if interactive and input_tokens > INTERACTIVE_CAP:20 # latency, not cost, is the binding constraint here21 return ("claude-sonnet-5", "retrieve_then_call")2223 if input_tokens > INTERACTIVE_CAP:24 return ("claude-sonnet-5", "retrieve_then_call")2526 return ("claude-sonnet-5", "single_call")272829def measured_tokens(client, model, system, tools, messages):30 """Never guess the size you are about to route on."""31 r = client.messages.count_tokens(32 model=model, system=system, tools=tools, messages=messages,33 )34 return r.input_tokensNotice that plan_request returns a strategy as well as a model. Choosing the model is only half the decision; the other half is what you send it. A routing function that only picks between a cheap and an expensive model, while sending the same bloated payload to both, leaves most of the available saving on the table.
The failure modes to watch for
| Failure | Symptom | Cause | Fix |
|---|---|---|---|
| Window mistaken for price | Team refuses a long-context model on cost grounds | Capacity confused with consumption | Compare per-token rates, not window sizes |
| Chunking that costs more | Bill rises after a "cost optimisation" | Instruction duplicated; output tokens multiplied | Price the reduce step before building it |
| Silent output truncation | Answers stop mid-sentence at high input sizes | Input plus output exceeds the window | Reserve headroom for max_tokens explicitly |
| Tier cliff | Cost per request doubles for no visible reason | Input drifted past a repricing threshold | Alert on requests within 10% of the boundary |
| Accuracy decay in long context | Correct facts are in the prompt but ignored | Relevant material buried mid-context | Retrieve and re-rank; put key material last |
| Routing on character length | Router picks the wrong tier for JSON payloads | Characters are not tokens | Route on a measured token count |
What this changes in your architecture
The practical shift is to stop treating context length as a setting and start treating it as a budget with three separate currencies: money, which scales linearly with tokens; latency, which scales worse than linearly and is what users feel; and accuracy, which does not simply improve as you add more material and frequently gets worse.
That reframing kills two habits worth killing. It kills "put everything in the prompt, the window is huge" — because the window being huge says nothing about whether the extra 100,000 tokens help, and the evidence usually says they hurt. And it kills "chunk everything to be safe" — because chunking a document that fits is a pure loss on all three currencies at once.
What replaces both is a question you ask per task: what is the smallest set of tokens that contains everything the answer depends on? If that set is the whole document, send the whole document and pay for it without embarrassment. If it is three paragraphs out of forty pages, build retrieval and send three paragraphs. The engineer in the opening story never asked that question — they optimised for a constraint (window size) that was not the one costing them anything, and their pipeline is now three days of code that makes every contract slower, worse, and 71 percent more expensive to read.