Course Content
Token Economics and Cost Optimization
3 sections · 7 lessons
Response Truncation and Caching — Paying Less for the Output Side
A classification endpoint had one job: read a support ticket and return one of eight category labels. Three tokens of output. The team set max_tokens to 4,096, copied from a quickstart, and wrote a prompt that ended with "classify this ticket."
The model, being helpful, wrote a paragraph. It restated the ticket, explained its reasoning, gave the label, and offered to help further. Average output: 180 tokens instead of 3. At 15.00 USD per million output tokens that is 0.0027 dollars per call rather than 0.000045 — sixty times more. Across 1.2 million classifications a month, 3,240 dollars instead of 54.
The reflexive fix is to set max_tokens=5. That is also wrong, and understanding why is the foundation for everything else here: the model does not know what max_tokens is. It is a ceiling enforced by the API after generation starts, not a hint the model plans around. Set it to 5 and you get the first five tokens of the paragraph — "I'll analyse this ticket" — and no label at all. You have not made the response shorter. You have cut it off.
Output tokens are the expensive half of most bills, and there are exactly two ways to spend fewer of them: make the model generate less, or avoid generating at all. This lesson covers both.
Why the output side is where the money is
Input tokens are processed in a single parallel pass over the whole sequence. Output tokens are produced one at a time, each requiring a complete forward pass through the model. Generating 1,000 tokens is 1,000 sequential model invocations. Providers price accordingly — output typically costs four to eight times input.
Run the division on a typical chat turn: 2,000 input tokens, 700 output, at 3.00 and 15.00 per million.
input : 2,000 x 3.00/1e6 = 0.006000 USD (36.4% of the call)output : 700 x 15.00/1e6 = 0.010500 USD (63.6% of the call)total 0.016500 USDUnder a third of the tokens, nearly two thirds of the cost. Any workload where the model writes prose back to a user has this shape, and it means output-side work — shorter formats, tighter limits, caching — has more leverage than prompt trimming.
Before choosing where to optimise, compute the input share and the output share. On a chat product the output is usually the larger number; on a document-analysis product it is usually the smaller. The two need opposite strategies.
Making the model generate less
Prompt for brevity; use max_tokens as the circuit breaker
The order matters. First, ask for what you want: "Respond with only the category label, no explanation." Second, constrain the format — a schema-constrained structured output leaves no room for a preamble at all. Third, set max_tokens as a safety net a little above the longest legitimate response, so that a pathological generation cannot run for 4,000 tokens.
For the classification endpoint, the corrected version prompts for a bare label, uses a structured output with an enum of eight values, and sets max_tokens=16. Output falls to 3 tokens, the ceiling never fires in normal operation, and if the model ever does something strange the damage is capped at 16 tokens rather than 4,096.
Task-specific limits
A single global max_tokens is always wrong for something. Set it per task.
| Task | Sensible max_tokens | Typical actual output | What a 4,096 default does |
|---|---|---|---|
| Single-label classification | 16 | 3–8 | Allows a 180-token essay |
| Yes/no with a one-line reason | 64 | 25–40 | Allows a three-paragraph justification |
| Structured extraction, 10 fields | 400 | 200–260 | Allows commentary around the JSON |
| Article summary | 320 | 180–240 | Allows a summary longer than the article |
| Conversational reply | 800 | 250–350 | Allows rambling on open questions |
| Code generation | 4,000 | 800–1,500 | Roughly right by accident |
| Long-form report | 16,000 | 5,000–7,000 | Truncates mid-document |
Note the last row. A global default is not merely wasteful on short tasks — it is a correctness bug on long ones. The symptom is a report that stops mid-sentence, which teams often misdiagnose as a model problem for weeks.
1TASK_LIMITS = {2 "classify": 16,3 "yes_no": 64,4 "extract": 400,5 "summarise": 320,6 "chat": 800,7 "code": 4000,8 "report": 16000,9}1011def call(client, task, messages, **kw):12 if task not in TASK_LIMITS:13 raise ValueError(f"no token limit defined for task {task!r}")14 return client.messages.create(15 model="claude-sonnet-5",16 max_tokens=TASK_LIMITS[task],17 messages=messages,18 **kw,19 )Raising ValueError for an unknown task is deliberate. A default value is how a new endpoint silently inherits a 4,096-token ceiling nobody chose.
Stop sequences cut at a meaning boundary
A token limit cuts at an arbitrary point. A stop sequence cuts where you say. Generation halts the moment the sequence appears, you are not billed for anything after it, and the sequence itself is not included in the returned text.
This is the right tool when the model reliably appends something you do not want. A model that ends every reply with "Let me know if you need anything else!" is adding roughly 12 tokens to every response. At 15.00 per million across 1.2 million calls, that trailer costs 216 dollars a month. A stop sequence removes it deterministically, where prompt instructions only mostly work.
1resp = client.messages.create(2 model="claude-sonnet-5",3 max_tokens=300,4 stop_sequences=["\n\nSources:", "\n\nLet me know"],5 messages=messages,6)7# resp.stop_reason == "stop_sequence" tells you it firedCheck stop_reason. If it comes back as "max_tokens" you have a truncation, not a completion, and you should treat that as an error rather than shipping half an answer to a user.
Format is a cost decision
Compare two ways of returning the same five extracted fields.
PROSE WITH PREAMBLE (~67 tokens)Certainly! I've analysed the record you provided. Here is whatI found: the customer's name is Jane Doe, their account numberis 88213, they are on the Pro plan, their last payment was on3 March 2026, and their support tier is Gold.STRUCTURED OUTPUT (~35 tokens){"name":"Jane Doe","account":"88213","plan":"Pro", "last_payment":"2026-03-03","tier":"Gold"}48 percent fewer output tokens, and the JSON is machine-readable so nothing downstream has to parse English. There is a second saving that is easy to miss: schema-constrained output cannot be malformed (unless it is cut off by max_tokens or the model refuses, so still check the stop reason). If 4 percent of your 1.2 million calls currently come back as unparseable JSON and get retried, that is 48,000 extra full-price calls at 0.0165 each — 792 dollars a month spent on retries that a schema constraint eliminates entirely.
Not generating at all: caching
The cheapest response is one you already have. But "caching" covers three genuinely different mechanisms, and confusing them is the most common source of disappointment.
| Response caching | Semantic caching | Provider prompt caching | |
|---|---|---|---|
| What is stored | The finished answer | The finished answer, plus an embedding | The model's internal state for a prompt prefix |
| Where | Your Redis / memory | Your vector store | The provider's infrastructure |
| Matches on | Exact key | Similarity above a threshold | Exact byte prefix |
| Cost on a hit | Zero model cost | Embedding call only | ~10% of the input rate |
| Cost on a miss | Full call | Embedding + full call | ~125% of input for the write |
| Best for | FAQs, deterministic tasks | Paraphrased user questions | Long fixed prefixes reused across varied queries |
Exact-match response caching
Hash a normalised request, store the response, serve it on a repeat. Simple, and it saves 100 percent of the model cost on a hit. The whole difficulty is in the key.
1import hashlib, json23PROMPT_VERSION = "support-v7" # bump on every prompt change45def cache_key(model, system, tools, messages, temperature):6 payload = {7 "v": PROMPT_VERSION,8 "model": model,9 "system": system,10 "tools": sorted(t["name"] for t in tools),11 "temperature": temperature,12 "messages": messages,13 }14 blob = json.dumps(payload, sort_keys=True, separators=(",", ":"))15 return "llm:" + hashlib.sha256(blob.encode()).hexdigest()Every field in there corresponds to a bug that happens if you leave it out. Omit model and a routing change serves answers from the wrong model. Omit PROMPT_VERSION and your prompt improvements are invisible for as long as the TTL lasts — the classic "I fixed it but production still does the old thing" incident. Omit temperature and creative and deterministic variants of the same request share a cache entry. Omit sort_keys and two identical requests hash differently because a dictionary iterated in a different order.
A cache key must contain everything that could change the correct answer. The cost of getting this wrong is not a slower system — it is confidently serving a stale answer that nobody can reproduce.
Caching only makes sense at low temperature. If you are deliberately sampling varied responses, caching defeats the point of the feature. (Some current models, including Claude Sonnet 5 and Claude Opus 5.5, reject temperature altogether; for those, put whatever settings they do take — such as the effort level — into the key instead.)
Semantic caching
Exact matching misses "how do I reset my password" against "I forgot my password, how do I reset it". Semantic caching embeds the query and returns a stored answer when the nearest neighbour is similar enough. The extra cost is one embedding call — a 30-token query at typical embedding rates is well under a thousandth of a cent, so the economics are not the issue. The threshold is the issue.
| Similarity threshold | Hit rate | Wrong answers served | Verdict |
|---|---|---|---|
| 0.99 | 12% | ~0% | Barely better than exact match |
| 0.95 | 34% | 0.4% | Reasonable default |
| 0.92 | 44% | 1.6% | Acceptable for low-stakes content |
| 0.90 | 51% | 3.2% | Only for a curated FAQ domain |
| 0.85 | 63% | 11% | Dangerous — one in nine answers is wrong |
The dangerous property of semantic caching is that lowering the threshold improves every metric on your cost dashboard while quietly degrading the product. "Cache hit rate up to 63 percent" reads as a success. It is a failure. Always measure the false-hit rate on a labelled sample before you tune the threshold, and never tune it in response to a cost target.
Two queries that are semantically near-identical but must not share an answer: "how do I cancel my subscription" and "how do I cancel my order". Similarity is high; the correct answers are completely different. Domains full of such pairs — anything with entities, dates, or account-specific data — are poor candidates for semantic caching at any threshold.
Provider-native prompt caching
This one is different in kind. It does not store answers; it stores the model's processed state for a prefix of your prompt, so that resending the same 20,000-token document does not require reprocessing it. Rates typically run at about 1.25 times the input rate to write the cache and about 0.10 times to read it.
The break-even is better than people expect. Let the normal input rate be 1.0. Serving n requests uncached costs n. Serving them with a cache costs 1.25 for the write plus 0.10 for each of the remaining n minus 1:
cached(n) = 1.25 + 0.10 x (n - 1) = 1.15 + 0.10nuncached(n)= 1.00nbreak-even: 1.15 + 0.10n = 1.00n -> n = 1.28Two requests against the same prefix within the cache lifetime already pays for itself. Concretely, a 20,000-token prefix at 3.00 input, 0.30 cached read and 3.75 cache write:
| Requests against the prefix | Uncached (USD) | Cached (USD) | Saving |
|---|---|---|---|
| 1 | 0.060 | 0.075 | −25% (a loss) |
| 2 | 0.120 | 0.081 | 32.5% |
| 10 | 0.600 | 0.129 | 78.5% |
| 100 | 6.000 | 0.669 | 88.9% |
The single-request row is the warning: caching a prefix you use once costs you 25 percent extra. Cache what repeats.
Two mechanical constraints decide whether this works at all. There is a minimum cacheable prefix, and it differs by model — as of September 2026 it is 1,024 tokens on Claude Sonnet 5 and on OpenAI's current models, but 4,096 on Claude Haiku 4.5. Shorter prefixes silently do not cache, and you will see a zero hit rate with no error message. The cache also expires: Anthropic's default entry lives for 5 minutes and each hit refreshes it (a 1-hour entry costs twice the input rate to write), so a prefix used once every ten minutes may never be read at all. And matching is byte-exact on the prefix: a timestamp, a request ID, a randomly ordered JSON key, or a user's name anywhere in the cached region invalidates everything after it. Order your prompt so that the frozen content comes first and everything volatile comes last.
1resp = client.messages.create(2 model="claude-sonnet-5",3 max_tokens=800,4 system=[5 {"type": "text",6 "text": FROZEN_INSTRUCTIONS + "\n\n" + POLICY_DOCUMENT,7 "cache_control": {"type": "ephemeral"}}, # everything above is cached8 ],9 messages=[10 # volatile content lives here, after the breakpoint11 {"role": "user", "content": f"[{now_iso()}] {user_question}"},12 ],13)14print(resp.usage.cache_read_input_tokens) # zero means something invalidated itWatch cache_read_input_tokens in production. A hit rate that was healthy and then went to zero overnight is the single clearest signal that someone added a variable to the cached region.
Deduplicating concurrent identical requests
A cache cannot help with requests that arrive before the first one finishes. When an article starts trending and 240 users hit the same summarisation endpoint within two seconds, all 240 miss the cache, all 240 call the model, and then 239 of them write the same value.
1import asyncio23_inflight: dict[str, asyncio.Future] = {}45async def cached_call(key, produce):6 hit = await redis.get(key)7 if hit is not None:8 return hit910 if key in _inflight:11 return await _inflight[key] # join the existing call1213 fut = asyncio.get_running_loop().create_future()14 _inflight[key] = fut15 try:16 value = await produce()17 await redis.set(key, value, ex=3600)18 fut.set_result(value)19 return value20 except Exception as exc:21 fut.set_exception(exc)22 raise23 finally:24 _inflight.pop(key, None)At 0.0165 dollars per call, that stampede costs 3.96 dollars instead of 0.0165. Once per trending article is trivial; on a news product it happens dozens of times a day.
Sharing the cache across instances
An in-process dictionary is not a cache in a system with twelve pods behind a load balancer — it is twelve caches, each of which must miss independently before it can hit.
Work through the mechanism. Suppose a given question is asked five times an hour and requests are distributed evenly across twelve pods. With a shared cache there is one miss and four hits: an 80 percent hit rate. With per-pod caches, those five requests almost certainly land on five different pods, so there are five misses and zero hits. The hit rate does not degrade gracefully with pod count — for any key requested fewer times than you have pods, it collapses to nothing.
Use a shared store with an explicit TTL. And set the TTL from how fast the underlying truth changes, not from how much you want to save: a product FAQ can sit for a day, a pricing answer for an hour, anything touching account state should probably not be cached at all.
Worked example: an FAQ bot at different hit rates
500,000 requests a month, 1,800 input and 260 output tokens each, at 3.00 and 15.00 per million. Uncached cost per request is 0.0054 plus 0.0039, or 0.0093 dollars, giving 4,650 dollars a month. Add a flat 40 dollars a month for a small Redis instance and a fraction of a dollar for embeddings.
| Cache hit rate | Model calls | Model cost (USD) | Infra (USD) | Total (USD) | Saving |
|---|---|---|---|---|---|
| 0% | 500,000 | 4,650.00 | 40.00 | 4,690.00 | — |
| 20% | 400,000 | 3,720.00 | 40.30 | 3,760.30 | 19.8% |
| 40% | 300,000 | 2,790.00 | 40.30 | 2,830.30 | 39.7% |
| 60% | 200,000 | 1,860.00 | 40.30 | 1,900.30 | 59.5% |
| 80% | 100,000 | 930.00 | 40.30 | 970.30 | 79.3% |
The saving tracks the hit rate almost exactly, because the infrastructure cost is negligible against the model cost. That is the important structural fact: with caching, cost is bounded by your miss rate, not your traffic. Traffic can triple and, if the new traffic is repetitive, the bill barely moves.
Now stack the two halves of this lesson. Change the response format so output falls from 260 to 140 tokens, and run at a 60 percent hit rate:
per-call cost : 1,800 x 3.00/1e6 + 140 x 15.00/1e6 = 0.007500 USDmodel calls : 500,000 x 0.40 = 200,000model cost : 200,000 x 0.0075 = 1,500.00 USDplus infra 40.30 USDtotal 1,540.30 USDversus the 4,690.00 baseline -> 67.2% savingWhere this goes wrong
| Belief | Reality | What it costs you |
|---|---|---|
| "max_tokens makes responses shorter" | It truncates them at an arbitrary point; the model never knew about it | Half-written answers shipped to users |
| "Caching always saves money" | Prefix caching a one-off costs 25% extra; low-repeat traffic never pays back | A cache layer that increases the bill |
| "A higher hit rate is always better" | Semantic hit rate rises by serving wrong answers | Silent quality collapse that looks like a win |
| "Prompt caching caches the answer" | It caches prefix computation; the model still generates every time | Expecting a 100% saving, getting 90% of the input side only |
| "The user's message is enough for a cache key" | Model, prompt version, tools and temperature all change the answer | Stale answers surviving a prompt fix |
| "Streaming avoids output cost" | You are billed for tokens generated, whether or not anyone reads them | Abandoned streams billed in full |
| "An in-memory cache is fine" | N replicas means N independent caches | Hit rate collapses to near zero under load balancing |
| "Set a long TTL for a better hit rate" | TTL is a correctness setting, not a cost setting | Serving last month's pricing to a customer |
The order to do this in
Output-side work has a natural sequence, and doing it out of order wastes effort.
Start by measuring your output distribution, not its average. Log output token counts and look at the p50, p95 and p99. An average of 180 tokens hides two different worlds: a system where every response is around 180, and a system where 90 percent are 40 tokens and 10 percent are 1,500. The second one is fixed by finding out what triggers the long tail; the first is fixed by changing the format.
Then fix the format and the prompt, and only then set the ceiling. Structured outputs and an explicit instruction to skip preamble do the real work. max_tokens is the guard rail that stops a bad day from being an expensive day; it is not the mechanism that makes responses concise.
Then measure your repeat rate before building any cache. Hash a day of production requests and count duplicates. If 3 percent repeat, a response cache will save you 3 percent and cost you a Redis cluster and a class of staleness bugs. If 55 percent repeat, it is the single best change available. That number takes an hour to obtain and decides whether the next two weeks are worth spending.
The classification endpoint from the opening ended up at 3 output tokens with a structured enum, a 16-token ceiling, and a shared cache that caught 71 percent of repeats because support tickets are far more repetitive than anyone expected. The 3,240-dollar line item became about 16 dollars. Nothing about the model changed — only what it was asked to produce, and how often it was asked at all.