Course Content
Token Economics and Cost Optimization
3 sections · 7 lessons
Prompt Compression — Shrinking the Input Before It Costs You
Someone finally opened the system prompt. It had been written eight months earlier, extended four times by three different people, and nobody had read it end to end since. It was 1,400 tokens long. It opened with two sentences thanking the model for its help, contained the instruction "be accurate" three times in slightly different wording, and carried a 300-token refund policy that applied to about eight percent of tickets.
That prompt was sent on every one of 900,000 requests a month. At an input rate of 3.00 USD per million tokens, those 1,400 tokens cost 1,400 times 900,000 times 3, divided by a million: 3,780 dollars a month, for a document nobody had read.
Rewriting it down to 420 tokens took an afternoon. The new cost was 1,134 dollars — a saving of 2,646 dollars a month, or 31,752 a year. Measured accuracy on the team's evaluation set went up by one and a half points, because two of the three "be accurate" instructions had drifted into contradicting each other about when to escalate.
Prompt compression is the practice of removing tokens from the input side without removing capability. It is the highest-leverage cost work available on most systems, for a reason that is arithmetic rather than clever: a token in a static part of your prompt is paid for on every single request. Cut one token from a system prompt sent a million times a month and you have removed a million tokens from the bill.
Where your input tokens actually are
You cannot compress what you have not itemised. Break a request into its components, count each one, and multiply by how often it is sent. Here is that itemisation for the workload above — 900,000 requests a month at 3.00 USD per million input tokens, which makes the monthly cost of any component simply its token count times 2.70.
| Component | Tokens per request | Monthly cost (USD) | Share of input spend | Varies per request? |
|---|---|---|---|---|
| Retrieved context (top-20 chunks) | 8,000 | 21,600 | 63.6% | Yes |
| Few-shot examples (8 of them) | 1,660 | 4,482 | 13.2% | No |
| System prompt | 1,400 | 3,780 | 11.1% | No |
| Tool schemas | 900 | 2,430 | 7.2% | No |
| Conversation history | 500 | 1,350 | 4.0% | Yes |
| User's actual message | 120 | 324 | 0.9% | Yes |
| Total | 12,580 | 33,966 | 100% |
Look at the last row properly. The thing the user typed — the entire point of the request — is under one percent of what you pay to send. Ninety-nine percent of the input bill is scaffolding you wrote. That is the opportunity, and it is also why "ask users to be more concise" is never the answer.
Rank every prompt component by tokens multiplied by frequency, and work down that list. Optimising in any other order is optimising by intuition, and intuition reliably picks the component that is most visible rather than the one that is most expensive.
Technique 1 — trim the static prompt
Static components are the free money. Here is a fragment of the original, and its replacement.
BEFORE (fragment, ~180 tokens)You are a highly capable and extremely helpful customer supportassistant. Thank you for your assistance with our customers. It isvery important that you are accurate in everything that you say.Please always try your very best to be helpful and accurate.When responding to the customer, you should always make sure thatyou are being polite and professional at all times. Please do notever be rude to the customer under any circumstances. You shouldnever insult the customer or use any inappropriate language.Please make sure your answers are accurate and factual and correct.If you are not sure about something then it is better to say thatyou do not know rather than guessing at an answer that might beincorrect, because accuracy is very important to us.AFTER (~34 tokens)You are a customer support assistant.- If you are not certain of a fact, say so. Never guess.- Escalate to a human when the customer asks for a refund above 200 dollars or mentions legal action.What made it shorter
- Deleted politeness aimed at the model. Thanking the model costs tokens and changes nothing. It is not a person and it is not reading its own praise.
- Deleted instructions the model already follows. "Do not insult the customer" and "do not use inappropriate language" describe default behaviour. You are paying to specify the baseline.
- Deduplicated. "Be accurate" appeared three times. Repetition does not increase compliance; it increases the chance that two versions disagree.
- Replaced prose with a list. Three sentences of hedged prose became one bullet with a concrete rule. Lists are both shorter and less ambiguous.
- Made vague rules specific. "Escalate when appropriate" is a rule the model has to interpret every time. "Above 200 dollars, or legal action" is a rule it can apply.
Conditional assembly — the technique most teams miss
The 300-token refund policy was relevant to about 8 percent of tickets and sent on 100 percent of them. Assemble it conditionally and the average cost of that block drops to 300 times 0.08, or 24 tokens per request. The 276 tokens you stop sending, across 900,000 requests at 3.00 per million, are worth 745.20 dollars a month.
1REFUND_POLICY = "..." # 300 tokens2SHIPPING_POLICY = "..." # 260 tokens3WARRANTY_POLICY = "..." # 310 tokens45TRIGGERS = {6 "refund": ("refund", "money back", "chargeback", "return my"),7 "shipping": ("shipping", "delivery", "tracking", "courier"),8 "warranty": ("warranty", "guarantee", "broken", "faulty"),9}10BLOCKS = {"refund": REFUND_POLICY,11 "shipping": SHIPPING_POLICY,12 "warranty": WARRANTY_POLICY}1314def policy_context(ticket_text):15 low = ticket_text.lower()16 return "\n\n".join(17 BLOCKS[name]18 for name, words in TRIGGERS.items()19 if any(w in low for w in words)20 )There is a trap here that is worth more than the saving if you fall into it. Do not put the conditional block inside a cached prefix. Prompt caching works on an exact prefix match — the moment the system prompt varies per request, every cached read becomes a full-price miss, and cached reads cost about a tenth of fresh input. Keep the frozen part frozen and inject the conditional block after the last cache breakpoint, in the message body. Get this the wrong way round and a 745-dollar saving buys you a several-thousand-dollar cache loss.
Technique 2 — structure beats prose
When you send structured data, the format you choose changes the token count dramatically, because verbose formats repeat their keys on every record.
JSON, 12 records, ~30 tokens each = 360 tokens{"product_name": "Widget A", "price_usd": 19.99, "in_stock": true, "category": "tools"}{"product_name": "Widget B", "price_usd": 24.50, "in_stock": false, "category": "tools"}...CSV, header 7 tokens + 12 records at ~12 tokens = 151 tokensname,price,stock,categoryWidget A,19.99,1,toolsWidget B,24.50,0,tools...Same information, 151 tokens instead of 360 — a 58 percent reduction — and the model reads a CSV table perfectly well. The saving comes entirely from not repeating "product_name" and "in_stock" twelve times.
The same principle applies to instructions. A heading and three bullets carry the structure of a rule in fewer tokens than a paragraph that has to describe the structure in words. But there is a floor, and this is where people overshoot:
| Compression | Tokens | Safe? | Why |
|---|---|---|---|
| Verbose JSON with full keys | 360 | Yes | Unambiguous but wasteful |
| CSV with a header row | 151 | Yes | Header carries the meaning once |
| CSV with abbreviated headers plus a legend | 138 | Usually | Legend costs tokens; saving is marginal |
| Headerless positional rows | 132 | No | Model has to infer column meaning; silent misreads |
| Removing units and currency symbols | 126 | No | 19.99 could be dollars, euros or cents |
The last two rows save 19 tokens out of 151 and introduce a class of error that produces confidently wrong answers rather than visible failures. That is the worst trade available in this entire subject.
Technique 3 — prune few-shot examples
Few-shot examples are demonstrations you put in the prompt to show the model the format and edge cases you want. They work. They also have sharply diminishing returns, and almost nobody measures where the returns stop.
Run the same evaluation set at different shot counts and price each one. At 3.00 USD per million input tokens, the cost per 1,000 requests is the prompt token count times 0.003.
| Shots | Prompt tokens | Accuracy | Cost per 1,000 requests | Marginal gain vs. previous |
|---|---|---|---|---|
| 0 | 380 | 71% | 1.14 USD | — |
| 1 | 540 | 86% | 1.62 USD | +15 points for 0.48 USD |
| 2 | 700 | 90% | 2.10 USD | +4 points for 0.48 USD |
| 4 | 1,020 | 91% | 3.06 USD | +1 point for 0.96 USD |
| 8 | 1,660 | 91% | 4.98 USD | 0 points for 1.92 USD |
If your quality bar is 90 percent, two shots clears it and eight shots costs 137 percent more for nothing at all. Across 900,000 requests a month, that is 1,890 dollars against 4,482 — 2,592 dollars a month being spent on six examples that changed no outcome.
Diversity beats quantity. Two examples covering genuinely different edge cases teach more than eight variations on the same easy case, and cost a quarter as much.
If your task really is varied enough to need many examples, use dynamic few-shot: keep a bank of thirty examples, embed the incoming query, and retrieve the two most similar ones. You pay for two examples per request and get coverage that would otherwise require all thirty in the prompt.
Technique 4 — summarise dynamic input before the expensive call
When the bulky part of your prompt is dynamic — a call transcript, a long email thread, a log dump — you cannot delete it in advance. You can, however, have a cheap model read it and hand a short version to the expensive one.
A concrete case: a 20,000-token support-call transcript, and a question that needs strong reasoning to answer well. Frontier model at 5.00 input and 25.00 output; small model at 1.00 and 5.00.
DIRECT (frontier model reads the whole transcript) input : 20,000 x 5.00/1e6 = 0.100000 output : 500 x 25.00/1e6 = 0.012500 total 0.112500 USDTWO-STAGE stage 1, small model summarises to 800 tokens input : 20,000 x 1.00/1e6 = 0.020000 output : 800 x 5.00/1e6 = 0.004000 -> 0.024000 stage 2, frontier model reads the 800-token summary + 150 instruction input : 950 x 5.00/1e6 = 0.004750 output : 500 x 25.00/1e6 = 0.012500 -> 0.017250 total 0.041250 USDThat is a 63.3 percent reduction. At 60,000 transcripts a month, 6,750 dollars becomes 2,475 — a saving of 4,275 dollars a month.
The break-even condition
Two-stage is not automatically cheaper. It adds a whole extra call, so there is a threshold. Write the original length as T and the summary length as S. Two-stage costs T times the cheap input rate, plus S times the cheap output rate, plus S times the expensive input rate. Direct costs T times the expensive input rate. Setting them equal and solving:
S/T = (expensive_input - cheap_input) / (cheap_output + expensive_input)with the rates above:S/T = (5.00 - 1.00) / (5.00 + 5.00) = 0.40The summary must be under 40 percent of the original length to be worth doing. At 800 tokens from 20,000 the ratio is 0.04, comfortably inside. But if your "summary" of a 5,000-token document comes back at 2,500 tokens, the two-stage pipeline is losing money and adding latency. Measure the ratio you actually achieve; do not assume it.
Hierarchical summarisation for very long inputs
When the input is far too large for one summarisation pass, summarise in levels: chunk the source, summarise each chunk with the cheap model, then summarise the summaries, then hand the result to the expensive model. For a 200,000-token archive split into 20 chunks producing 400-token summaries each:
level 1 (small model, 20 chunks) input : 200,000 x 1.00/1e6 = 0.200000 output : 8,000 x 5.00/1e6 = 0.040000 -> 0.240000level 2 (small model, summarise the 8,000 tokens of summaries) input : 8,000 x 1.00/1e6 = 0.008000 output : 1,000 x 5.00/1e6 = 0.005000 -> 0.013000final (frontier model, 1,000-token digest + 150 instruction) input : 1,150 x 5.00/1e6 = 0.005750 output : 500 x 25.00/1e6 = 0.012500 -> 0.018250total 0.271250 USDdirect on the frontier model: 200,000 x 5.00/1e6 + 500 x 25.00/1e6 = 1.012500 USDA 73.2 percent reduction. The price you pay is information loss at every level — a detail dropped in level 1 cannot be recovered later — so this is right for "what happened overall" questions and wrong for "find the one clause that mentions indemnity" questions. For the latter, retrieval is the tool.
Technique 5 — send only what is relevant
Retrieval pipelines commonly fetch the top 20 chunks by vector similarity and paste all of them into the prompt. Twenty chunks at 400 tokens is 8,000 tokens — and in the itemisation at the top of this lesson, that single decision was 63.6 percent of the entire input bill.
Two cheap changes fix most of it:
- Rerank, then truncate. Vector similarity is a coarse first pass. Run the 20 candidates through a reranker — a cross-encoder model, or a cheap LLM scoring pass — and keep the top 3. That is 1,200 tokens instead of 8,000.
- Use a relevance threshold, not a fixed k. If only one chunk scores above the threshold, send one. Fixed
k=20means you always pay for twenty chunks even when the answer is entirely inside the first.
Dropping 6,800 tokens per request across 900,000 requests at 3.00 per million saves 18,360 dollars a month. Accuracy typically improves at the same time, because seventeen irrelevant passages were competing for the model's attention with the three that mattered.
Measuring compression honestly
Three numbers, and a rule about reporting them.
compression ratio = original_tokens / compressed_tokenstoken reduction = (1 - compressed_tokens / original_tokens) x 100 %monthly saving = (original - compressed) x requests x rate / 1,000,0001def compression_report(original_tokens, compressed_tokens,2 requests_per_month, input_rate_per_m,3 accuracy_before, accuracy_after):4 saved = original_tokens - compressed_tokens5 monthly = saved * requests_per_month * input_rate_per_m / 1_000_0006 return {7 "ratio": round(original_tokens / compressed_tokens, 2),8 "reduction_pct": round(100 * saved / original_tokens, 1),9 "monthly_saving_usd": round(monthly, 2),10 "accuracy_delta_points": round(11 100 * (accuracy_after - accuracy_before), 1),12 }1314print(compression_report(12_580, 3_584, 900_000, 3.00, 0.883, 0.891))15# {'ratio': 3.51, 'reduction_pct': 71.5,16# 'monthly_saving_usd': 24289.2, 'accuracy_delta_points': 0.8}Never report a compression number without the quality number beside it. "We cut tokens by 70 percent" is not a result. "We cut tokens by 70 percent and accuracy moved from 88.3 to 89.1" is a result.
Where compression goes wrong
| Mistake | Symptom | Why it happens | Fix |
|---|---|---|---|
| Compressing the wrong component | Weeks of work, single-digit saving | Optimised the visible part, not the expensive part | Itemise by tokens × frequency first |
| Cutting a load-bearing instruction | Quality drops silently, weeks later | No evaluation set gating the change | Re-run evals on every prompt edit |
| Conditional prompt inside the cached prefix | Cache hit rate collapses; bill rises | Prefix no longer matches byte for byte | Inject conditional blocks after the breakpoint |
| Over-abbreviating data | Confidently wrong answers | Removed headers, units, or currency | Keep the legend; stop at CSV with headers |
| Compressing input when output dominates | Big effort, small effect | Never computed the input/output split | Divide input cost by total cost before starting |
| Summariser output too long | Two-stage pipeline costs more than direct | No length cap on the summary | Check the ratio against the break-even formula |
| Measuring in characters | Reported saving does not appear on the invoice | Characters are not tokens, especially for JSON | Measure with the provider's token counter |
| Compressing once and walking away | Prompt bloats back within two quarters | No ownership after the project ends | Assert a token ceiling in CI |
What a full pass actually looks like
Apply all five techniques to the workload we itemised and the numbers land like this — same evaluation set, accuracy from 88.3 to 89.1 percent.
| Component | Before | After | Technique |
|---|---|---|---|
| Retrieved context | 8,000 | 1,200 | Rerank to top 3 |
| Few-shot examples | 1,660 | 700 | Prune 8 shots to 2 |
| System prompt | 1,400 | 444 | Trim, plus conditional policy blocks |
| Tool schemas | 900 | 620 | Terser descriptions; drop two unused tools |
| Conversation history | 500 | 500 | Unchanged |
| User message | 120 | 120 | Unchanged |
| Total input tokens | 12,580 | 3,584 | 71.5% reduction |
| Monthly input cost | 33,966 USD | 9,676.80 USD | 24,289.20 USD saved |
That is 291,470 dollars a year, and not one line of model-serving code changed. Every technique was applied to text that the team wrote themselves and had stopped looking at.
Two habits keep it from growing back. First, put a token ceiling in your test suite — a test that counts the assembled prompt for a fixed fixture and fails if it exceeds an agreed number. Prompts bloat the way logging bloats: one reasonable addition at a time, none of which anyone would object to individually. Second, whenever someone proposes adding to a prompt, make them state which component it lands in and what that component's monthly cost becomes. A 200-token addition to a system prompt on this workload is a 540-dollar-a-month decision, and it should be made as one.