Token Economics and Cost Optimization

Prompt Compression — Shrinking the Input Before It Costs You


Someone finally opened the system prompt. It had been written eight months earlier, extended four times by three different people, and nobody had read it end to end since. It was 1,400 tokens long. It opened with two sentences thanking the model for its help, contained the instruction "be accurate" three times in slightly different wording, and carried a 300-token refund policy that applied to about eight percent of tickets.

That prompt was sent on every one of 900,000 requests a month. At an input rate of 3.00 USD per million tokens, those 1,400 tokens cost 1,400 times 900,000 times 3, divided by a million: 3,780 dollars a month, for a document nobody had read.

Rewriting it down to 420 tokens took an afternoon. The new cost was 1,134 dollars — a saving of 2,646 dollars a month, or 31,752 a year. Measured accuracy on the team's evaluation set went up by one and a half points, because two of the three "be accurate" instructions had drifted into contradicting each other about when to escalate.

Prompt compression is the practice of removing tokens from the input side without removing capability. It is the highest-leverage cost work available on most systems, for a reason that is arithmetic rather than clever: a token in a static part of your prompt is paid for on every single request. Cut one token from a system prompt sent a million times a month and you have removed a million tokens from the bill.

Where the 6,820 input tokens actually areSystemprompt — 1,400Few-shotexamples — 900Retrievedcontext — 2,600History — 1,800User turn — 120topbottomThe user's actual question is 1.8 percent of what you send and pay for on every call.
Compression only pays where the tokens are, and they are almost never in the message the user typed.

Where your input tokens actually are

You cannot compress what you have not itemised. Break a request into its components, count each one, and multiply by how often it is sent. Here is that itemisation for the workload above — 900,000 requests a month at 3.00 USD per million input tokens, which makes the monthly cost of any component simply its token count times 2.70.

ComponentTokens per requestMonthly cost (USD)Share of input spendVaries per request?
Retrieved context (top-20 chunks)8,00021,60063.6%Yes
Few-shot examples (8 of them)1,6604,48213.2%No
System prompt1,4003,78011.1%No
Tool schemas9002,4307.2%No
Conversation history5001,3504.0%Yes
User's actual message1203240.9%Yes
Total12,58033,966100%

Look at the last row properly. The thing the user typed — the entire point of the request — is under one percent of what you pay to send. Ninety-nine percent of the input bill is scaffolding you wrote. That is the opportunity, and it is also why "ask users to be more concise" is never the answer.

Rank every prompt component by tokens multiplied by frequency, and work down that list. Optimising in any other order is optimising by intuition, and intuition reliably picks the component that is most visible rather than the one that is most expensive.

Technique 1 — trim the static prompt

Static components are the free money. Here is a fragment of the original, and its replacement.

Text
BEFORE (fragment, ~180 tokens)You are a highly capable and extremely helpful customer supportassistant. Thank you for your assistance with our customers. It isvery important that you are accurate in everything that you say.Please always try your very best to be helpful and accurate.When responding to the customer, you should always make sure thatyou are being polite and professional at all times. Please do notever be rude to the customer under any circumstances. You shouldnever insult the customer or use any inappropriate language.Please make sure your answers are accurate and factual and correct.If you are not sure about something then it is better to say thatyou do not know rather than guessing at an answer that might beincorrect, because accuracy is very important to us.
Text
AFTER (~34 tokens)You are a customer support assistant.- If you are not certain of a fact, say so. Never guess.- Escalate to a human when the customer asks for a refund  above 200 dollars or mentions legal action.

What made it shorter

  • Deleted politeness aimed at the model. Thanking the model costs tokens and changes nothing. It is not a person and it is not reading its own praise.
  • Deleted instructions the model already follows. "Do not insult the customer" and "do not use inappropriate language" describe default behaviour. You are paying to specify the baseline.
  • Deduplicated. "Be accurate" appeared three times. Repetition does not increase compliance; it increases the chance that two versions disagree.
  • Replaced prose with a list. Three sentences of hedged prose became one bullet with a concrete rule. Lists are both shorter and less ambiguous.
  • Made vague rules specific. "Escalate when appropriate" is a rule the model has to interpret every time. "Above 200 dollars, or legal action" is a rule it can apply.

Conditional assembly — the technique most teams miss

The 300-token refund policy was relevant to about 8 percent of tickets and sent on 100 percent of them. Assemble it conditionally and the average cost of that block drops to 300 times 0.08, or 24 tokens per request. The 276 tokens you stop sending, across 900,000 requests at 3.00 per million, are worth 745.20 dollars a month.

Python
REFUND_POLICY = "..."      # 300 tokensSHIPPING_POLICY = "..."    # 260 tokensWARRANTY_POLICY = "..."    # 310 tokensTRIGGERS = {    "refund":   ("refund", "money back", "chargeback", "return my"),    "shipping": ("shipping", "delivery", "tracking", "courier"),    "warranty": ("warranty", "guarantee", "broken", "faulty"),}BLOCKS = {"refund": REFUND_POLICY,          "shipping": SHIPPING_POLICY,          "warranty": WARRANTY_POLICY}def policy_context(ticket_text):    low = ticket_text.lower()    return "\n\n".join(        BLOCKS[name]        for name, words in TRIGGERS.items()        if any(w in low for w in words)    )

There is a trap here that is worth more than the saving if you fall into it. Do not put the conditional block inside a cached prefix. Prompt caching works on an exact prefix match — the moment the system prompt varies per request, every cached read becomes a full-price miss, and cached reads cost about a tenth of fresh input. Keep the frozen part frozen and inject the conditional block after the last cache breakpoint, in the message body. Get this the wrong way round and a 745-dollar saving buys you a several-thousand-dollar cache loss.

Technique 2 — structure beats prose

When you send structured data, the format you choose changes the token count dramatically, because verbose formats repeat their keys on every record.

Text
JSON, 12 records, ~30 tokens each = 360 tokens{"product_name": "Widget A", "price_usd": 19.99, "in_stock": true, "category": "tools"}{"product_name": "Widget B", "price_usd": 24.50, "in_stock": false, "category": "tools"}...
Text
CSV, header 7 tokens + 12 records at ~12 tokens = 151 tokensname,price,stock,categoryWidget A,19.99,1,toolsWidget B,24.50,0,tools...

Same information, 151 tokens instead of 360 — a 58 percent reduction — and the model reads a CSV table perfectly well. The saving comes entirely from not repeating "product_name" and "in_stock" twelve times.

The same principle applies to instructions. A heading and three bullets carry the structure of a rule in fewer tokens than a paragraph that has to describe the structure in words. But there is a floor, and this is where people overshoot:

CompressionTokensSafe?Why
Verbose JSON with full keys360YesUnambiguous but wasteful
CSV with a header row151YesHeader carries the meaning once
CSV with abbreviated headers plus a legend138UsuallyLegend costs tokens; saving is marginal
Headerless positional rows132NoModel has to infer column meaning; silent misreads
Removing units and currency symbols126No19.99 could be dollars, euros or cents

The last two rows save 19 tokens out of 151 and introduce a class of error that produces confidently wrong answers rather than visible failures. That is the worst trade available in this entire subject.

Technique 3 — prune few-shot examples

Few-shot examples are demonstrations you put in the prompt to show the model the format and edge cases you want. They work. They also have sharply diminishing returns, and almost nobody measures where the returns stop.

Run the same evaluation set at different shot counts and price each one. At 3.00 USD per million input tokens, the cost per 1,000 requests is the prompt token count times 0.003.

ShotsPrompt tokensAccuracyCost per 1,000 requestsMarginal gain vs. previous
038071%1.14 USD—
154086%1.62 USD+15 points for 0.48 USD
270090%2.10 USD+4 points for 0.48 USD
41,02091%3.06 USD+1 point for 0.96 USD
81,66091%4.98 USD0 points for 1.92 USD

If your quality bar is 90 percent, two shots clears it and eight shots costs 137 percent more for nothing at all. Across 900,000 requests a month, that is 1,890 dollars against 4,482 — 2,592 dollars a month being spent on six examples that changed no outcome.

Diversity beats quantity. Two examples covering genuinely different edge cases teach more than eight variations on the same easy case, and cost a quarter as much.

If your task really is varied enough to need many examples, use dynamic few-shot: keep a bank of thirty examples, embed the incoming query, and retrieve the two most similar ones. You pay for two examples per request and get coverage that would otherwise require all thirty in the prompt.

Technique 4 — summarise dynamic input before the expensive call

When the bulky part of your prompt is dynamic — a call transcript, a long email thread, a log dump — you cannot delete it in advance. You can, however, have a cheap model read it and hand a short version to the expensive one.

A concrete case: a 20,000-token support-call transcript, and a question that needs strong reasoning to answer well. Frontier model at 5.00 input and 25.00 output; small model at 1.00 and 5.00.

Text
DIRECT (frontier model reads the whole transcript)  input  : 20,000 x  5.00/1e6 = 0.100000  output :    500 x 25.00/1e6 = 0.012500  total                         0.112500 USDTWO-STAGE  stage 1, small model summarises to 800 tokens    input  : 20,000 x 1.00/1e6 = 0.020000    output :    800 x 5.00/1e6 = 0.004000   ->  0.024000  stage 2, frontier model reads the 800-token summary + 150 instruction    input  :    950 x  5.00/1e6 = 0.004750    output :    500 x 25.00/1e6 = 0.012500  ->  0.017250  total                                        0.041250 USD

That is a 63.3 percent reduction. At 60,000 transcripts a month, 6,750 dollars becomes 2,475 — a saving of 4,275 dollars a month.

The break-even condition

Two-stage is not automatically cheaper. It adds a whole extra call, so there is a threshold. Write the original length as T and the summary length as S. Two-stage costs T times the cheap input rate, plus S times the cheap output rate, plus S times the expensive input rate. Direct costs T times the expensive input rate. Setting them equal and solving:

Text
S/T  =  (expensive_input - cheap_input) / (cheap_output + expensive_input)with the rates above:S/T  =  (5.00 - 1.00) / (5.00 + 5.00)  =  0.40

The summary must be under 40 percent of the original length to be worth doing. At 800 tokens from 20,000 the ratio is 0.04, comfortably inside. But if your "summary" of a 5,000-token document comes back at 2,500 tokens, the two-stage pipeline is losing money and adding latency. Measure the ratio you actually achieve; do not assume it.

Hierarchical summarisation for very long inputs

When the input is far too large for one summarisation pass, summarise in levels: chunk the source, summarise each chunk with the cheap model, then summarise the summaries, then hand the result to the expensive model. For a 200,000-token archive split into 20 chunks producing 400-token summaries each:

Text
level 1 (small model, 20 chunks)  input  : 200,000 x 1.00/1e6 = 0.200000  output :   8,000 x 5.00/1e6 = 0.040000   ->  0.240000level 2 (small model, summarise the 8,000 tokens of summaries)  input  :   8,000 x 1.00/1e6 = 0.008000  output :   1,000 x 5.00/1e6 = 0.005000   ->  0.013000final   (frontier model, 1,000-token digest + 150 instruction)  input  :   1,150 x  5.00/1e6 = 0.005750  output :     500 x 25.00/1e6 = 0.012500  ->  0.018250total                                          0.271250 USDdirect on the frontier model:  200,000 x 5.00/1e6 + 500 x 25.00/1e6      =  1.012500 USD

A 73.2 percent reduction. The price you pay is information loss at every level — a detail dropped in level 1 cannot be recovered later — so this is right for "what happened overall" questions and wrong for "find the one clause that mentions indemnity" questions. For the latter, retrieval is the tool.

Technique 5 — send only what is relevant

Retrieval pipelines commonly fetch the top 20 chunks by vector similarity and paste all of them into the prompt. Twenty chunks at 400 tokens is 8,000 tokens — and in the itemisation at the top of this lesson, that single decision was 63.6 percent of the entire input bill.

Two cheap changes fix most of it:

  • Rerank, then truncate. Vector similarity is a coarse first pass. Run the 20 candidates through a reranker — a cross-encoder model, or a cheap LLM scoring pass — and keep the top 3. That is 1,200 tokens instead of 8,000.
  • Use a relevance threshold, not a fixed k. If only one chunk scores above the threshold, send one. Fixed k=20 means you always pay for twenty chunks even when the answer is entirely inside the first.

Dropping 6,800 tokens per request across 900,000 requests at 3.00 per million saves 18,360 dollars a month. Accuracy typically improves at the same time, because seventeen irrelevant passages were competing for the model's attention with the three that mattered.

Measuring compression honestly

Three numbers, and a rule about reporting them.

Text
compression ratio  = original_tokens / compressed_tokenstoken reduction    = (1 - compressed_tokens / original_tokens) x 100 %monthly saving     = (original - compressed) x requests x rate / 1,000,000
Python
def compression_report(original_tokens, compressed_tokens,                       requests_per_month, input_rate_per_m,                       accuracy_before, accuracy_after):    saved = original_tokens - compressed_tokens    monthly = saved * requests_per_month * input_rate_per_m / 1_000_000    return {        "ratio": round(original_tokens / compressed_tokens, 2),        "reduction_pct": round(100 * saved / original_tokens, 1),        "monthly_saving_usd": round(monthly, 2),        "accuracy_delta_points": round(            100 * (accuracy_after - accuracy_before), 1),    }print(compression_report(12_580, 3_584, 900_000, 3.00, 0.883, 0.891))# {'ratio': 3.51, 'reduction_pct': 71.5,#  'monthly_saving_usd': 24289.2, 'accuracy_delta_points': 0.8}

Never report a compression number without the quality number beside it. "We cut tokens by 70 percent" is not a result. "We cut tokens by 70 percent and accuracy moved from 88.3 to 89.1" is a result.

Where compression goes wrong

MistakeSymptomWhy it happensFix
Compressing the wrong componentWeeks of work, single-digit savingOptimised the visible part, not the expensive partItemise by tokens × frequency first
Cutting a load-bearing instructionQuality drops silently, weeks laterNo evaluation set gating the changeRe-run evals on every prompt edit
Conditional prompt inside the cached prefixCache hit rate collapses; bill risesPrefix no longer matches byte for byteInject conditional blocks after the breakpoint
Over-abbreviating dataConfidently wrong answersRemoved headers, units, or currencyKeep the legend; stop at CSV with headers
Compressing input when output dominatesBig effort, small effectNever computed the input/output splitDivide input cost by total cost before starting
Summariser output too longTwo-stage pipeline costs more than directNo length cap on the summaryCheck the ratio against the break-even formula
Measuring in charactersReported saving does not appear on the invoiceCharacters are not tokens, especially for JSONMeasure with the provider's token counter
Compressing once and walking awayPrompt bloats back within two quartersNo ownership after the project endsAssert a token ceiling in CI

What a full pass actually looks like

Apply all five techniques to the workload we itemised and the numbers land like this — same evaluation set, accuracy from 88.3 to 89.1 percent.

ComponentBeforeAfterTechnique
Retrieved context8,0001,200Rerank to top 3
Few-shot examples1,660700Prune 8 shots to 2
System prompt1,400444Trim, plus conditional policy blocks
Tool schemas900620Terser descriptions; drop two unused tools
Conversation history500500Unchanged
User message120120Unchanged
Total input tokens12,5803,58471.5% reduction
Monthly input cost33,966 USD9,676.80 USD24,289.20 USD saved

That is 291,470 dollars a year, and not one line of model-serving code changed. Every technique was applied to text that the team wrote themselves and had stopped looking at.

Two habits keep it from growing back. First, put a token ceiling in your test suite — a test that counts the assembled prompt for a fixed fixture and fails if it exceeds an agreed number. Prompts bloat the way logging bloats: one reasonable addition at a time, none of which anyone would object to individually. Second, whenever someone proposes adding to a prompt, make them state which component it lands in and what that component's monthly cost becomes. A 200-token addition to a system prompt on this workload is a 540-dollar-a-month decision, and it should be made as one.