Context Management and Memory

Context Windows and Token Limits


A refunds bot at a mid-sized retailer goes into production on a Tuesday. It works beautifully for a week. Then a support lead files a bug: on long conversations, the bot starts inventing refund amounts.

The transcript is damning. At turn 12 the bot correctly quotes a refund of 84.50 for a returned pair of boots. At turn 34, the customer asks "so remind me what the refund comes to?" and the bot answers "your refund is 129.00" — a number that appears nowhere in the conversation, nowhere in the order, nowhere at all.

The application logs explain it. The request at turn 34 was 203,880 tokens against a 200,000-token limit. A helper function in the middleware caught the overflow and did the obvious thing: dropped messages from the front of the list until it fit. It dropped 19 messages, including turn 12. The model was then asked to recall a figure that had been deleted from its input three milliseconds earlier, and it did what language models do when asked a question they cannot answer from evidence — it produced a plausible-looking number.

Nothing crashed. No exception was raised. No alert fired. The only symptom was a wrong answer. This is the defining property of context-window failures: they are silent, and they look like hallucination. If you cannot do the token arithmetic for your own application, you cannot tell the difference between a model that is confused and a model that was never shown the answer.

A 16k budget with every claimant namedSystem andrules — 800Toolschemas — 1,200Retrieveddocuments — 4,000Conversationhistory — 6,000Reserved forthe reply — 2,000topbottomThat leaves about 2,000 tokens of slack; the history is the only block that grows on its own.
The window is a budget, not a container — decide each claimant's share before the conversation, or the history will silently evict the rules.

The window is not "the conversation"

The single most common mental model — "the context window is how much chat history the model remembers" — is wrong in a way that causes real outages. A language model is stateless. It has no memory between calls. On every single request you resend the entire input, and the context window is the ceiling on everything you send plus everything it may generate.

That total includes, at minimum:

  • The system prompt — persona, policy, formatting rules, safety instructions. Sent in full on every call.
  • Tool and function definitions — the JSON schema for every tool you expose. Also sent in full on every call, whether or not any tool is used.
  • Few-shot examples — demonstration input/output pairs baked into the prompt.
  • Retrieved documents — chunks pulled from a search index or vector store for this particular turn.
  • Conversation history — every prior user and assistant message you choose to include.
  • The current user message.
  • Reserved output space — the tokens the model will generate. Most APIs make you declare a ceiling (max_tokens), and that reservation is subtracted from what is left for input.

The context window is a fixed-size budget shared by your instructions, your tools, your retrieved evidence, your chat history and the reply. History is only one line item, and it is usually not the biggest one.

Tool definitions are the line item that catches people out. A well-documented tool schema — a description, six parameters, each with its own description and enum values — runs 400 to 900 tokens. Eleven tools is comfortably 6,000-plus tokens, spent on every request in the conversation, forever, even on the turn where the user says "thanks".

What a token actually is

A token is not a word and not a character. Modern models use subword tokenisation — usually a variant of byte-pair encoding (BPE), which is a compression scheme learned from a large text corpus. The algorithm starts with individual bytes and repeatedly merges the most frequent adjacent pair into a new symbol, some tens of thousands of times. The result is a vocabulary where common words are one token and rare words are assembled from fragments.

So a frequent word like the is a single token. A rarer one like unbelievable splits into something like un + believ + able. A truly unusual string — a UUID such as 8f14e45f-ceea-467a-9c1e-3b2a5c7d0e91 — has no learned merges at all and shatters into twenty-something tokens for thirty-six characters.

The reason this design won is that it never fails. A word-level vocabulary hits an unknown word and has nothing to emit. A character-level vocabulary handles everything but makes sequences four to five times longer, which is catastrophic when attention cost grows with the square of sequence length. Subwords sit in the middle: unlimited coverage, sequences roughly a quarter the length of the raw characters.

The rules of thumb, and where they break

For ordinary English prose, two approximations get you close:

  • Roughly 4 characters per token.
  • Roughly 0.75 words per token — so 1,000 words is about 1,330 tokens.

Both are calibrated on English prose and both fall apart the moment your content is not English prose.

Content typeApprox. chars per tokenWhy
English prose~4.0What the tokeniser was optimised for
Code~3.0Indentation, brackets, underscores and operators fragment
JSON payloads~2.5Braces, quotes and colons each cost a token; keys repeat on every record
Numbers, IDs, hashes~2.0 or worseNo frequent merges exist for arbitrary digit strings
Non-Latin scripts~1.0 or worseUnder-represented in the merge corpus; often one token per character or per byte
Base64 / encoded blobs~2.0Effectively random; the worst case for any tokeniser

The practical consequence is sharp. Take the same fact expressed two ways:

Text
JSON:  {"customer": "Alice Chen", "plan": "pro", "renews": "2026-03-14"}       67 characters, roughly 26 tokensProse: Alice Chen is on the pro plan, renewing on 2026-03-14.       53 characters, roughly 14 tokens

Same information, and the JSON costs close to twice as much. Multiply by 200 retrieved records per request and the formatting decision alone is the difference between fitting and not fitting. If you are stuffing structured data into a prompt for the model to read rather than to parse programmatically, flatten it to prose or to a compact delimited form.

Tokenisers are not interchangeable

Every model family ships its own tokeniser with its own vocabulary and its own merge table. The same paragraph counted with OpenAI's o200k_base, Anthropic's tokeniser and a Llama tokeniser will give three different numbers — typically within 10 to 35 percent of each other, occasionally further apart on code or non-English text. Tokenisers also change within a family: Anthropic's documentation says Claude Opus 4.7 and later models produce roughly 30 percent more tokens for the same text than earlier Claude models, so a budget measured on last year's model is wrong for this year's.

Counting tokens for one model's tokeniser and using the result to budget for a different model is a bug, not an approximation. Build the margin in, or count with the right tool.

Counting tokens for real

There are three counting strategies, and a serious system uses all three.

1. Count before you send

For OpenAI-family models, the tokeniser is published, so you can count locally and offline. Ask tiktoken for the encoding by model name rather than hard-coding one: current models (GPT-4o, GPT-4.1, GPT-5, the o-series) use o200k_base, while the older cl100k_base belongs to GPT-4 and GPT-3.5.

Python
import tiktokenenc = tiktoken.encoding_for_model("gpt-4o")   # resolves to o200k_basedef count(text: str) -> int:    return len(enc.encode(text))print(count("The quick brown fox jumps over the lazy dog."))print(count("8f14e45f-ceea-467a-9c1e-3b2a5c7d0e91"))

Note the trap: this counts the text, not the request. Chat APIs wrap every message in role markers and delimiters, so a conversation costs several tokens more per message than the sum of its message bodies. If you budget from raw text length you will under-count by a few tokens per message, which is invisible at turn 3 and is 200 tokens at turn 50.

For Claude models the tokeniser is not published, and tiktoken will simply give you a wrong number. Use the token-counting endpoint, which counts the whole request — system prompt, tools, messages and all — server-side. It is free (it has its own rate limit), and the documentation describes the result as a close estimate of what the real call will use:

Python
from anthropic import Anthropicclient = Anthropic()result = client.messages.count_tokens(    model="claude-opus-5",    system=SYSTEM_PROMPT,    tools=TOOL_DEFINITIONS,    messages=conversation,)print(result.input_tokens)

This is the number that matters, because it includes the tool schemas and the message scaffolding that local counting silently omits.

2. Read the actual usage off the response

Every response carries a usage block. Trust it over your own estimate — it is the number you are billed on.

Python
response = client.messages.create(    model="claude-opus-5",    max_tokens=4096,    system=SYSTEM_PROMPT,    messages=conversation,)u = response.usageprint(u.input_tokens, u.output_tokens)print(u.cache_read_input_tokens, u.cache_creation_input_tokens)

Log these on every call. The gap between your pre-count estimate and the reported input_tokens is the size of your blind spot, and you want it on a dashboard rather than in a post-mortem.

3. Track cumulative growth across the conversation

The single most useful production metric is not the size of any one request; it is the slope. Record input tokens per turn and you can see the conversation heading for the wall several turns before it arrives, which is enough warning to summarise or evict on purpose rather than by accident.

Python
class ContextMeter:    def __init__(self, window: int, reserved_output: int, margin: float = 0.05):        self.usable = int(window * (1 - margin)) - reserved_output        self.history = []    def record(self, input_tokens: int) -> dict:        self.history.append(input_tokens)        headroom = self.usable - input_tokens        growth = 0        if len(self.history) >= 3:            growth = (self.history[-1] - self.history[-3]) / 2        turns_left = headroom // growth if growth > 0 else None        return {"headroom": headroom, "growth_per_turn": growth,                "turns_before_overflow": turns_left}

When turns_before_overflow drops below about five, you act. Acting early is cheap; acting at the boundary means dropping content you needed.

Building a context budget

A budget is an explicit allocation of the window across line items, decided before the first request rather than discovered during an incident. Here is one for an agent with tools and retrieval, running on a 200,000-token window.

Line itemTokensFixed or variable?
Reserved for the reply (max_tokens)8,000Fixed reservation
System prompt and policy2,400Fixed, every call
Tool definitions (11 tools)6,900Fixed, every call
Few-shot examples3,100Fixed, every call
Retrieved chunks (8 × 700)5,600Variable, per turn
Safety margin (5% of window)10,000Fixed
Available for conversation history164,000

The arithmetic: 8,000 + 2,400 + 6,900 + 3,100 + 5,600 + 10,000 = 36,000, and 200,000 − 36,000 = 164,000. At an average of 500 tokens per exchange (user message plus reply), that is 328 turns of history before anything must be evicted. Comfortable.

Now run exactly the same application against a 32,000-token window — a very common size for smaller and self-hosted models — with the output reservation and margin scaled down proportionally:

Line item200K window32K window
Reserved for the reply8,0002,000
System prompt2,4002,400
Tool definitions6,9006,900
Few-shot examples3,1003,100
Retrieved chunks5,6005,600
Safety margin10,0001,600
Fixed overhead total36,00021,600
Left for history164,00010,400
Turns before eviction (at 500 tok/turn)32820

Read the last row carefully. The fixed overhead did not shrink — the system prompt, the tools and the few-shot examples are the same 12,400 tokens in both columns, but on the 32K model they consume 39% of the entire window before a single word of conversation. Turn 21 begins evicting, and this is precisely where the refunds bot failed.

On a small window, your fixed prompt overhead is the enemy. Cutting three tools and two few-shot examples can buy you more usable history than any clever summarisation strategy.

What actually happens when you exceed the window

There are three distinct failure modes, and they are frequently confused with each other.

Failure modeSymptomCauseFix
Hard rejectionHTTP 400, "prompt is too long"The input alone exceeds the window (on older models and some providers, input plus max_tokens is enough)Pre-count and trim before sending; treat 400 as a bug in your budgeting, not a retryable error
Silent evictionModel "forgets" facts stated earlier; invents plausible replacementsYour own middleware dropped messages to make the request fitLog what you evict; keep pinned facts outside the evictable history
Output truncationReply stops mid-sentence; stop_reason is max_tokens, or model_context_window_exceeded when the reply ran into the window itselfReserved output space was too small, or the input left too little roomRaise max_tokens or trim the input; stream and continue
Degraded recallFacts are present in the input but the model misses them"Lost in the middle" — attention favours the start and end of very long inputsPut critical facts near the top or bottom; retrieve less, retrieve better

The first and third rows changed recently on Claude. Since the Claude 4.5 models, a request whose input fits but whose input plus max_tokens does not is accepted; if the reply then reaches the edge of the window, it stops with stop_reason: "model_context_window_exceeded". Older models, and some other providers, reject that request with a 400 instead. Either way, handle both stop reasons in code.

The fourth row is the one that surprises people. Fitting inside the window is necessary but not sufficient. Liu et al. named the effect "lost in the middle" in 2023: information buried in the middle of a long input was recalled less reliably than the same information at the beginning or the end. Anthropic's documentation calls the broader pattern context rot: as the token count grows, accuracy and recall degrade. Windows of 1,000,000 tokens are now standard on large current models, but that means you can send a million tokens, not that you should. Filling a huge window with marginally relevant retrieved text reliably makes answers worse and always makes them slower and more expensive.

The cost arithmetic nobody runs first

Because you resend the whole history on every call, the cost of a naive conversation grows with the square of its length. Let FF be the fixed prefix (system prompt, tools, examples) and dd the tokens added per exchange. Over NN turns, total input tokens billed are:

total input=N⋅F+d⋅N(N−1)2\text{total input} = N \cdot F + d \cdot \frac{N(N-1)}{2}

Put real numbers in. Take F=1,500F = 1{,}500, d=500d = 500, and a 200-turn conversation — a single long support session:

  • Fixed part: 200×1,500=300,000200 \times 1{,}500 = 300{,}000 tokens.
  • Growing part: 500×200×1992=500×19,900=9,950,000500 \times \frac{200 \times 199}{2} = 500 \times 19{,}900 = 9{,}950{,}000 tokens.
  • Total: 10,250,000 input tokens for one conversation.

At an input price of 3 dollars per million tokens, that is 30.75 dollars for a single conversation, of which the user only ever typed about 50,000 tokens. Now cap the history at the last 20 exchanges — 10,000 tokens plus the 1,500-token prefix, so 11,500 per request from turn 21 onward:

  • Turns 1–20 (still growing): 20×1,500+500×20×192=30,000+95,000=125,00020 \times 1{,}500 + 500 \times \frac{20 \times 19}{2} = 30{,}000 + 95{,}000 = 125{,}000 tokens.
  • Turns 21–200 (capped): 180×11,500=2,070,000180 \times 11{,}500 = 2{,}070{,}000 tokens.
  • Total: 2,195,000 tokens, or 6.59 dollars.

A 79% cost reduction from one bounded list. And the same change caps latency, because time to first token scales with input size.

Caching the fixed prefix

The fixed 1,500 tokens were resent 200 times: 300,000 tokens, or 90 cents at 3 dollars per million. Prompt caching lets the provider store that prefix and charge roughly a tenth of the input price to read it back. The same 300,000 tokens then cost about 9 cents. On the 6,900-token tool-definition block across 200 turns — 1,380,000 tokens, 4.14 dollars uncached — caching brings it to roughly 41 cents.

Three details decide whether you get that saving on Anthropic's API. Caching is opt-in: you mark the prefix with cache_control, or set it once at the top level of the request. Writing the cache costs more than a normal read, 1.25 times the input price for the default five-minute lifetime, so it pays off from the second request onwards. And there is a minimum cacheable length that depends on the model: 512 tokens on Claude Opus 5, 4,096 on Claude Haiku 4.5. A prefix below the minimum silently does not cache, so the 1,500-token prefix above caches on one model and not on the other. Other providers trigger and price caching differently, so check yours.

Caching is a prefix match, so it only works if the prefix is byte-identical every time. A timestamp in the system prompt, a request ID, or a tool list built from an unordered dictionary will invalidate the cache on every call and you will pay full price while believing you are caching. The check is direct: if cache_read_input_tokens is zero across repeated requests, something in your prefix is changing.

Where the limit bites hardest

ApplicationPressure comes fromWhat to watch
Long support conversationsLinear history growth over 50–300 turnsCumulative input tokens per session
Document Q&AOne document can exceed the whole windowChunk size × number of retrieved chunks
Coding agentsFile contents and tool output dominateTool result size; truncate long outputs at the source
Multi-tool agentsTool schemas are a fixed tax on every callNumber of tools × average schema size
RAG over many sourcesRetrieving 20 chunks "to be safe"Precision of retrieval, not just recall

Designing a budget you can defend

When you build something real, the token budget is a design artefact, written down before the first request, not a number you discover during an incident. Four commitments make it defensible.

Write the allocation down and enforce it in code. A dictionary of line items with token ceilings, checked before every send. If retrieval wants 9,000 tokens and its allocation is 5,600, retrieval gets truncated — not the conversation history, and not silently.

Reserve output space explicitly, and subtract it. Depending on the model, an input that fits perfectly plus a max_tokens that pushes the total over either fails with a 400 or produces a reply cut off at the edge of the window. Neither is what you want. Your usable input budget is window − max_tokens − margin, never window.

Make eviction an explicit, logged decision. The refunds bot failed because a utility function silently deleted messages. Whatever strategy you choose, emit a structured log line every time content leaves the context: what was dropped, how many tokens, which turn. When someone reports that the bot forgot something, that log is the difference between a five-minute diagnosis and a week of guessing.

Separate facts from transcript. The refund amount of 84.50 should never have lived only inside a chat message that eviction could reach. Extracted facts — order numbers, quoted amounts, stated preferences, decisions — belong in a small pinned block that is rebuilt into every prompt and is not part of the evictable history. A pinned block of 300 tokens costing 0.09 cents per turn is a trivial price for never inventing a refund figure again.

Do the arithmetic once, on paper, with your real system prompt and your real tool schemas measured rather than guessed. Nearly every context failure in production traces back to a budget that was never written down.