Context Management and Memory

Sliding Window and Summarization Methods


A nutrition coaching bot ships with the simplest memory anyone reaches for first: keep the last ten messages. It is one line of configuration and it never overflows the context window.

On the second message of the conversation, a user writes: "I'm vegetarian, and I'm allergic to peanuts." The bot acknowledges it. Twenty-eight messages later, deep into a discussion about high-protein lunches, the bot enthusiastically recommends a chicken satay bowl.

The engineer on call checks the request payload. It is perfectly valid, well under the token limit, no errors anywhere. The allergy message is simply not in it. It fell out of the window at message 12 and the bot has had no idea since. Every reply after that point was a confident answer produced from a truncated record.

This is the central tension of short-term context management. You cannot keep everything — history grows without bound, and you resend all of it on every call. But whatever you throw away, you throw away permanently, and the model will never tell you it is missing. The whole discipline is about choosing what to discard with more care than "the oldest thing".

Two ways to stop a 200-exchange chat overflowingLast ten messages• One line ofconfiguration, never overflows• Message count ignores token size• One pasted log can blow the window• Forgets the allergy stated in message 3Summary plus verbatim tail• Old turns compressed to about 400 tokens• Recent eight turns kept word for word• Costs one extra call per compaction• Detail is lost gradually, not at a cliff
A window forgets by position and a summary forgets by importance — which is why the important facts belong in neither.

Why you cannot just keep everything

Start with the naive approach so its failure is concrete. Buffer every message and resend the lot. Suppose an average exchange — one user message plus one assistant reply — is 480 tokens, and your fixed prefix (system prompt, tool schemas, examples) is 3,000 tokens.

Request size at turn NN is 3,000+480N3{,}000 + 480N. On a 128,000-token window with 4,000 reserved for the reply and a 5% margin, your usable input is about 117,600 tokens, so you overflow at:

N=117,600−3,000480≈238 exchangesN = \frac{117{,}600 - 3{,}000}{480} \approx 238 \text{ exchanges}

That sounds fine until you price it. Total input tokens billed across those 238 turns is

238×3,000+480×238×2372=714,000+13,537,440≈14.25M tokens238 \times 3{,}000 + 480 \times \frac{238 \times 237}{2} = 714{,}000 + 13{,}537{,}440 \approx 14.25\text{M tokens}

At 3 dollars per million input tokens, that is roughly 42.75 dollars for one conversation, and the last request alone takes several seconds just to read. You do not run out of window first — you run out of money and patience first.

Fitting inside the window is a hard constraint. Staying well below it is an economic and quality decision, and you should make it deliberately rather than letting the limit decide for you.

So set a budget by choice. For the rest of this discussion, allocate 12,000 tokens to conversation history — enough for genuine continuity, small enough that requests stay fast and cheap. At 480 tokens per exchange, that holds about 25 exchanges.

Sliding windows

A sliding window keeps the most recent kk units of conversation and drops everything older. It rests on one assumption: recency correlates with relevance. In most dialogue that assumption is broadly right — the answer to "what about the second option?" is nearly always three messages back, not three hundred.

The message-count trap

The obvious implementation counts messages: keep the last 10. It is also the version that broke the nutrition bot, and it has a second, subtler problem: messages are not a unit of size.

Conversation shape10 messages =Result
Short chat turns (~240 tokens each)~2,400 tokensWasting 9,600 tokens of a 12,000 budget
Normal mixed conversation~4,800 tokensRoughly what you intended
One turn contains a 9,000-token tool result~11,400 tokensNearly the entire budget in one message
Two large tool results~20,000 tokensOverflow, despite "only 10 messages"

A fivefold variance in the size of "10 messages" means a message-count window either wastes most of your budget or blows through it, depending on what the user happened to ask. Count tokens, not messages.

A token-based sliding window, done properly

The logic is: walk backwards from the newest message, accumulate token counts, and stop when adding the next message would exceed the allocation. Simple — but there are four ways to get it wrong that produce API errors or silent corruption.

Python
def trim_to_budget(messages, budget_tokens, count_fn):    """Keep the most recent messages that fit in budget_tokens.    messages: list of dicts with 'role' and 'content'    count_fn: returns token count for one message    """    kept, total = [], 0    for msg in reversed(messages):        cost = count_fn(msg)        if total + cost > budget_tokens:            break        kept.append(msg)        total += cost    kept.reverse()    # Trap 1: a tool_result must never be orphaned from its tool_use.    # If the first kept message is a tool result, drop it.    while kept and _is_orphan_tool_result(kept[0], kept):        kept.pop(0)    # Trap 2: most chat APIs require the first message to be from the user.    while kept and kept[0]["role"] != "user":        kept.pop(0)    return kept, total

The four traps, stated plainly:

  • Orphaned tool results. If you cut between an assistant message containing a tool_use block and the user message containing its tool_result, the API returns a 400. Trim at exchange boundaries, not message boundaries.
  • Leading assistant message. Many APIs require the conversation to begin with a user turn. Trimming can leave an assistant message first.
  • Trimming the system prompt. It is not part of history and must never enter the eviction list. Keep it in a separate variable so it is structurally impossible to drop.
  • Cache invalidation. This one costs real money and is almost never noticed.

Prompt caching is a prefix match. If you evict one message from the front on every single turn, the prefix changes on every single turn and your cache hit rate is zero — you pay full price while believing you are caching.

The fix is to trim in chunks, not continuously. Rather than evicting whenever you exceed 12,000 tokens, let history grow to 12,000 and then cut back to 6,000 in one operation. The prefix then stays byte-identical for roughly a dozen turns, cache reads land at about a tenth of the input price, and you pay the invalidation cost once instead of every turn.

What a sliding window is good and bad at

Sliding window
Cost per turnZero extra API calls — pure bookkeeping
Latency addedEffectively zero (milliseconds of counting)
Fidelity of what it keepsPerfect — kept messages are verbatim
Fidelity of what it dropsZero — gone completely, no trace
Fails whenImportant facts were stated early: names, constraints, allergies, order numbers, decisions
Best forTask-focused, short-horizon chat where old turns genuinely stop mattering

Summarisation

Summarisation replaces old messages with a compressed description of them. Instead of deleting turns 1 to 30, you spend one extra model call to turn them into a paragraph and keep the paragraph.

The arithmetic is the appeal. Thirty exchanges at 480 tokens each is 14,400 tokens. A good summary of them runs 600 tokens. That is a 24:1 compression ratio, and it means those thirty exchanges now cost 600 tokens on every subsequent request instead of 14,400.

The three shapes of summarisation

TypeHow it worksCost per compactionWeakness
Rolling / recursiveNew summary = summarise(old summary + new messages)One call over a small inputDrift — each generation is a summary of a summary; errors compound
Map-reduceSummarise each block of raw messages independently, then combine the block summariesN parallel calls plus one combineMore expensive; blocks lose cross-block context
HierarchicalLevel 1 summarises raw blocks; level 2 summarises level-1 summaries; and so onAmortised — most turns cost nothingMore machinery; older levels are very lossy

Writing the summarisation prompt is the whole job

This is where summarisation quietly fails, and it is worth being blunt about it. The default instruction people write is:

Text
Summarise the conversation so far.

A general-purpose summariser optimises for readability, and readable summaries drop exactly what a coaching bot or a support bot needs: numbers, identifiers, explicit constraints. "The user discussed their dietary requirements and asked about lunch options" is a fine English summary and a catastrophic memory record — it has thrown away "vegetarian" and "peanut allergy", the only two facts that mattered.

Constrain the output instead:

Python
SUMMARY_PROMPT = """You are compressing a conversation into a memory record.Preserve verbatim, in every case:- Names, IDs, order numbers, account numbers, dates, amounts- Stated constraints, preferences, allergies, restrictions- Decisions made and commitments given by the assistant- Questions the user asked that are still unansweredDiscard: pleasantries, restatements, the assistant's own explanations,anything already superseded by a later message.If a fact was later corrected, record only the corrected version andnote that it was a correction.Output under 600 tokens using these headings:FACTS / CONSTRAINTS / DECISIONS / OPENPREVIOUS SUMMARY:{previous_summary}NEW MESSAGES:{new_messages}"""

Compare what the two prompts produce from the same source material:

Generic summaryConstrained summary
"The user discussed dietary requirements and asked about high-protein lunch options. The assistant offered several suggestions.""FACTS: user is vegetarian; peanut allergy (severe). Target 120g protein/day. CONSTRAINTS: no meat, no peanuts or peanut oil. DECISIONS: assistant agreed to suggest only vegetarian, peanut-free meals. OPEN: user asked about protein powder brands — not yet answered."
18 tokens. Useless.~70 tokens. Actually a memory.

What summarisation costs

Two costs, and people budget for neither.

Money. Compressing 14,400 tokens into 600 costs one call: 14,400 input tokens at 3 dollars per million is 4.32 cents, plus 600 output tokens at 15 dollars per million is 0.9 cents. About 5.2 cents per compaction. Cheap, but not free, and it is charged on top of the conversation itself.

Latency. The compaction call takes two to five seconds. If you run it synchronously on the turn that triggers it, one unlucky user waits five extra seconds for a reply and has no idea why. Run compaction on the previous turn's boundary, in the background, or on a cheaper and faster model — summarisation is a much easier task than the conversation itself, and a small model does it well.

The hybrid: summary plus verbatim tail

Neither pure strategy is what production systems use. The standard design keeps both:

Text
[ system prompt          ]  fixed, cached[ pinned facts           ]  ~300 tokens, never evicted[ running summary        ]  ~600 tokens, covers turns 1..M[ verbatim recent turns  ]  ~4,800-9,600 tokens, turns M+1..N[ current user message   ]

The reasoning is that the two halves fail in opposite directions. Recent turns need exact wording — the user says "the second one", and only the verbatim text tells you what that refers to. Old turns need only their durable content, and compression is nearly free there.

Python
class HybridMemory:    def __init__(self, keep_min=10, keep_max=20, summariser=None, count_fn=None):        self.keep_min = keep_min      # exchanges kept verbatim after compaction        self.keep_max = keep_max      # trigger point        self.summary = ""        self.exchanges = []           # list of (user_msg, assistant_msg)        self.summarise = summariser        self.count = count_fn    def add(self, user_msg, assistant_msg):        self.exchanges.append((user_msg, assistant_msg))        if len(self.exchanges) > self.keep_max:            self._compact()    def _compact(self):        n = len(self.exchanges) - self.keep_min        old, self.exchanges = self.exchanges[:n], self.exchanges[n:]        self.summary = self.summarise(self.summary, old)    def build(self):        blocks = []        if self.summary:            blocks.append({"role": "user",                           "content": f"[Earlier conversation]\n{self.summary}"})            blocks.append({"role": "assistant",                           "content": "Understood, I have that context."})        for u, a in self.exchanges:            blocks.append(u)            blocks.append(a)        return blocks

Note the chunked compaction: it triggers at 20 exchanges and cuts back to 10, so it fires once every ten turns rather than every turn. That is the cache-friendly shape.

The numbers for a 200-exchange conversation

Same parameters as before: fixed prefix 3,000 tokens, 480 tokens per exchange, 3 dollars per million input, 15 per million output.

StrategyInput tokens billedExtra callsTotal costFacts from turn 3 survive?
Full buffer10,152,000030.46 dollarsYes — until it overflows
Sliding window (25 exchanges)2,844,00008.53 dollarsNo
Hybrid (summary + 10–20 verbatim)2,180,00018 compactions7.05 dollarsYes, in compressed form

The hybrid arithmetic, worked: the request averages 3,000 fixed + 700 summary + about 7,200 verbatim (oscillating between 10 and 20 exchanges) = 10,900 tokens. Across 200 turns that is 2,180,000 tokens, or 6.54 dollars. Compaction fires 18 times (at exchange 21, then every tenth exchange), each summarising 11 exchanges (5,280 tokens) plus the previous summary (700) into a new 700-token summary: (5,980×3+700×15)/106=0.0284(5{,}980 \times 3 + 700 \times 15) / 10^6 = 0.0284 dollars each, so 0.51 dollars total. Sum: 7.05 dollars, a 77% reduction against the full buffer — and unlike the sliding window, the vegetarian and the peanut allergy are still there.

Failure modes worth naming

FailureWhat you seeRoot causeFix
Early-fact amnesiaBot violates a constraint stated at turn 2Pure recency windowPin extracted facts outside the evictable history
Summary driftAfter many compactions, details are subtly wrong or inventedRecursively summarising summaries; each pass is lossy and each pass can hallucinateKeep raw messages in a store and re-summarise from originals; or hold facts in an append-only list rather than free text
Number lossAmounts, IDs and dates disappearUnconstrained summarisation promptExplicit preserve-verbatim instruction and a fixed output schema
Stale correctionsBot uses a value the user already correctedSummary appended both the original and the correctionInstruct the summariser to record only the corrected value
Cache thrashCosts far above the estimate; cache_read_input_tokens is zeroEvicting one message per turn changes the prefix every turnCompact in chunks, not continuously
Latency spikeOne turn in ten takes five seconds longerSynchronous compaction on the triggering turnCompact asynchronously, or on a smaller and faster model
Overflow despite a window400 error with a message-count window in placeA single tool result was 9,000 tokensBudget in tokens; truncate tool output at the source

Frameworks and managed options

Older LangChain code used memory classes such as ConversationBufferWindowMemory, ConversationTokenBufferMemory, ConversationSummaryMemory and ConversationSummaryBufferMemory. They map exactly onto the strategies above: message-count window, token-count window, pure rolling summary, and hybrid. In LangChain 1.x they survive only in the langchain-classic compatibility package. Current code keeps the thread's messages in a LangGraph checkpointer (for example InMemorySaver in development, a Postgres saver in production), and shapes what the model sees with middleware on create_agent: a before_model hook that trims, or the built-in SummarizationMiddleware, which summarises once a token or message threshold is reached and keeps the most recent messages verbatim. That is the hybrid above, as configuration. The change is a good one — the class-based versions hid the eviction decision, and hidden eviction is precisely how the chicken satay reached a vegetarian.

Model providers now offer managed versions too. The Claude API has server-side compaction (in beta): the API writes the summary of older turns for you and hands back a compaction block that replaces them on the next request. It also has context editing, which clears old tool results or thinking blocks by rule instead of summarising them — useful for agents whose context fills up with tool output rather than chat. Both save you writing the summariser; neither removes the need to pin facts that must never be lost, or the need to log what was compacted.

Whatever you use, the requirement is the same: you must be able to answer "what is currently in the context, and what was dropped?" from your logs. If the abstraction will not tell you that, replace it.

Choosing a strategy for a real system

Work through it in this order.

Ask how long a session actually runs. If your median conversation is six turns, a token window is the entire answer and summarisation is machinery you do not need. Measure the distribution before building anything; a great many chatbots that shipped elaborate summarisation had a 95th-percentile session of nine turns.

Identify what must never be lost, and take it out of the transcript. Before choosing between window and summary, write the list: order number, stated allergies, budget ceiling, the plan the user agreed to. Extract those into a pinned block of a few hundred tokens that is rebuilt into every prompt and is not part of the evictable history. This single step fixes most of the failures above, and it costs about 0.09 cents per turn. Doing it means your window or summary strategy only has to handle conversational flow, which is a much easier problem.

Then choose by session length. Under 20 exchanges: token window, nothing else. Twenty to a hundred: hybrid, compacting in chunks. Over a hundred, or across sessions: hybrid plus a persistent store, because at that point the raw transcript belongs in a database rather than in a variable.

Log every eviction. One structured line per compaction — turn number, messages dropped, tokens before and after, a hash of the resulting summary. When a user reports that the bot forgot something, you want to answer in thirty seconds rather than reconstruct the session from scratch. The nutrition bot's incident took two days to diagnose, and every minute of it was spent establishing something a single log line would have said.