Course Content
Context Management and Memory
3 sections · 6 lessons
Sliding Window and Summarization Methods
A nutrition coaching bot ships with the simplest memory anyone reaches for first: keep the last ten messages. It is one line of configuration and it never overflows the context window.
On the second message of the conversation, a user writes: "I'm vegetarian, and I'm allergic to peanuts." The bot acknowledges it. Twenty-eight messages later, deep into a discussion about high-protein lunches, the bot enthusiastically recommends a chicken satay bowl.
The engineer on call checks the request payload. It is perfectly valid, well under the token limit, no errors anywhere. The allergy message is simply not in it. It fell out of the window at message 12 and the bot has had no idea since. Every reply after that point was a confident answer produced from a truncated record.
This is the central tension of short-term context management. You cannot keep everything — history grows without bound, and you resend all of it on every call. But whatever you throw away, you throw away permanently, and the model will never tell you it is missing. The whole discipline is about choosing what to discard with more care than "the oldest thing".
Why you cannot just keep everything
Start with the naive approach so its failure is concrete. Buffer every message and resend the lot. Suppose an average exchange — one user message plus one assistant reply — is 480 tokens, and your fixed prefix (system prompt, tool schemas, examples) is 3,000 tokens.
Request size at turn N is 3,000+480N. On a 128,000-token window with 4,000 reserved for the reply and a 5% margin, your usable input is about 117,600 tokens, so you overflow at:
That sounds fine until you price it. Total input tokens billed across those 238 turns is
At 3 dollars per million input tokens, that is roughly 42.75 dollars for one conversation, and the last request alone takes several seconds just to read. You do not run out of window first — you run out of money and patience first.
Fitting inside the window is a hard constraint. Staying well below it is an economic and quality decision, and you should make it deliberately rather than letting the limit decide for you.
So set a budget by choice. For the rest of this discussion, allocate 12,000 tokens to conversation history — enough for genuine continuity, small enough that requests stay fast and cheap. At 480 tokens per exchange, that holds about 25 exchanges.
Sliding windows
A sliding window keeps the most recent k units of conversation and drops everything older. It rests on one assumption: recency correlates with relevance. In most dialogue that assumption is broadly right — the answer to "what about the second option?" is nearly always three messages back, not three hundred.
The message-count trap
The obvious implementation counts messages: keep the last 10. It is also the version that broke the nutrition bot, and it has a second, subtler problem: messages are not a unit of size.
| Conversation shape | 10 messages = | Result |
|---|---|---|
| Short chat turns (~240 tokens each) | ~2,400 tokens | Wasting 9,600 tokens of a 12,000 budget |
| Normal mixed conversation | ~4,800 tokens | Roughly what you intended |
| One turn contains a 9,000-token tool result | ~11,400 tokens | Nearly the entire budget in one message |
| Two large tool results | ~20,000 tokens | Overflow, despite "only 10 messages" |
A fivefold variance in the size of "10 messages" means a message-count window either wastes most of your budget or blows through it, depending on what the user happened to ask. Count tokens, not messages.
A token-based sliding window, done properly
The logic is: walk backwards from the newest message, accumulate token counts, and stop when adding the next message would exceed the allocation. Simple — but there are four ways to get it wrong that produce API errors or silent corruption.
1def trim_to_budget(messages, budget_tokens, count_fn):2 """Keep the most recent messages that fit in budget_tokens.34 messages: list of dicts with 'role' and 'content'5 count_fn: returns token count for one message6 """7 kept, total = [], 089 for msg in reversed(messages):10 cost = count_fn(msg)11 if total + cost > budget_tokens:12 break13 kept.append(msg)14 total += cost1516 kept.reverse()1718 # Trap 1: a tool_result must never be orphaned from its tool_use.19 # If the first kept message is a tool result, drop it.20 while kept and _is_orphan_tool_result(kept[0], kept):21 kept.pop(0)2223 # Trap 2: most chat APIs require the first message to be from the user.24 while kept and kept[0]["role"] != "user":25 kept.pop(0)2627 return kept, totalThe four traps, stated plainly:
- Orphaned tool results. If you cut between an assistant message containing a
tool_useblock and the user message containing itstool_result, the API returns a 400. Trim at exchange boundaries, not message boundaries. - Leading assistant message. Many APIs require the conversation to begin with a user turn. Trimming can leave an assistant message first.
- Trimming the system prompt. It is not part of history and must never enter the eviction list. Keep it in a separate variable so it is structurally impossible to drop.
- Cache invalidation. This one costs real money and is almost never noticed.
Prompt caching is a prefix match. If you evict one message from the front on every single turn, the prefix changes on every single turn and your cache hit rate is zero — you pay full price while believing you are caching.
The fix is to trim in chunks, not continuously. Rather than evicting whenever you exceed 12,000 tokens, let history grow to 12,000 and then cut back to 6,000 in one operation. The prefix then stays byte-identical for roughly a dozen turns, cache reads land at about a tenth of the input price, and you pay the invalidation cost once instead of every turn.
What a sliding window is good and bad at
| Sliding window | |
|---|---|
| Cost per turn | Zero extra API calls — pure bookkeeping |
| Latency added | Effectively zero (milliseconds of counting) |
| Fidelity of what it keeps | Perfect — kept messages are verbatim |
| Fidelity of what it drops | Zero — gone completely, no trace |
| Fails when | Important facts were stated early: names, constraints, allergies, order numbers, decisions |
| Best for | Task-focused, short-horizon chat where old turns genuinely stop mattering |
Summarisation
Summarisation replaces old messages with a compressed description of them. Instead of deleting turns 1 to 30, you spend one extra model call to turn them into a paragraph and keep the paragraph.
The arithmetic is the appeal. Thirty exchanges at 480 tokens each is 14,400 tokens. A good summary of them runs 600 tokens. That is a 24:1 compression ratio, and it means those thirty exchanges now cost 600 tokens on every subsequent request instead of 14,400.
The three shapes of summarisation
| Type | How it works | Cost per compaction | Weakness |
|---|---|---|---|
| Rolling / recursive | New summary = summarise(old summary + new messages) | One call over a small input | Drift — each generation is a summary of a summary; errors compound |
| Map-reduce | Summarise each block of raw messages independently, then combine the block summaries | N parallel calls plus one combine | More expensive; blocks lose cross-block context |
| Hierarchical | Level 1 summarises raw blocks; level 2 summarises level-1 summaries; and so on | Amortised — most turns cost nothing | More machinery; older levels are very lossy |
Writing the summarisation prompt is the whole job
This is where summarisation quietly fails, and it is worth being blunt about it. The default instruction people write is:
Summarise the conversation so far.A general-purpose summariser optimises for readability, and readable summaries drop exactly what a coaching bot or a support bot needs: numbers, identifiers, explicit constraints. "The user discussed their dietary requirements and asked about lunch options" is a fine English summary and a catastrophic memory record — it has thrown away "vegetarian" and "peanut allergy", the only two facts that mattered.
Constrain the output instead:
1SUMMARY_PROMPT = """You are compressing a conversation into a memory record.23Preserve verbatim, in every case:4- Names, IDs, order numbers, account numbers, dates, amounts5- Stated constraints, preferences, allergies, restrictions6- Decisions made and commitments given by the assistant7- Questions the user asked that are still unanswered89Discard: pleasantries, restatements, the assistant's own explanations,10anything already superseded by a later message.1112If a fact was later corrected, record only the corrected version and13note that it was a correction.1415Output under 600 tokens using these headings:16FACTS / CONSTRAINTS / DECISIONS / OPEN1718PREVIOUS SUMMARY:19{previous_summary}2021NEW MESSAGES:22{new_messages}"""Compare what the two prompts produce from the same source material:
| Generic summary | Constrained summary |
|---|---|
| "The user discussed dietary requirements and asked about high-protein lunch options. The assistant offered several suggestions." | "FACTS: user is vegetarian; peanut allergy (severe). Target 120g protein/day. CONSTRAINTS: no meat, no peanuts or peanut oil. DECISIONS: assistant agreed to suggest only vegetarian, peanut-free meals. OPEN: user asked about protein powder brands — not yet answered." |
| 18 tokens. Useless. | ~70 tokens. Actually a memory. |
What summarisation costs
Two costs, and people budget for neither.
Money. Compressing 14,400 tokens into 600 costs one call: 14,400 input tokens at 3 dollars per million is 4.32 cents, plus 600 output tokens at 15 dollars per million is 0.9 cents. About 5.2 cents per compaction. Cheap, but not free, and it is charged on top of the conversation itself.
Latency. The compaction call takes two to five seconds. If you run it synchronously on the turn that triggers it, one unlucky user waits five extra seconds for a reply and has no idea why. Run compaction on the previous turn's boundary, in the background, or on a cheaper and faster model — summarisation is a much easier task than the conversation itself, and a small model does it well.
The hybrid: summary plus verbatim tail
Neither pure strategy is what production systems use. The standard design keeps both:
[ system prompt ] fixed, cached[ pinned facts ] ~300 tokens, never evicted[ running summary ] ~600 tokens, covers turns 1..M[ verbatim recent turns ] ~4,800-9,600 tokens, turns M+1..N[ current user message ]The reasoning is that the two halves fail in opposite directions. Recent turns need exact wording — the user says "the second one", and only the verbatim text tells you what that refers to. Old turns need only their durable content, and compression is nearly free there.
1class HybridMemory:2 def __init__(self, keep_min=10, keep_max=20, summariser=None, count_fn=None):3 self.keep_min = keep_min # exchanges kept verbatim after compaction4 self.keep_max = keep_max # trigger point5 self.summary = ""6 self.exchanges = [] # list of (user_msg, assistant_msg)7 self.summarise = summariser8 self.count = count_fn910 def add(self, user_msg, assistant_msg):11 self.exchanges.append((user_msg, assistant_msg))12 if len(self.exchanges) > self.keep_max:13 self._compact()1415 def _compact(self):16 n = len(self.exchanges) - self.keep_min17 old, self.exchanges = self.exchanges[:n], self.exchanges[n:]18 self.summary = self.summarise(self.summary, old)1920 def build(self):21 blocks = []22 if self.summary:23 blocks.append({"role": "user",24 "content": f"[Earlier conversation]\n{self.summary}"})25 blocks.append({"role": "assistant",26 "content": "Understood, I have that context."})27 for u, a in self.exchanges:28 blocks.append(u)29 blocks.append(a)30 return blocksNote the chunked compaction: it triggers at 20 exchanges and cuts back to 10, so it fires once every ten turns rather than every turn. That is the cache-friendly shape.
The numbers for a 200-exchange conversation
Same parameters as before: fixed prefix 3,000 tokens, 480 tokens per exchange, 3 dollars per million input, 15 per million output.
| Strategy | Input tokens billed | Extra calls | Total cost | Facts from turn 3 survive? |
|---|---|---|---|---|
| Full buffer | 10,152,000 | 0 | 30.46 dollars | Yes — until it overflows |
| Sliding window (25 exchanges) | 2,844,000 | 0 | 8.53 dollars | No |
| Hybrid (summary + 10–20 verbatim) | 2,180,000 | 18 compactions | 7.05 dollars | Yes, in compressed form |
The hybrid arithmetic, worked: the request averages 3,000 fixed + 700 summary + about 7,200 verbatim (oscillating between 10 and 20 exchanges) = 10,900 tokens. Across 200 turns that is 2,180,000 tokens, or 6.54 dollars. Compaction fires 18 times (at exchange 21, then every tenth exchange), each summarising 11 exchanges (5,280 tokens) plus the previous summary (700) into a new 700-token summary: (5,980×3+700×15)/106=0.0284 dollars each, so 0.51 dollars total. Sum: 7.05 dollars, a 77% reduction against the full buffer — and unlike the sliding window, the vegetarian and the peanut allergy are still there.
Failure modes worth naming
| Failure | What you see | Root cause | Fix |
|---|---|---|---|
| Early-fact amnesia | Bot violates a constraint stated at turn 2 | Pure recency window | Pin extracted facts outside the evictable history |
| Summary drift | After many compactions, details are subtly wrong or invented | Recursively summarising summaries; each pass is lossy and each pass can hallucinate | Keep raw messages in a store and re-summarise from originals; or hold facts in an append-only list rather than free text |
| Number loss | Amounts, IDs and dates disappear | Unconstrained summarisation prompt | Explicit preserve-verbatim instruction and a fixed output schema |
| Stale corrections | Bot uses a value the user already corrected | Summary appended both the original and the correction | Instruct the summariser to record only the corrected value |
| Cache thrash | Costs far above the estimate; cache_read_input_tokens is zero | Evicting one message per turn changes the prefix every turn | Compact in chunks, not continuously |
| Latency spike | One turn in ten takes five seconds longer | Synchronous compaction on the triggering turn | Compact asynchronously, or on a smaller and faster model |
| Overflow despite a window | 400 error with a message-count window in place | A single tool result was 9,000 tokens | Budget in tokens; truncate tool output at the source |
Frameworks and managed options
Older LangChain code used memory classes such as ConversationBufferWindowMemory, ConversationTokenBufferMemory, ConversationSummaryMemory and ConversationSummaryBufferMemory. They map exactly onto the strategies above: message-count window, token-count window, pure rolling summary, and hybrid. In LangChain 1.x they survive only in the langchain-classic compatibility package. Current code keeps the thread's messages in a LangGraph checkpointer (for example InMemorySaver in development, a Postgres saver in production), and shapes what the model sees with middleware on create_agent: a before_model hook that trims, or the built-in SummarizationMiddleware, which summarises once a token or message threshold is reached and keeps the most recent messages verbatim. That is the hybrid above, as configuration. The change is a good one — the class-based versions hid the eviction decision, and hidden eviction is precisely how the chicken satay reached a vegetarian.
Model providers now offer managed versions too. The Claude API has server-side compaction (in beta): the API writes the summary of older turns for you and hands back a compaction block that replaces them on the next request. It also has context editing, which clears old tool results or thinking blocks by rule instead of summarising them — useful for agents whose context fills up with tool output rather than chat. Both save you writing the summariser; neither removes the need to pin facts that must never be lost, or the need to log what was compacted.
Whatever you use, the requirement is the same: you must be able to answer "what is currently in the context, and what was dropped?" from your logs. If the abstraction will not tell you that, replace it.
Choosing a strategy for a real system
Work through it in this order.
Ask how long a session actually runs. If your median conversation is six turns, a token window is the entire answer and summarisation is machinery you do not need. Measure the distribution before building anything; a great many chatbots that shipped elaborate summarisation had a 95th-percentile session of nine turns.
Identify what must never be lost, and take it out of the transcript. Before choosing between window and summary, write the list: order number, stated allergies, budget ceiling, the plan the user agreed to. Extract those into a pinned block of a few hundred tokens that is rebuilt into every prompt and is not part of the evictable history. This single step fixes most of the failures above, and it costs about 0.09 cents per turn. Doing it means your window or summary strategy only has to handle conversational flow, which is a much easier problem.
Then choose by session length. Under 20 exchanges: token window, nothing else. Twenty to a hundred: hybrid, compacting in chunks. Over a hundred, or across sessions: hybrid plus a persistent store, because at that point the raw transcript belongs in a database rather than in a variable.
Log every eviction. One structured line per compaction — turn number, messages dropped, tokens before and after, a hash of the resulting summary. When a user reports that the bot forgot something, you want to answer in thirty seconds rather than reconstruct the session from scratch. The nutrition bot's incident took two days to diagnose, and every minute of it was spent establishing something a single log line would have said.