Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to create a LangChain chain with summary memory.
What you need to know
Why summarise
A verbatim history grows every turn. After 50 turns you might send 15,000 tokens per call. A summary of the first 40 turns might be 300 tokens, plus the last 10 turns verbatim — perhaps 3,000 tokens in total. Long-range facts ("the customer's router is a TP-200") survive in compressed form.
The 1.x way: middleware
1from langchain.agents import create_agent2from langchain.agents.middleware import SummarizationMiddleware3from langgraph.checkpoint.postgres import PostgresSaver45def build_chat(saver: PostgresSaver):6 return create_agent(7 model=settings.chat_model, # model names live in config8 tools=[],9 system_prompt="You are a help-centre assistant for a broadband provider.",10 checkpointer=saver,11 middleware=[SummarizationMiddleware(12 model=settings.small_model, # cheap model writes the summary13 trigger=("tokens", 4000), # summarise when history passes 4k tokens14 keep=("messages", 10), # keep the last 10 messages verbatim15 )],16 )Before each model call the middleware counts the history. Past the trigger, it summarises the older messages, replaces them in the saved state with one summary message, and keeps the last 10. It avoids splitting a tool call from its result.
In your own graph: write it yourself
1from langchain_core.messages import HumanMessage2from langchain_core.messages.utils import count_tokens_approximately34def compress(messages, summariser, max_tokens=4000, keep=6):5 if count_tokens_approximately(messages) <= max_tokens or len(messages) <= keep:6 return messages7 old, recent = messages[:-keep], messages[-keep:]8 text = "\n".join(f"{m.type}: {m.content}" for m in old)9 summary = summariser.invoke(10 "Summarise this support chat in at most 6 bullets. Keep names, IDs, "11 "numbers, and anything the customer asked us to do.\n\n" + text)12 return [HumanMessage(f"[Summary of earlier conversation]\n{summary.content}")] + recentCall it in a graph node that runs before your chain, and write the result back to state so the summary is made once per threshold crossing, not on every call:
1from langchain.messages import RemoveMessage2from langgraph.graph.message import REMOVE_ALL_MESSAGES34def summarise_node(state: MessagesState):5 new = compress(state["messages"], small_model)6 if new is state["messages"]:7 return {} # under budget: no change8 return {"messages": [RemoveMessage(id=REMOVE_ALL_MESSAGES), *new]}RemoveMessage(id=REMOVE_ALL_MESSAGES) clears the stored list and the new messages replace it, so the checkpointer now holds the summary plus recent turns. The summary is stored as a clearly labelled human-role message because some providers accept a system message only at the very start of the conversation. Older code did the same with history.clear() and history.add_messages(...) on a chat message history.
Choosing the settings
- Trigger on tokens, not turns. One pasted log can be larger than 30 short turns.
- Keep enough recent turns (6 to 10 messages) for pronouns and follow-ups to resolve.
- Tell the summariser what matters. IDs, amounts, dates, promises made. A generic "summarise" drops exactly these.
- Use a cheaper model for summaries; they do not need the strongest reasoning.
The trade-offs
| Gain | Cost |
|---|---|
| Near-constant prompt size | One extra model call each time it triggers |
| Long chats never overflow | Details the summary drops are gone |
| Lower cost on long chats | A summary can contain a mistake that then persists |
A real-life example
A broadband help-centre bot has long troubleshooting chats: customers paste router logs, try steps, report back. The median chat is 8 turns, but the top 10% run past 40 turns and 20,000 tokens.
With SummarizationMiddleware (trigger 4,000 tokens, keep 10 messages), those long chats settle at about 4,500 tokens per call. The first summary prompt was generic, and in testing it dropped the customer's account number, so the bot asked for it again. The team changed the summary prompt to always keep IDs, device models and steps already tried. In a review of 100 long chats, "asked for information already given" fell from 14 cases to 2.
Follow-up questions to expect
- "Summary memory versus a sliding window?" — A window drops old turns completely; a summary keeps a compressed version, at the cost of an extra call.
- "How do you stop the summary drifting over many rounds?" — Summarise the previous summary plus new messages with a prompt that preserves IDs and commitments, and keep critical facts in structured state instead.
- "What was
ConversationSummaryBufferMemory?" — The legacy class doing the same: summary of old turns plus a recent-turns buffer up to a token limit.