LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to create a LangChain chain with summary memory.


What the model sees in a 40-turn chatNew user messageLast 10messages, verbatimSummary ofeverything olderSystem prompttopbottomAbout 4,500 tokens instead of 20,000; the summary prompt must keep IDs and steps tried.
Old turns shrink to a few hundred tokens while recent ones stay exact, so follow-ups still resolve.

What you need to know

Why summarise

A verbatim history grows every turn. After 50 turns you might send 15,000 tokens per call. A summary of the first 40 turns might be 300 tokens, plus the last 10 turns verbatim — perhaps 3,000 tokens in total. Long-range facts ("the customer's router is a TP-200") survive in compressed form.

The 1.x way: middleware

Python
from langchain.agents import create_agentfrom langchain.agents.middleware import SummarizationMiddlewarefrom langgraph.checkpoint.postgres import PostgresSaverdef build_chat(saver: PostgresSaver):    return create_agent(        model=settings.chat_model,       # model names live in config        tools=[],        system_prompt="You are a help-centre assistant for a broadband provider.",        checkpointer=saver,        middleware=[SummarizationMiddleware(            model=settings.small_model,      # cheap model writes the summary            trigger=("tokens", 4000),        # summarise when history passes 4k tokens            keep=("messages", 10),           # keep the last 10 messages verbatim        )],    )

Before each model call the middleware counts the history. Past the trigger, it summarises the older messages, replaces them in the saved state with one summary message, and keeps the last 10. It avoids splitting a tool call from its result.

In your own graph: write it yourself

Python
from langchain_core.messages import HumanMessagefrom langchain_core.messages.utils import count_tokens_approximatelydef compress(messages, summariser, max_tokens=4000, keep=6):    if count_tokens_approximately(messages) <= max_tokens or len(messages) <= keep:        return messages    old, recent = messages[:-keep], messages[-keep:]    text = "\n".join(f"{m.type}: {m.content}" for m in old)    summary = summariser.invoke(        "Summarise this support chat in at most 6 bullets. Keep names, IDs, "        "numbers, and anything the customer asked us to do.\n\n" + text)    return [HumanMessage(f"[Summary of earlier conversation]\n{summary.content}")] + recent

Call it in a graph node that runs before your chain, and write the result back to state so the summary is made once per threshold crossing, not on every call:

Python
from langchain.messages import RemoveMessagefrom langgraph.graph.message import REMOVE_ALL_MESSAGESdef summarise_node(state: MessagesState):    new = compress(state["messages"], small_model)    if new is state["messages"]:        return {}                                   # under budget: no change    return {"messages": [RemoveMessage(id=REMOVE_ALL_MESSAGES), *new]}

RemoveMessage(id=REMOVE_ALL_MESSAGES) clears the stored list and the new messages replace it, so the checkpointer now holds the summary plus recent turns. The summary is stored as a clearly labelled human-role message because some providers accept a system message only at the very start of the conversation. Older code did the same with history.clear() and history.add_messages(...) on a chat message history.

Choosing the settings

  • Trigger on tokens, not turns. One pasted log can be larger than 30 short turns.
  • Keep enough recent turns (6 to 10 messages) for pronouns and follow-ups to resolve.
  • Tell the summariser what matters. IDs, amounts, dates, promises made. A generic "summarise" drops exactly these.
  • Use a cheaper model for summaries; they do not need the strongest reasoning.

The trade-offs

GainCost
Near-constant prompt sizeOne extra model call each time it triggers
Long chats never overflowDetails the summary drops are gone
Lower cost on long chatsA summary can contain a mistake that then persists

A real-life example

A broadband help-centre bot has long troubleshooting chats: customers paste router logs, try steps, report back. The median chat is 8 turns, but the top 10% run past 40 turns and 20,000 tokens.

With SummarizationMiddleware (trigger 4,000 tokens, keep 10 messages), those long chats settle at about 4,500 tokens per call. The first summary prompt was generic, and in testing it dropped the customer's account number, so the bot asked for it again. The team changed the summary prompt to always keep IDs, device models and steps already tried. In a review of 100 long chats, "asked for information already given" fell from 14 cases to 2.

Follow-up questions to expect

  • "Summary memory versus a sliding window?" — A window drops old turns completely; a summary keeps a compressed version, at the cost of an extra call.
  • "How do you stop the summary drifting over many rounds?" — Summarise the previous summary plus new messages with a prompt that preserves IDs and commitments, and keep critical facts in structured state instead.
  • "What was ConversationSummaryBufferMemory?" — The legacy class doing the same: summary of old turns plus a recent-turns buffer up to a token limit.