Course Content
LangChain Mastery
7 sections · 109 lessons
How do you handle memory overflow in LangChain?
What you need to know
What fills the context window
The context window is the maximum number of tokens a model can take in one call, input and output together. History is only one part:
system prompt 800tool schemas 1,500retrieved documents 4,000conversation history ?????space for the answer 1,000On a 32,000-token model, history can use about 24,000 before the call fails. On a 200,000-token model it rarely fails, but a 150,000-token prompt is slow and costly on every turn, and long prompts can make models pay less attention to instructions in the middle. So "overflow" is also a cost and quality problem, not only an error.
The layers
- Trim —
trim_messages(strategy="last", max_tokens=...)keeps the newest turns that fit. Cheap, fast, the default. - Summarise — past a token threshold,
SummarizationMiddlewarereplaces older turns with a summary. Keeps long-range facts at some loss of detail. - Offload — store big things (pasted logs, long documents, old turns) outside the prompt and retrieve relevant parts by search when needed.
- Shrink tool output — tools return projections, not raw payloads; old tool results can be replaced with a short note.
- Cap — a maximum session length, then "Let's start a new chat — here is a summary of where we were."
Putting it together in create_agent
1from langchain.agents import create_agent2from langchain.agents.middleware import SummarizationMiddleware, wrap_model_call3from langchain_core.messages import trim_messages4from langchain_core.messages.utils import count_tokens_approximately56HARD_LIMIT = 24_00078@wrap_model_call9def guard(request, handler):10 n = count_tokens_approximately(request.messages)11 if n > HARD_LIMIT: # safety net after summarising12 log.warning("history_trimmed tokens=%s", n)13 request = request.override(messages=trim_messages(14 request.messages, max_tokens=HARD_LIMIT, token_counter="approximate",15 strategy="last", start_on="human"))16 return handler(request)1718agent = create_agent(model, tools=TOOLS, checkpointer=saver,19 middleware=[SummarizationMiddleware(model=settings.small_model,20 trigger=("tokens", 8000), keep=("messages", 12)),21 guard])Summarisation handles the normal case; the guard is a safety net for what summarising cannot fix in time, such as several large tool results inside one run. A single message bigger than the whole budget cannot be trimmed usefully — trim_messages never cuts a message in half by default, so it would drop it — so huge pastes must be handled before they enter the chat, as in the example below. Logging each time the guard fires tells you how often your design is being stretched.
Why not let the provider handle it
Some providers return an error; some APIs can truncate for you. Either way you lose control over what is dropped — possibly the system prompt or the tool call the model needs. Deciding yourself is the only way to be predictable.
A real-life example
A broadband help-centre bot let customers paste router diagnostics. Some pastes were 25,000 tokens. On the 32,000-token model the team used for cost reasons, about 1.5% of conversations crashed with a context-length error, and those were the frustrated customers most likely to churn.
They added three things: pasted text over 2,000 tokens is saved to storage and replaced in the chat with "Customer pasted a diagnostic log (saved as log-8812); key lines: ..." from a small extraction step; SummarizationMiddleware at 8,000 tokens; and the guard above. Context errors went to zero. The guard fired on 0.2% of runs, which the team reviews weekly.
Follow-up questions to expect
- "Just use a model with a bigger window?" — It moves the limit but not the cost; every token is paid for on every turn, and very long prompts can reduce attention to instructions.
- "Trim or summarise?" — Trim by default; add a summary when users refer back to things from long ago, such as order IDs or earlier decisions.
- "How do you know it is working?" — Track prompt tokens per turn; it should level off as conversations get longer, not grow in a straight line.