Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 3: Memory Saturation Issue
Scenario: a LangChain chat app keeps the whole history in memory, and long conversations become slow, expensive and eventually exceed the context window. How do you fix it?
What you need to know
The old langchain.memory classes (ConversationBufferMemory, ConversationSummaryBufferMemory) hid memory inside a chain object. They now live only in the langchain-classic package for backward compatibility. The current approach makes memory visible:
Three tiers
- Recent turns, verbatim — trimmed to a token budget.
- Rolling summary — older turns compressed by a cheap model when the history crosses a threshold.
- Structured facts — order ids, entitlements, stated preferences, kept in state and injected exactly each turn.
Trimming in a chain or graph node
1from langchain_core.messages import SystemMessage, trim_messages23def call_model(state):4 recent = trim_messages(state["messages"], max_tokens=6000, strategy="last",5 token_counter=llm, include_system=True, start_on="human")6 facts = SystemMessage(f"Known facts: {state['facts']}")7 return {"messages": [llm.invoke([facts, *recent])]}start_on="human" matters: cutting in the middle of a tool-call sequence leaves a tool result without its request, and some providers reject that message list outright.
Summarisation, built in
In a LangChain 1.x agent, SummarizationMiddleware does tier 2 for you:
1from langchain.agents import create_agent2from langchain.agents.middleware import SummarizationMiddleware3from langgraph.checkpoint.postgres import PostgresSaver45with PostgresSaver.from_conn_string(DB_URL) as checkpointer: # checkpointer.setup() once6 agent = create_agent(7 model=llm, tools=tools,8 middleware=[SummarizationMiddleware(model=small_llm, trigger=("tokens", 8000),9 keep=("messages", 12))],10 checkpointer=checkpointer)11 agent.invoke({"messages": [("user", text)]},12 config={"configurable": {"thread_id": chat_id}})When the conversation passes 8,000 tokens, older messages are replaced by a summary and the last 12 are kept. The checkpointer stores everything per thread_id, so the conversation survives restarts and works across replicas.
Why facts get their own place
Summaries paraphrase. "Order ORD-48212, delivered to the Pune address, damaged box" becomes "the user had an issue with an order". Keep identifiers and decisions as structured state, updated by a small extraction step or by tools, and inject them every turn.
Options compared
| Approach | Tokens per turn | Keeps exact ids | Persistence |
|---|---|---|---|
| Full history in memory (legacy) | Grows without limit | Yes, until it overflows | Lost on restart |
| Trim only | Fixed | Only recent ones | With a checkpointer |
| Trim plus summary plus facts | Small and bounded | Yes | With a checkpointer |
A real-life example
Scenario (illustrative numbers). An insurance company's claims chatbot uses ConversationBufferMemory in process memory. Claims conversations often run 50 turns. By turn 40, prompts reach 45,000 tokens, replies take 7 seconds, and a pod restart loses the conversation entirely.
The team moves to a LangGraph agent with a Postgres checkpointer, summarisation middleware at 8,000 tokens, and a facts dictionary holding the claim number, policy number and incident date. Median tokens per turn at turn 40 fall to about 7,500, reply time to 1.8 seconds, and restarts no longer lose conversations. On 30 long test transcripts asking about early details ("what date did I say the accident happened?"), 29 are answered correctly.
Follow-up questions to expect
- "What about memory across conversations?" — That is long-term memory: a LangGraph store (or your own database) holding user facts, separate from the per-thread checkpointer.
- "Where does the summary live?" — In graph state, persisted by the checkpointer alongside the messages.
- "How do you test it?" — Long synthetic or replayed transcripts with questions whose answers sit in the summarised region; track tokens per turn and cost per conversation.