Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 3: Memory Saturation Issue


Scenario: a LangChain chat app keeps the whole history in memory, and long conversations become slow, expensive and eventually exceed the context window. How do you fix it?

What you need to know

The old langchain.memory classes (ConversationBufferMemory, ConversationSummaryBufferMemory) hid memory inside a chain object. They now live only in the langchain-classic package for backward compatibility. The current approach makes memory visible:

Three tiers

  1. Recent turns, verbatim — trimmed to a token budget.
  2. Rolling summary — older turns compressed by a cheap model when the history crosses a threshold.
  3. Structured facts — order ids, entitlements, stated preferences, kept in state and injected exactly each turn.

Trimming in a chain or graph node

Python
from langchain_core.messages import SystemMessage, trim_messagesdef call_model(state):    recent = trim_messages(state["messages"], max_tokens=6000, strategy="last",                           token_counter=llm, include_system=True, start_on="human")    facts = SystemMessage(f"Known facts: {state['facts']}")    return {"messages": [llm.invoke([facts, *recent])]}

start_on="human" matters: cutting in the middle of a tool-call sequence leaves a tool result without its request, and some providers reject that message list outright.

Summarisation, built in

In a LangChain 1.x agent, SummarizationMiddleware does tier 2 for you:

Python
from langchain.agents import create_agentfrom langchain.agents.middleware import SummarizationMiddlewarefrom langgraph.checkpoint.postgres import PostgresSaverwith PostgresSaver.from_conn_string(DB_URL) as checkpointer:   # checkpointer.setup() once    agent = create_agent(        model=llm, tools=tools,        middleware=[SummarizationMiddleware(model=small_llm, trigger=("tokens", 8000),                                            keep=("messages", 12))],        checkpointer=checkpointer)    agent.invoke({"messages": [("user", text)]},                 config={"configurable": {"thread_id": chat_id}})

When the conversation passes 8,000 tokens, older messages are replaced by a summary and the last 12 are kept. The checkpointer stores everything per thread_id, so the conversation survives restarts and works across replicas.

Why facts get their own place

Summaries paraphrase. "Order ORD-48212, delivered to the Pune address, damaged box" becomes "the user had an issue with an order". Keep identifiers and decisions as structured state, updated by a small extraction step or by tools, and inject them every turn.

Options compared

ApproachTokens per turnKeeps exact idsPersistence
Full history in memory (legacy)Grows without limitYes, until it overflowsLost on restart
Trim onlyFixedOnly recent onesWith a checkpointer
Trim plus summary plus factsSmall and boundedYesWith a checkpointer

A real-life example

Scenario (illustrative numbers). An insurance company's claims chatbot uses ConversationBufferMemory in process memory. Claims conversations often run 50 turns. By turn 40, prompts reach 45,000 tokens, replies take 7 seconds, and a pod restart loses the conversation entirely.

The team moves to a LangGraph agent with a Postgres checkpointer, summarisation middleware at 8,000 tokens, and a facts dictionary holding the claim number, policy number and incident date. Median tokens per turn at turn 40 fall to about 7,500, reply time to 1.8 seconds, and restarts no longer lose conversations. On 30 long test transcripts asking about early details ("what date did I say the accident happened?"), 29 are answered correctly.

Follow-up questions to expect

  • "What about memory across conversations?" — That is long-term memory: a LangGraph store (or your own database) holding user facts, separate from the per-thread checkpointer.
  • "Where does the summary live?" — In graph state, persisted by the checkpointer alongside the messages.
  • "How do you test it?" — Long synthetic or replayed transcripts with questions whose answers sit in the summarised region; track tokens per turn and cost per conversation.