LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you debug LangChain memory issues in long conversations?


Input tokens per turn in a long leave-planning chat9002100330046005900140026000123456trigger at 6,000summarylost the datesThe drop is the summarisation middleware firing; the next turn's trace shows what survived.
Logging tokens per turn turns 'the bot forgets' into a visible event you can open and read.

What you need to know

Look at what was sent

For an agent with a checkpointer, read the stored state for the thread:

Python
config = {"configurable": {"thread_id": "emp-10233"}}state = agent.get_state(config)messages = state.values["messages"]from langchain_core.messages.utils import count_tokens_approximatelyprint("messages:", len(messages), "tokens:", count_tokens_approximately(messages))for m in messages[:2] + messages[-4:]:    print(m.type, "|", str(m.content)[:80])

In LangSmith, open the model run for the failing turn: the input shows the exact messages after any trimming or summarising middleware ran — which may differ from what is stored.

The usual suspects

SymptomLikely causeCheck
Cost and latency rise every turn, then a context errorTrimming or summarisation not runningPlot tokens per turn — should level off
Forgets something said 2 turns agoTrimmed from the wrong end, or budget too smallstrategy="last", larger budget
Forgets the system instructionsSystem message trimmed awayinclude_system=True
Facts slowly change ("12 days" becomes "10 days")Summaries of summariesKeep key facts in structured state; summarise from original messages
Provider error about tool messagesTrimming split a tool call from its resultstart_on="human", trim at message-pair boundaries
Sees another user's historyShared or wrong thread_idBuild ids from the authenticated user, never from input

Store important facts outside the chat

Facts that must never be lost — employee id, approved dates, booking reference — belong in structured state (custom agent state fields or a database), not only in chat text that may be trimmed or summarised.

A real-life example

An HR bot handles leave planning in long chats. Around turn 25, employees report it "forgets" that they already chose dates and asks again.

The engineer logs tokens per turn: they grow until turn 24, then drop sharply — SummarizationMiddleware fires at 6,000 tokens. The LangSmith trace for turn 25 shows the summary: "User discussed leave options." The chosen dates, 14–18 October, were lost in the summary.

Two fixes: the summary prompt now says "keep all dates, numbers and decisions exactly", and the chosen dates are saved into a leave_draft field in the agent state when the user confirms them, so they no longer depend on chat history. The repeat-question complaints stopped.

Follow-up questions to expect

  • "How do you test memory behaviour?" — Scripted multi-turn conversations in the eval set that ask about something from turn 3 at turn 30.
  • "Trim or summarise?" — Trim for short support chats; summarise when older context matters, and keep key facts in structured state either way.
  • "What about legacy memory classes?" — ConversationBufferMemory, ConversationSummaryMemory and friends are in langchain-classic; the debugging approach is the same — inspect what reaches the model.