LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

What is memory in LangChain, and how is it used in NLP?


What you need to know

Why memory is needed

Turn 1: "What is the warranty on the X200 laptop?" Turn 2: "Does it cover the battery?" Without memory, turn 2 arrives alone, and the model has no idea what "it" is. With memory, your code sends turn 1, the answer, and turn 2 together.

So "memory" is not magic inside the model. It is: store messages, select which ones to send, inject them into the prompt.

Short-term and long-term memory

Short-term (conversation)Long-term (across conversations)
HoldsMessages and tool results of one threadFacts: preferences, profile, past decisions
LifetimeOne conversationWeeks or months
LangChain 1.xCheckpointer + thread_idLangGraph store + namespace
Example"it" means the X200"Prefers replies in Hindi"

How the APIs map

SituationCurrent approachLegacy equivalent
Agent or chat appcreate_agent(..., checkpointer=...)ConversationBufferMemory on an agent
Plain LCEL chainRun it inside a StateGraph node with a checkpointerConversationChain(memory=...), later RunnableWithMessageHistory
Long historiesSummarizationMiddleware or trim_messagesConversationSummaryBufferMemory, ConversationTokenBufferMemory
Facts across sessionsLangGraph store (InMemoryStore, PostgresStore)ConversationEntityMemory

The legacy memory classes moved to the langchain-classic package in 1.0 and are marked for removal in 2.0. RunnableWithMessageHistory, the LCEL-era replacement, was itself deprecated in langchain-core 1.3.3 (its warning says "Use LangGraph's built-in persistence instead"), and InMemoryChatMessageHistory in 1.6.4. You will still see both in 2024–2025 code, so know how they work.

The current code, minimal

Python
from langchain.agents import create_agentfrom langgraph.checkpoint.memory import InMemorySaver   # Postgres in productionbot = create_agent(model, tools=[], checkpointer=InMemorySaver(),                   system_prompt="You answer questions about our laptops.")cfg = {"configurable": {"thread_id": "cust-42"}}bot.invoke({"messages": [{"role": "user", "content": "Warranty on the X200?"}]}, cfg)bot.invoke({"messages": [{"role": "user", "content": "Does it cover the battery?"}]}, cfg)

Even with no tools, create_agent gives you a chat app with saved history. On the second call you send only the new message; the checkpointer loads the rest.

The core trade-off

Everything you remember costs tokens on every call. A 40-turn conversation might be 12,000 tokens, sent again each turn. So every memory design balances how much is remembered against cost, latency and context-window limits. That is why trimming, summarising and retrieval of old turns exist.

A real-life example

An electronics store's product Q&A assistant launched without memory, treating each message separately. Analytics showed 35% of conversations had a second message containing "it", "that one" or "the cheaper one" — and the assistant answered those badly, often asking "Which product do you mean?".

Adding a checkpointer with the last 20 messages fixed follow-ups. The prompt grew from about 800 tokens to about 2,500 tokens on an average conversation, which the team judged worth it. Later they added SummarizationMiddleware for the few long chats that went past 8,000 tokens.

Follow-up questions to expect

  • "Does the model remember anything by itself?" — No. Each call is independent; memory is your application re-sending context.
  • "What is the difference between a checkpointer and a store?" — A checkpointer saves one thread's state; a store holds data shared across threads, such as a user profile.
  • "Is RunnableWithMessageHistory still the way to do it?" — No. It is deprecated since langchain-core 1.3.3 and due for removal in 2.0; new code uses LangGraph persistence.