Course Content
LangChain Mastery
7 sections · 109 lessons
What is memory in LangChain, and how is it used in NLP?
What you need to know
Why memory is needed
Turn 1: "What is the warranty on the X200 laptop?" Turn 2: "Does it cover the battery?" Without memory, turn 2 arrives alone, and the model has no idea what "it" is. With memory, your code sends turn 1, the answer, and turn 2 together.
So "memory" is not magic inside the model. It is: store messages, select which ones to send, inject them into the prompt.
Short-term and long-term memory
| Short-term (conversation) | Long-term (across conversations) | |
|---|---|---|
| Holds | Messages and tool results of one thread | Facts: preferences, profile, past decisions |
| Lifetime | One conversation | Weeks or months |
| LangChain 1.x | Checkpointer + thread_id | LangGraph store + namespace |
| Example | "it" means the X200 | "Prefers replies in Hindi" |
How the APIs map
| Situation | Current approach | Legacy equivalent |
|---|---|---|
| Agent or chat app | create_agent(..., checkpointer=...) | ConversationBufferMemory on an agent |
| Plain LCEL chain | Run it inside a StateGraph node with a checkpointer | ConversationChain(memory=...), later RunnableWithMessageHistory |
| Long histories | SummarizationMiddleware or trim_messages | ConversationSummaryBufferMemory, ConversationTokenBufferMemory |
| Facts across sessions | LangGraph store (InMemoryStore, PostgresStore) | ConversationEntityMemory |
The legacy memory classes moved to the langchain-classic package in 1.0 and are marked for removal in 2.0. RunnableWithMessageHistory, the LCEL-era replacement, was itself deprecated in langchain-core 1.3.3 (its warning says "Use LangGraph's built-in persistence instead"), and InMemoryChatMessageHistory in 1.6.4. You will still see both in 2024–2025 code, so know how they work.
The current code, minimal
1from langchain.agents import create_agent2from langgraph.checkpoint.memory import InMemorySaver # Postgres in production34bot = create_agent(model, tools=[], checkpointer=InMemorySaver(),5 system_prompt="You answer questions about our laptops.")6cfg = {"configurable": {"thread_id": "cust-42"}}7bot.invoke({"messages": [{"role": "user", "content": "Warranty on the X200?"}]}, cfg)8bot.invoke({"messages": [{"role": "user", "content": "Does it cover the battery?"}]}, cfg)Even with no tools, create_agent gives you a chat app with saved history. On the second call you send only the new message; the checkpointer loads the rest.
The core trade-off
Everything you remember costs tokens on every call. A 40-turn conversation might be 12,000 tokens, sent again each turn. So every memory design balances how much is remembered against cost, latency and context-window limits. That is why trimming, summarising and retrieval of old turns exist.
A real-life example
An electronics store's product Q&A assistant launched without memory, treating each message separately. Analytics showed 35% of conversations had a second message containing "it", "that one" or "the cheaper one" — and the assistant answered those badly, often asking "Which product do you mean?".
Adding a checkpointer with the last 20 messages fixed follow-ups. The prompt grew from about 800 tokens to about 2,500 tokens on an average conversation, which the team judged worth it. Later they added SummarizationMiddleware for the few long chats that went past 8,000 tokens.
Follow-up questions to expect
- "Does the model remember anything by itself?" — No. Each call is independent; memory is your application re-sending context.
- "What is the difference between a checkpointer and a store?" — A checkpointer saves one thread's state; a store holds data shared across threads, such as a user profile.
- "Is
RunnableWithMessageHistorystill the way to do it?" — No. It is deprecated sincelangchain-core1.3.3 and due for removal in 2.0; new code uses LangGraph persistence.