Course Content
LangChain Mastery
7 sections · 109 lessons
How do you add memory to a LangChain chain?
What you need to know
The idea in one line
Memory for a chain means: load this conversation's past messages, inject them into the prompt, run, then save the new question and answer. The only question is which component does the loading and saving.
The current way: the chain inside a graph with a checkpointer
1from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder2from langgraph.graph import StateGraph, MessagesState, START3from langgraph.checkpoint.memory import InMemorySaver # PostgresSaver in production45prompt = ChatPromptTemplate.from_messages([6 ("system", "You answer questions from our help centre. Be brief."),7 MessagesPlaceholder("messages"), # the whole conversation, newest last8])9chain = prompt | model # your existing LCEL chain1011def respond(state: MessagesState):12 return {"messages": [chain.invoke({"messages": state["messages"]})]}1314g = StateGraph(MessagesState)15g.add_node("respond", respond)16g.add_edge(START, "respond")17app = g.compile(checkpointer=InMemorySaver())1819cfg = {"configurable": {"thread_id": "cust-42"}}20app.invoke({"messages": [("user", "How do I reset my router?")]}, cfg)21app.invoke({"messages": [("user", "And if the light stays red?")]}, cfg)What happens on each invoke:
- Load — the checkpointer reads the saved state for
thread_id="cust-42". - Append — the new user message is added to
messages(theMessagesStatereducer appends). - Run —
respondcalls your chain with the full message list. - Save — the AI reply is appended and the new state is checkpointed.
The chain itself is unchanged; the graph adds memory around it. InMemorySaver is for development only. In production use PostgresSaver (or SQLite, Redis, MongoDB) so any server can load any thread and restarts lose nothing. Add trim_messages inside respond before the history outgrows the context window.
If you will add tools, human approval or limits later, create_agent(model, tools=[], system_prompt=..., checkpointer=...) gives you the same memory with middleware support.
The legacy way: RunnableWithMessageHistory
You will meet this in code from 2024 and 2025, and interviewers may ask about it by name:
1from langchain_core.runnables.history import RunnableWithMessageHistory # deprecated23chat = RunnableWithMessageHistory(4 chain, # prompt must have MessagesPlaceholder("history")5 get_history, # session_id -> a BaseChatMessageHistory6 input_messages_key="question", # which input is the new user turn7 history_messages_key="history", # must match the placeholder name8)9chat.invoke({"question": "How do I reset my router?"},10 config={"configurable": {"session_id": "cust-42"}})It wrapped the chain: before the call it loaded history from get_history(session_id); after the call it saved the new human and AI messages. It raised an error if session_id was missing. Since langchain-core 1.3.3 it emits a deprecation warning saying to use LangGraph's built-in persistence, and it is due for removal in 2.0. Before that, LLMChain(memory=ConversationBufferMemory()) and ConversationChain were the way; both are in langchain-classic now.
Why the graph version is better
- It saves more than text. Tool calls, tool results and any custom state fields are checkpointed too.
- It can pause and resume, for human approval or after a crash.
- One mechanism everywhere. Chains, agents and multi-agent graphs all use the same checkpointer and
thread_id.
A real-life example
A broadband company's help-centre bot started in 2024 as a single LCEL chain wrapped in RunnableWithMessageHistory, with histories in Redis. After upgrading to LangChain 1.x, the logs filled with deprecation warnings, and the team also wanted a check_outage tool and a "talk to an agent" handover — neither fits the wrapper.
They moved the same prompt and chain into a one-node StateGraph with a Postgres checkpointer, keyed by thread_id = f"{user_id}:{chat_id}". The chain code did not change. A month later, swapping the node for create_agent with the outage tool and HumanInTheLoopMiddleware took one afternoon, because memory, pausing and resuming were already handled by the checkpointer.
Follow-up questions to expect
- "What replaced
RunnableWithMessageHistory?" — LangGraph persistence: a checkpointer plus athread_id, throughcreate_agentor your ownStateGraph. - "Why not keep history in a global list?" — All users would share one conversation, and it would vanish on restart.
- "How do you add memory to a RAG chain?" — Same graph; add a node that rewrites the follow-up into a standalone question before retrieval, so "and the red light?" searches for "router red light".