LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you add memory to a LangChain chain?


What you need to know

The idea in one line

Memory for a chain means: load this conversation's past messages, inject them into the prompt, run, then save the new question and answer. The only question is which component does the loading and saving.

The current way: the chain inside a graph with a checkpointer

Python
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholderfrom langgraph.graph import StateGraph, MessagesState, STARTfrom langgraph.checkpoint.memory import InMemorySaver      # PostgresSaver in productionprompt = ChatPromptTemplate.from_messages([    ("system", "You answer questions from our help centre. Be brief."),    MessagesPlaceholder("messages"),          # the whole conversation, newest last])chain = prompt | model                        # your existing LCEL chaindef respond(state: MessagesState):    return {"messages": [chain.invoke({"messages": state["messages"]})]}g = StateGraph(MessagesState)g.add_node("respond", respond)g.add_edge(START, "respond")app = g.compile(checkpointer=InMemorySaver())cfg = {"configurable": {"thread_id": "cust-42"}}app.invoke({"messages": [("user", "How do I reset my router?")]}, cfg)app.invoke({"messages": [("user", "And if the light stays red?")]}, cfg)

What happens on each invoke:

  1. Load — the checkpointer reads the saved state for thread_id="cust-42".
  2. Append — the new user message is added to messages (the MessagesState reducer appends).
  3. Run — respond calls your chain with the full message list.
  4. Save — the AI reply is appended and the new state is checkpointed.

The chain itself is unchanged; the graph adds memory around it. InMemorySaver is for development only. In production use PostgresSaver (or SQLite, Redis, MongoDB) so any server can load any thread and restarts lose nothing. Add trim_messages inside respond before the history outgrows the context window.

If you will add tools, human approval or limits later, create_agent(model, tools=[], system_prompt=..., checkpointer=...) gives you the same memory with middleware support.

The legacy way: RunnableWithMessageHistory

You will meet this in code from 2024 and 2025, and interviewers may ask about it by name:

Python
from langchain_core.runnables.history import RunnableWithMessageHistory   # deprecatedchat = RunnableWithMessageHistory(    chain,                               # prompt must have MessagesPlaceholder("history")    get_history,                         # session_id -> a BaseChatMessageHistory    input_messages_key="question",       # which input is the new user turn    history_messages_key="history",      # must match the placeholder name)chat.invoke({"question": "How do I reset my router?"},            config={"configurable": {"session_id": "cust-42"}})

It wrapped the chain: before the call it loaded history from get_history(session_id); after the call it saved the new human and AI messages. It raised an error if session_id was missing. Since langchain-core 1.3.3 it emits a deprecation warning saying to use LangGraph's built-in persistence, and it is due for removal in 2.0. Before that, LLMChain(memory=ConversationBufferMemory()) and ConversationChain were the way; both are in langchain-classic now.

Why the graph version is better

  • It saves more than text. Tool calls, tool results and any custom state fields are checkpointed too.
  • It can pause and resume, for human approval or after a crash.
  • One mechanism everywhere. Chains, agents and multi-agent graphs all use the same checkpointer and thread_id.

A real-life example

A broadband company's help-centre bot started in 2024 as a single LCEL chain wrapped in RunnableWithMessageHistory, with histories in Redis. After upgrading to LangChain 1.x, the logs filled with deprecation warnings, and the team also wanted a check_outage tool and a "talk to an agent" handover — neither fits the wrapper.

They moved the same prompt and chain into a one-node StateGraph with a Postgres checkpointer, keyed by thread_id = f"{user_id}:{chat_id}". The chain code did not change. A month later, swapping the node for create_agent with the outage tool and HumanInTheLoopMiddleware took one afternoon, because memory, pausing and resuming were already handled by the checkpointer.

Follow-up questions to expect

  • "What replaced RunnableWithMessageHistory?" — LangGraph persistence: a checkpointer plus a thread_id, through create_agent or your own StateGraph.
  • "Why not keep history in a global list?" — All users would share one conversation, and it would vanish on restart.
  • "How do you add memory to a RAG chain?" — Same graph; add a node that rewrites the follow-up into a standalone question before retrieval, so "and the red light?" searches for "router red light".