Course Content
LangChain Mastery
7 sections · 109 lessons
What is ConversationBufferMemory in LangChain?
What you need to know
What it did
1# Legacy — for reading old code only2from langchain_classic.memory import ConversationBufferMemory3from langchain_classic.chains import ConversationChain45memory = ConversationBufferMemory(return_messages=True)6chain = ConversationChain(llm=llm, memory=memory)7chain.invoke({"input": "My router model is TP-200."})8chain.invoke({"input": "How do I reset it?"})9memory.load_memory_variables({}) # {"history": [HumanMessage, AIMessage, ...]}- It stored turns on the memory object.
- Before each call,
load_memory_variablesreturned the history under the keyhistory(set bymemory_key). - After each call,
save_contextappended the new input and output. return_messages=Truereturned message objects;Falsereturned one long string like"Human: ...\nAI: ...".
Its family of legacy classes
| Legacy class | What it kept | Current equivalent |
|---|---|---|
ConversationBufferMemory | Everything | Checkpointer + thread_id |
ConversationBufferWindowMemory | Last k turns | trim_messages or a before_model middleware |
ConversationTokenBufferMemory | Last N tokens | trim_messages(max_tokens=...) |
ConversationSummaryMemory | A running summary | SummarizationMiddleware |
ConversationSummaryBufferMemory | Summary + recent turns | SummarizationMiddleware with keep= |
ConversationEntityMemory | Facts per entity | LangGraph store + structured extraction |
Why it was replaced
- State inside the object. One memory object per chain meant one chain per user, or users sharing history by accident.
- Poor fit with LCEL. It predates runnables;
|pipelines have no memory slot. - Lost tool steps. With agents it saved the input and final output, not the intermediate tool calls.
- No persistence or resume. It lived in process memory unless you wired a history backend.
The current equivalent
1from langchain.agents import create_agent2from langgraph.checkpoint.postgres import PostgresSaver34with PostgresSaver.from_conn_string(DB_URI) as saver:5 bot = create_agent(model, tools=[], checkpointer=saver)6 bot.invoke({"messages": [("user", "My router model is TP-200.")]},7 {"configurable": {"thread_id": "cust-42"}})Every message in the thread is kept and sent, just like the buffer, but saved per thread in Postgres.
When "keep everything" is fine
Short sessions — under about 20 turns, or where total history stays well under the context window — such as a form-filling assistant. For long or open-ended chats, add a limit.
A real-life example
Think of a court stenographer who reads the entire transcript aloud before every new question. For a 10-minute hearing that is fine. For a three-day trial, most of the day is spent re-reading, and eventually there is no time left for the new question. The analogy breaks in one way: the stenographer gets tired, but the model gets charged — every re-read transcript is billed as input tokens.
In production: a help-centre bot using ConversationBufferMemory worked in testing, where chats were 4 to 6 turns. In production, 5% of chats ran past 60 turns — customers pasting router logs — and hit the 128,000-token context limit with an error. Those same chats were also the most expensive: their prompts averaged 40,000 tokens per turn. Moving to a checkpointer with SummarizationMiddleware (trigger at 6,000 tokens) removed the errors and cut the cost of long chats by about 80%.
Follow-up questions to expect
- "What is
memory_key?" — The prompt variable name the history is injected under; it defaults tohistoryand must match the prompt. - "Buffer versus window memory?" — Buffer keeps everything; window keeps only the last k turns, which caps cost but forgets older facts.
- "What should I use instead today?" — A LangGraph checkpointer, through
create_agentor your ownStateGraph.RunnableWithMessageHistory, the in-between answer from 2024, is deprecated too.