LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

What is ConversationBufferMemory in LangChain?


Prompt tokens per turn with a full buffer6001,8004,2009,00040,00001234turn 1turn 10turn 60,pasted logsEvery turn re-sends the whole transcript, so cost grows until the 128,000-token window is hit.
Buffer memory is perfect for a short chat and a cost incident for a long one, because nothing is ever dropped.

What you need to know

What it did

Python
# Legacy — for reading old code onlyfrom langchain_classic.memory import ConversationBufferMemoryfrom langchain_classic.chains import ConversationChainmemory = ConversationBufferMemory(return_messages=True)chain = ConversationChain(llm=llm, memory=memory)chain.invoke({"input": "My router model is TP-200."})chain.invoke({"input": "How do I reset it?"})memory.load_memory_variables({})   # {"history": [HumanMessage, AIMessage, ...]}
  • It stored turns on the memory object.
  • Before each call, load_memory_variables returned the history under the key history (set by memory_key).
  • After each call, save_context appended the new input and output.
  • return_messages=True returned message objects; False returned one long string like "Human: ...\nAI: ...".

Its family of legacy classes

Legacy classWhat it keptCurrent equivalent
ConversationBufferMemoryEverythingCheckpointer + thread_id
ConversationBufferWindowMemoryLast k turnstrim_messages or a before_model middleware
ConversationTokenBufferMemoryLast N tokenstrim_messages(max_tokens=...)
ConversationSummaryMemoryA running summarySummarizationMiddleware
ConversationSummaryBufferMemorySummary + recent turnsSummarizationMiddleware with keep=
ConversationEntityMemoryFacts per entityLangGraph store + structured extraction

Why it was replaced

  • State inside the object. One memory object per chain meant one chain per user, or users sharing history by accident.
  • Poor fit with LCEL. It predates runnables; | pipelines have no memory slot.
  • Lost tool steps. With agents it saved the input and final output, not the intermediate tool calls.
  • No persistence or resume. It lived in process memory unless you wired a history backend.

The current equivalent

Python
from langchain.agents import create_agentfrom langgraph.checkpoint.postgres import PostgresSaverwith PostgresSaver.from_conn_string(DB_URI) as saver:    bot = create_agent(model, tools=[], checkpointer=saver)    bot.invoke({"messages": [("user", "My router model is TP-200.")]},               {"configurable": {"thread_id": "cust-42"}})

Every message in the thread is kept and sent, just like the buffer, but saved per thread in Postgres.

When "keep everything" is fine

Short sessions — under about 20 turns, or where total history stays well under the context window — such as a form-filling assistant. For long or open-ended chats, add a limit.

A real-life example

Think of a court stenographer who reads the entire transcript aloud before every new question. For a 10-minute hearing that is fine. For a three-day trial, most of the day is spent re-reading, and eventually there is no time left for the new question. The analogy breaks in one way: the stenographer gets tired, but the model gets charged — every re-read transcript is billed as input tokens.

In production: a help-centre bot using ConversationBufferMemory worked in testing, where chats were 4 to 6 turns. In production, 5% of chats ran past 60 turns — customers pasting router logs — and hit the 128,000-token context limit with an error. Those same chats were also the most expensive: their prompts averaged 40,000 tokens per turn. Moving to a checkpointer with SummarizationMiddleware (trigger at 6,000 tokens) removed the errors and cut the cost of long chats by about 80%.

Follow-up questions to expect

  • "What is memory_key?" — The prompt variable name the history is injected under; it defaults to history and must match the prompt.
  • "Buffer versus window memory?" — Buffer keeps everything; window keeps only the last k turns, which caps cost but forgets older facts.
  • "What should I use instead today?" — A LangGraph checkpointer, through create_agent or your own StateGraph. RunnableWithMessageHistory, the in-between answer from 2024, is deprecated too.