LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to optimize LangChain memory usage.


The two kinds of memory a chat app spendsPrompt memory• History sent to the model each turn• Drives cost, latency, context errors• Fix: trim_messages or summarisation• Budget in tokens, not messagesProcess memory• Conversations held in worker RAM• Drives out-of-memory crashes• Fix: database checkpointer• Expire idle threads on a schedule
A strong answer names both: one grows the bill every turn, the other grows the worker until it is killed.

What you need to know

"Memory usage" has two meanings, and a strong answer covers both:

  • Prompt memory — how much history is sent to the model each turn. It drives cost, latency and context-length errors.
  • Process memory — how much RAM your service uses to hold conversations. It drives out-of-memory crashes.

For an LCEL chain: trim before the prompt

Python
from langchain_core.messages import trim_messagesfrom langchain_core.messages.utils import count_tokens_approximatelyfrom langchain_core.runnables import RunnablePassthroughdef with_bounded_history(chain, *, max_tokens: int = 2000):    """Trim input["history"] to max_tokens before running chain."""    trimmer = trim_messages(        max_tokens=max_tokens, strategy="last",        token_counter=count_tokens_approximately,        include_system=True, start_on="human",    )    return RunnablePassthrough.assign(        history=lambda x: trimmer.invoke(x["history"])) | chainbounded = with_bounded_history(prompt | llm, max_tokens=2000)bounded.invoke({"history": past_messages, "question": "And for next year?"})
  • Called without messages, trim_messages(...) returns a runnable trimmer.
  • strategy="last" keeps the newest messages; include_system=True keeps instructions; start_on="human" avoids starting with an orphan AI or tool message.
  • count_tokens_approximately is fast; pass the chat model as token_counter when you need exact counts.

For an agent: summarise old turns

Python
from langchain.agents import create_agentfrom langchain.agents.middleware import SummarizationMiddlewareagent = create_agent(    model, tools=TOOLS, checkpointer=checkpointer,    middleware=[SummarizationMiddleware(        model=small_model,               # a cheap model writes the summary        trigger=("tokens", 4000),        # summarise when history passes 4,000 tokens        keep=("messages", 20),           # keep the latest 20 messages verbatim    )],)

Summaries keep older facts that trimming would drop, at the cost of an extra model call when the trigger fires — and some risk that details are lost. Keep critical facts (ids, dates, decisions) in structured state as well.

Process memory

  • Don't keep sessions = {} in a module; it grows forever and vanishes on restart.
  • Use a database-backed checkpointer (Postgres) or Redis for history, and delete or expire old threads with a scheduled cleanup.
  • Load only the thread you need for the current request.

Verify

Log input tokens per turn. The line should rise, then level off around your budget. If it keeps rising, trimming or summarisation isn't running.

A real-life example

An HR assistant's average chat is 12 turns, but some managers keep one thread open for weeks while planning team leave — 300+ turns. Those threads cost 20 times more per message and eventually failed with context-length errors. Separately, a worker's RAM grew by about 1 GB per day because histories were cached in a dict.

The team added SummarizationMiddleware (trigger at 4,000 tokens, keep 20 messages) and moved history to the Postgres checkpointer with a nightly job deleting threads idle for 30 days. Input tokens per message now level off at about 4,500. Cost for long threads dropped by 85%, and worker memory stayed flat across a week of monitoring.

Follow-up questions to expect

  • "How do you choose the budget?" — Leave room for the system prompt, retrieved context, tools and the answer inside the context window; then pick the smallest budget that passes multi-turn eval tests.
  • "Why not always summarise?" — Each summary costs a call and can drift; for short chats trimming is enough.
  • "What about the old memory classes?" — ConversationBufferWindowMemory and ConversationSummaryMemory did this in legacy chains; they are in langchain-classic now.