Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to optimize LangChain memory usage.
What you need to know
"Memory usage" has two meanings, and a strong answer covers both:
- Prompt memory — how much history is sent to the model each turn. It drives cost, latency and context-length errors.
- Process memory — how much RAM your service uses to hold conversations. It drives out-of-memory crashes.
For an LCEL chain: trim before the prompt
1from langchain_core.messages import trim_messages2from langchain_core.messages.utils import count_tokens_approximately3from langchain_core.runnables import RunnablePassthrough45def with_bounded_history(chain, *, max_tokens: int = 2000):6 """Trim input["history"] to max_tokens before running chain."""7 trimmer = trim_messages(8 max_tokens=max_tokens, strategy="last",9 token_counter=count_tokens_approximately,10 include_system=True, start_on="human",11 )12 return RunnablePassthrough.assign(13 history=lambda x: trimmer.invoke(x["history"])) | chain1415bounded = with_bounded_history(prompt | llm, max_tokens=2000)16bounded.invoke({"history": past_messages, "question": "And for next year?"})- Called without messages,
trim_messages(...)returns a runnable trimmer. strategy="last"keeps the newest messages;include_system=Truekeeps instructions;start_on="human"avoids starting with an orphan AI or tool message.count_tokens_approximatelyis fast; pass the chat model astoken_counterwhen you need exact counts.
For an agent: summarise old turns
1from langchain.agents import create_agent2from langchain.agents.middleware import SummarizationMiddleware34agent = create_agent(5 model, tools=TOOLS, checkpointer=checkpointer,6 middleware=[SummarizationMiddleware(7 model=small_model, # a cheap model writes the summary8 trigger=("tokens", 4000), # summarise when history passes 4,000 tokens9 keep=("messages", 20), # keep the latest 20 messages verbatim10 )],11)Summaries keep older facts that trimming would drop, at the cost of an extra model call when the trigger fires — and some risk that details are lost. Keep critical facts (ids, dates, decisions) in structured state as well.
Process memory
- Don't keep
sessions = {}in a module; it grows forever and vanishes on restart. - Use a database-backed checkpointer (Postgres) or Redis for history, and delete or expire old threads with a scheduled cleanup.
- Load only the thread you need for the current request.
Verify
Log input tokens per turn. The line should rise, then level off around your budget. If it keeps rising, trimming or summarisation isn't running.
A real-life example
An HR assistant's average chat is 12 turns, but some managers keep one thread open for weeks while planning team leave — 300+ turns. Those threads cost 20 times more per message and eventually failed with context-length errors. Separately, a worker's RAM grew by about 1 GB per day because histories were cached in a dict.
The team added SummarizationMiddleware (trigger at 4,000 tokens, keep 20 messages) and moved history to the Postgres checkpointer with a nightly job deleting threads idle for 30 days. Input tokens per message now level off at about 4,500. Cost for long threads dropped by 85%, and worker memory stayed flat across a week of monitoring.
Follow-up questions to expect
- "How do you choose the budget?" — Leave room for the system prompt, retrieved context, tools and the answer inside the context window; then pick the smallest budget that passes multi-turn eval tests.
- "Why not always summarise?" — Each summary costs a call and can drift; for short chats trimming is enough.
- "What about the old memory classes?" —
ConversationBufferWindowMemoryandConversationSummaryMemorydid this in legacy chains; they are inlangchain-classicnow.