LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to handle memory errors in LangChain?


What you need to know

1. Context overflow

Every model has a context limit. A long chat eventually exceeds it and the provider returns a 400 error ("context length exceeded"). Retrying does not help. Keep the history under a token budget:

Python
from langchain_core.messages import trim_messagesdef fit_history(messages, max_tokens: int = 3000):    return trim_messages(        messages, max_tokens=max_tokens,        strategy="last",            # keep the most recent turns        token_counter="approximate",        include_system=True,        # never drop the system prompt        start_on="human",           # don't start with an orphan AI or tool message    )

token_counter can also be the chat model itself (exact but slower) or count_tokens_approximately. For agents, SummarizationMiddleware(model=..., trigger=("tokens", 4000), keep=("messages", 20)) summarises older turns automatically when the history grows too large.

2. History store failure

In current LangChain, an agent's conversation lives in a LangGraph checkpointer — InMemorySaver for tests, PostgresSaver (package langgraph-checkpoint-postgres) in production — keyed by a thread_id. If the database is down, the call fails. Degrade to a stateless answer instead:

Python
import loggingimport psycopgfrom langchain.agents import create_agentlog = logging.getLogger(__name__)stateful = create_agent(model, tools=TOOLS, checkpointer=checkpointer)stateless = create_agent(model, tools=TOOLS)          # same agent, no memorydef chat(thread_id: str, text: str):    request = {"messages": [{"role": "user", "content": text}]}    try:        return stateful.invoke(request, {"configurable": {"thread_id": thread_id}})    except psycopg.OperationalError:        log.warning("memory_store_down", extra={"thread_id": thread_id})        return stateless.invoke(request)              # answer this turn without history

One caution: if the store fails after a tool with side effects has run, the stateless retry runs it again. Tools that change things (filing leave, booking) need idempotency keys before you add this fallback.

The older message-history classes (RunnableWithMessageHistory, InMemoryChatMessageHistory, RedisChatMessageHistory) still exist, but recent langchain-core releases mark RunnableWithMessageHistory and InMemoryChatMessageHistory as deprecated in favour of LangGraph persistence (checkpointers). The same rule applies to them: catch the failure where the history is read, log it, and continue.

3. Process memory

A module-level sessions = {} grows with every user and is lost on restart. Use Redis or Postgres with a TTL, and keep only the recent window in the worker.

A real-life example

An HR bot supports leave questions over long chats. Two incidents hit in the same week.

First, a manager pasted a 40-page policy into chat, and every later message failed with a context-length error, because the paste stayed in history. The fix was fit_history with a 6,000-token budget plus a 4,000-character limit on single messages.

Second, a database failover took 90 seconds. Every chat returned HTTP 500, because loading the conversation from the checkpointer raised a connection error. After the change above, users got answers without history during the failover — "What is my leave balance?" still worked, because it calls the HR API, not memory. The only visible effect was that follow-up questions like "and next year?" lacked context for those 90 seconds.

Follow-up questions to expect

  • "Trim or summarise?" — Trim is free and predictable; summarising keeps older facts but costs a model call and can drift. Many apps trim first and summarise only past a threshold.
  • "How do you avoid cutting a tool call from its result?" — Use start_on="human" and trimming helpers that respect message pairs; an orphan tool message causes provider errors.
  • "What is legacy here?" — ConversationBufferMemory and the other memory classes are in langchain-classic; current apps use agent checkpointers with trimming or SummarizationMiddleware.