Course Content
LangChain Mastery
7 sections · 109 lessons
How do you optimize memory for long conversations in LangChain?
What you need to know
The shape of a good long-conversation prompt
[system prompt + tool schemas] stable ~2,000 tokens (cacheable prefix)[user profile / key facts] small ~200 tokens[summary of older turns] refreshed ~300 tokens[retrieved older details, if any] optional ~500 tokens[last 6-10 messages, verbatim] changing ~2,000 tokens[new user message]Total: about 5,000 tokens whether the chat is 10 turns or 200.
The tiers, and the tool for each
| Tier | Purpose | LangChain 1.x |
|---|---|---|
| Recent turns verbatim | Pronouns, follow-ups | keep= in SummarizationMiddleware, or trim_messages |
| Running summary | Long-range context | SummarizationMiddleware(trigger=("tokens", N)) |
| Structured facts | IDs, decisions, preferences | Custom state_schema fields; LangGraph store |
| Retrieval of old detail | Rare look-backs | Store with an embedding index, store.search |
| Hard budget | Never overflow | trim_messages in wrap_model_call |
Cost optimisations beyond size
- Prompt caching. Providers such as Anthropic, OpenAI and Google give a large discount on input tokens that repeat an earlier prompt's prefix exactly. Put stable content first; do not put timestamps or changing facts at the top of the system prompt, or the cache never hits.
- Summarise with a cheap model, and only when the threshold is crossed — not every turn.
- Shrink old tool results. A 3,000-token search result from 20 turns ago can be replaced with "searched help centre for router reset; used article 118".
- Stream the answer so long prompts do not feel slow.
Measure it
In LangSmith, chart prompt tokens per turn against turn number for a sample of long chats. A flat line after the threshold means the design works. Also track answer quality on long chats specifically — for example, how often the bot asks for information the user already gave.
A real-life example
A broadband company's help-centre bot had a long tail: 8% of chats ran over 50 turns, mostly multi-day fault investigations. Before optimising, prompt tokens rose steadily to about 30,000 by turn 60, and those chats made up 35% of the model bill.
The team applied the tiers: SummarizationMiddleware at 6,000 tokens keeping 10 messages; the account number, router model and open ticket ID moved into state fields; old diagnostic outputs replaced by one-line notes; and the system prompt reordered so the date and account facts came after the stable instructions. Prompt tokens now level off at about 5,500. The cache hit rate on the stable prefix rose from near zero to about 70%, and the long-chat share of the bill fell from 35% to 11%. In a review of 60 long chats, "asked again for information already given" fell from 12 to 2.
Follow-up questions to expect
- "What gets lost with this design?" — Exact wording of old turns; keep the full transcript in storage for audits, even though the model sees a compressed view.
- "How many recent turns should you keep?" — Usually 6 to 10 messages; test on your own follow-up questions, since tool-heavy chats need more.
- "Why does the order of the prompt matter?" — Prompt caching works on an exact matching prefix; anything that changes near the top breaks the cache for everything after it.