Course Content
LangChain Mastery
7 sections · 109 lessons
Implement a function to merge multiple memory contexts.
What you need to know
Typical sources
| Source | Where it lives | Size | Changes |
|---|---|---|---|
| User profile and preferences | LangGraph store | Small | Rarely |
| Facts from past conversations | Store with semantic search | Varies | Grows over time |
| Summary of this thread's older turns | Thread state | Small | Every so often |
| Recent turns of this thread | Checkpointer messages | Large | Every turn |
Rules that make merging work
- Label every block. The model must know where a fact came from: "Profile", "From past chats", "This chat".
- Budget per source, then check the total. If you trim only the merged list, a long current chat can push out the profile entirely, or a big retrieval result can push out recent turns.
- Fixed order. Stable content first (system rules, profile), changing content last. This also helps provider prompt caching, which reuses an unchanged prefix.
- State precedence. "If the user says something that contradicts stored facts, trust the user and say what changed."
- Keep conversation turns as messages. Do not flatten them into text; the model uses the human and AI roles.
The code
1from langchain.agents import create_agent2from langchain.agents.middleware import wrap_model_call3from langchain.messages import SystemMessage4from langchain_core.messages import trim_messages5from langchain_core.messages.utils import count_tokens_approximately as count67def fit(lines: list[str], budget: int) -> str:8 out, used = [], 09 for line in lines: # lines are in priority order10 t = count([("user", line)])11 if used + t > budget:12 break13 out.append(line); used += t14 return "\n".join(out)1516@wrap_model_call17def merged_memory(request, handler):18 uid, store = request.runtime.context.user_id, request.runtime.store19 profile = store.get(("users", uid), "profile")20 question = request.messages[-1].text21 past = store.search(("users", uid, "facts"), query=question, limit=10)22 system = (f"{BASE_RULES}\n\n## Profile\n{fit([str(profile.value)] if profile else [], 300)}"23 f"\n\n## From past chats\n{fit([p.value['text'] for p in past], 600)}"24 "\n\nIf this chat contradicts the blocks above, trust this chat.")25 recent = trim_messages(request.messages, max_tokens=3000,26 token_counter="approximate", strategy="last", start_on="human")27 return handler(request.override(system_message=SystemMessage(system), messages=recent))2829agent = create_agent(model, tools=TOOLS, middleware=[merged_memory],30 store=store, checkpointer=saver, context_schema=Ctx)Each source gets its own budget — 300 tokens for the profile, 600 for past facts, 3,000 for recent turns — so the total prompt is predictable. request.override changes only what this model call sees; the stored thread and store are unchanged.
In your own StateGraph the same idea works in the node that calls the chain: build profile, past_facts and a trimmed message list, and inject each into its own labelled slot of the ChatPromptTemplate.
A real-life example
A law firm's contract assistant merges three sources for each associate: their current matter's profile (client, jurisdiction, contract types), notes saved from earlier research sessions, and the current chat. In the first version everything was concatenated and trimmed as one list. On long research chats the matter profile was trimmed away first, and the assistant started citing Delhi High Court cases for a Singapore-law contract.
With per-source budgets (profile always kept, 600 tokens of notes, 4,000 of chat) and the rule "the current chat overrides notes", jurisdiction errors in a 40-question review dropped from 7 to 0. The prompt size also became stable at around 6,000 tokens instead of swinging between 2,000 and 15,000.
Follow-up questions to expect
- "Why not trim everything together?" — The biggest or newest source wins and small but important context, like the profile, disappears; per-source budgets protect it.
- "How do you handle conflicting facts?" — Label sources, add timestamps, and give an explicit precedence rule in the system prompt; update the store when the user corrects a fact.
- "Where does merging run in LangChain 1.x?" — In middleware (
dynamic_promptorwrap_model_call), so it runs before every model call without changing stored state.