LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Implement a function to merge multiple memory contexts.


One prompt built from four sources, each with its own budgetThis chat:last turns, 3,000From pastchats, 600Profile, 300Rules andprecedencetopbottomTrimmed as one list, the profile was dropped first and Delhi cases appeared for a Singapore contract.
Per-source budgets keep small but critical context from being pushed out by whichever source is longest.

What you need to know

Typical sources

SourceWhere it livesSizeChanges
User profile and preferencesLangGraph storeSmallRarely
Facts from past conversationsStore with semantic searchVariesGrows over time
Summary of this thread's older turnsThread stateSmallEvery so often
Recent turns of this threadCheckpointer messagesLargeEvery turn

Rules that make merging work

  • Label every block. The model must know where a fact came from: "Profile", "From past chats", "This chat".
  • Budget per source, then check the total. If you trim only the merged list, a long current chat can push out the profile entirely, or a big retrieval result can push out recent turns.
  • Fixed order. Stable content first (system rules, profile), changing content last. This also helps provider prompt caching, which reuses an unchanged prefix.
  • State precedence. "If the user says something that contradicts stored facts, trust the user and say what changed."
  • Keep conversation turns as messages. Do not flatten them into text; the model uses the human and AI roles.

The code

Python
from langchain.agents import create_agentfrom langchain.agents.middleware import wrap_model_callfrom langchain.messages import SystemMessagefrom langchain_core.messages import trim_messagesfrom langchain_core.messages.utils import count_tokens_approximately as countdef fit(lines: list[str], budget: int) -> str:    out, used = [], 0    for line in lines:                         # lines are in priority order        t = count([("user", line)])        if used + t > budget:            break        out.append(line); used += t    return "\n".join(out)@wrap_model_calldef merged_memory(request, handler):    uid, store = request.runtime.context.user_id, request.runtime.store    profile = store.get(("users", uid), "profile")    question = request.messages[-1].text    past = store.search(("users", uid, "facts"), query=question, limit=10)    system = (f"{BASE_RULES}\n\n## Profile\n{fit([str(profile.value)] if profile else [], 300)}"              f"\n\n## From past chats\n{fit([p.value['text'] for p in past], 600)}"              "\n\nIf this chat contradicts the blocks above, trust this chat.")    recent = trim_messages(request.messages, max_tokens=3000,                           token_counter="approximate", strategy="last", start_on="human")    return handler(request.override(system_message=SystemMessage(system), messages=recent))agent = create_agent(model, tools=TOOLS, middleware=[merged_memory],                     store=store, checkpointer=saver, context_schema=Ctx)

Each source gets its own budget — 300 tokens for the profile, 600 for past facts, 3,000 for recent turns — so the total prompt is predictable. request.override changes only what this model call sees; the stored thread and store are unchanged.

In your own StateGraph the same idea works in the node that calls the chain: build profile, past_facts and a trimmed message list, and inject each into its own labelled slot of the ChatPromptTemplate.

A real-life example

A law firm's contract assistant merges three sources for each associate: their current matter's profile (client, jurisdiction, contract types), notes saved from earlier research sessions, and the current chat. In the first version everything was concatenated and trimmed as one list. On long research chats the matter profile was trimmed away first, and the assistant started citing Delhi High Court cases for a Singapore-law contract.

With per-source budgets (profile always kept, 600 tokens of notes, 4,000 of chat) and the rule "the current chat overrides notes", jurisdiction errors in a 40-question review dropped from 7 to 0. The prompt size also became stable at around 6,000 tokens instead of swinging between 2,000 and 15,000.

Follow-up questions to expect

  • "Why not trim everything together?" — The biggest or newest source wins and small but important context, like the profile, disappears; per-source budgets protect it.
  • "How do you handle conflicting facts?" — Label sources, add timestamps, and give an explicit precedence rule in the system prompt; update the store when the user corrects a fact.
  • "Where does merging run in LangChain 1.x?" — In middleware (dynamic_prompt or wrap_model_call), so it runs before every model call without changing stored state.