Course Content
LangChain Mastery
7 sections · 109 lessons
How do you implement token-limited memory in LangChain?
What you need to know
Why tokens and not turns
Turn counts are a poor limit. "Thanks!" is 2 tokens; a pasted error log can be 3,000. A token budget directly controls what matters: context-window fit, cost and latency.
trim_messages
1from langchain_core.messages import trim_messages23trimmed = trim_messages(4 messages,5 max_tokens=3000,6 strategy="last", # keep the newest messages7 token_counter=model, # exact count via the model; or "approximate"8 include_system=True, # always keep the first SystemMessage9 start_on="human", # result must start with a human turn10 end_on=("human", "tool"), # and end where the model is expected to reply11)| Parameter | Why it matters |
|---|---|
strategy="last" | Recent turns matter most for follow-ups |
token_counter | A model object counts with its real tokenizer (slower, may call an API); "approximate" uses about 4 characters per token (fast, local) |
include_system=True | Without it, the system prompt is the first thing dropped |
start_on="human" | Many providers reject a history that starts with an AI message or a tool result without its call |
allow_partial=False (default) | Never cuts a message in half |
It returns a new list; your stored history is not changed.
Where to put it
In your own graph node — trim just before calling the chain:
1def respond(state: MessagesState):2 recent = trim_messages(state["messages"], max_tokens=3000,3 token_counter="approximate", strategy="last", start_on="human")4 return {"messages": [chain.invoke({"messages": recent})]}In a create_agent app — trim in middleware:
1from langchain.agents.middleware import wrap_model_call23@wrap_model_call4def token_window(request, handler):5 kept = trim_messages(request.messages, max_tokens=3000, token_counter="approximate",6 strategy="last", start_on="human")7 return handler(request.override(messages=kept))This changes only what the model receives on this call; the full history stays in the checkpointer, so nothing is lost for audits or later summarisation. Use before_model with RemoveMessage instead if you want the stored history itself to shrink.
Choosing the budget
Add up what must fit: system prompt (say 800 tokens), tool schemas (1,500), retrieved documents (4,000), the answer (1,000), and a safety margin. What is left is the history budget. On a 32,000-token model this might be 8,000; on a 200,000-token model you could keep much more, but you still pay for every token on every call, so a smaller budget is often right anyway.
A real-life example
An electronics store's product Q&A assistant allowed 12,000 tokens of history. Most chats were short, but customers comparing laptops often pasted full spec sheets. Those chats cost 5 times the average and, oddly, got worse answers: the model mixed up specs from different pasted sheets.
The team set a 4,000-token history budget with trim_messages, plus a rule to store pasted spec sheets as a short structured summary. Cost per long chat fell by 60%. In a check of 50 comparison chats, spec mix-ups dropped from 9 to 3, because the model was no longer reading three old spec sheets at once.
Follow-up questions to expect
- "Exact or approximate token counting?" — Approximate for trimming, where being off by 10% is fine; exact when you are close to a hard provider limit.
- "What happens to trimmed messages?" — With a request-level trim, they stay in storage; the model just does not see them. Combine with a summary if old facts matter.
- "What was
ConversationTokenBufferMemory?" — The legacy class that kept only the most recent messages undermax_token_limit.