LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you implement token-limited memory in LangChain?


Keeping the newest messages under a 3,000-token budget800sys1202,600paste90450602100123456kept:include_systemdropped wholestrategy last walks back from the newest message; start_on human keeps the kept slice valid.
Trimming by tokens drops the one huge paste instead of twenty short turns, which a message-count limit would get backwards.

What you need to know

Why tokens and not turns

Turn counts are a poor limit. "Thanks!" is 2 tokens; a pasted error log can be 3,000. A token budget directly controls what matters: context-window fit, cost and latency.

trim_messages

Python
from langchain_core.messages import trim_messagestrimmed = trim_messages(    messages,    max_tokens=3000,    strategy="last",          # keep the newest messages    token_counter=model,      # exact count via the model; or "approximate"    include_system=True,      # always keep the first SystemMessage    start_on="human",         # result must start with a human turn    end_on=("human", "tool"), # and end where the model is expected to reply)
ParameterWhy it matters
strategy="last"Recent turns matter most for follow-ups
token_counterA model object counts with its real tokenizer (slower, may call an API); "approximate" uses about 4 characters per token (fast, local)
include_system=TrueWithout it, the system prompt is the first thing dropped
start_on="human"Many providers reject a history that starts with an AI message or a tool result without its call
allow_partial=False (default)Never cuts a message in half

It returns a new list; your stored history is not changed.

Where to put it

In your own graph node — trim just before calling the chain:

Python
def respond(state: MessagesState):    recent = trim_messages(state["messages"], max_tokens=3000,                           token_counter="approximate", strategy="last", start_on="human")    return {"messages": [chain.invoke({"messages": recent})]}

In a create_agent app — trim in middleware:

Python
from langchain.agents.middleware import wrap_model_call@wrap_model_calldef token_window(request, handler):    kept = trim_messages(request.messages, max_tokens=3000, token_counter="approximate",                         strategy="last", start_on="human")    return handler(request.override(messages=kept))

This changes only what the model receives on this call; the full history stays in the checkpointer, so nothing is lost for audits or later summarisation. Use before_model with RemoveMessage instead if you want the stored history itself to shrink.

Choosing the budget

Add up what must fit: system prompt (say 800 tokens), tool schemas (1,500), retrieved documents (4,000), the answer (1,000), and a safety margin. What is left is the history budget. On a 32,000-token model this might be 8,000; on a 200,000-token model you could keep much more, but you still pay for every token on every call, so a smaller budget is often right anyway.

A real-life example

An electronics store's product Q&A assistant allowed 12,000 tokens of history. Most chats were short, but customers comparing laptops often pasted full spec sheets. Those chats cost 5 times the average and, oddly, got worse answers: the model mixed up specs from different pasted sheets.

The team set a 4,000-token history budget with trim_messages, plus a rule to store pasted spec sheets as a short structured summary. Cost per long chat fell by 60%. In a check of 50 comparison chats, spec mix-ups dropped from 9 to 3, because the model was no longer reading three old spec sheets at once.

Follow-up questions to expect

  • "Exact or approximate token counting?" — Approximate for trimming, where being off by 10% is fine; exact when you are close to a hard provider limit.
  • "What happens to trimmed messages?" — With a request-level trim, they stay in storage; the model just does not see them. Combine with a summary if old facts matter.
  • "What was ConversationTokenBufferMemory?" — The legacy class that kept only the most recent messages under max_token_limit.