LangGraph Agents

Course Content

LangGraph Agents

7 sections · 49 lessons

How do you prevent state from growing unbounded?


What the model actually receives on turn 41systemsummaryq38a38q39a39q400123456replaces 36old messageswindow startson a human turnOld messages were removed with RemoveMessage, so the checkpoint shrank too.
Trimming only shrinks the prompt; summarising and deleting is what shrinks the saved state.

What you need to know

Two different problems

ProblemSymptomFix
Prompt sizeCost and latency rise each turn; context limit errorsTrim or summarise what the model sees
Checkpoint sizeSlow writes, big databaseSmaller state, retention policy, delete old threads

Trim before the model call

Python
from langchain_core.messages import trim_messagesdef call_model(state):    window = trim_messages(        state["messages"], max_tokens=4000, strategy="last",        token_counter="approximate", include_system=True, start_on="human",    )    return {"messages": [llm.invoke(window)]}

This trims only what is sent; the full history stays in state. start_on="human" keeps the window from starting with an orphan tool result.

Summarise and delete

A summarise node runs when the history passes a threshold. It writes one summary message and returns RemoveMessage(id=m.id) for the old ones, so state itself shrinks. With create_agent, SummarizationMiddleware does this for you with a token or message trigger.

Other controls

  • Bounded reducers — lambda old, new: (old + new)[-50:].
  • Reset scratch keys — {"drafts": Overwrite([]), "attempts": 0} at the end of a loop.
  • References — keep document_ids, fetch text inside the node that needs it.
  • Long-term Store — move "prefers Hindi, has a dual-band router" out of the transcript.
  • Retention — delete finished threads with checkpointer.delete_thread(thread_id), or set a TTL on the hosted Agent Server.

A real-life example

A broadband support bot keeps one thread per customer "for continuity". After six months, one busy customer's thread holds 1,900 messages. Each new turn sends about 90,000 tokens, costs roughly 40 times a fresh chat, and the checkpoint row is 3 MB. The team changes three things: a new thread per support case, a summary message when a case passes 40 messages, and facts like router model saved to the Store. Average prompt size falls to 3,500 tokens and the checkpoint table shrinks by 85% after a retention job deletes closed cases older than 90 days.

Follow-up questions to expect

  • "Trim or summarise?" — Trim for chat where old turns rarely matter; summarise when earlier details, like an order number, still matter.
  • "Won't trimming break tool calls?" — It can if an AI tool call is kept without its tool result. Use start_on="human" and keep whole turns.
  • "Does trimming reduce checkpoint size?" — No. Trimming the prompt leaves state unchanged; only RemoveMessage or new threads shrink it.