Course Content
LangGraph Agents
7 sections · 49 lessons
How do you model short-term vs long-term memory using state and persistence?
What you need to know
Short-term (checkpointer)
- Scope: one thread
- Holds: messages, working keys
- Written: automatically every super-step
- Class:
InMemorySaver,SqliteSaver,PostgresSaver
Long-term (Store)
- Scope: across threads, per user or org
- Holds: facts, preferences, past outcomes
- Written: explicitly by your code
- Class:
InMemoryStore,PostgresStore
1from dataclasses import dataclass2from langgraph.runtime import Runtime34@dataclass5class Ctx:6 customer_id: str78def load_profile(state, runtime: Runtime[Ctx]):9 ns = ("customers", runtime.context.customer_id)10 hits = runtime.store.search(ns, query=state["messages"][-1].content, limit=3)11 return {"profile_notes": [h.value["note"] for h in hits]}1213def save_fact(state, runtime: Runtime[Ctx]):14 ns = ("customers", runtime.context.customer_id)15 runtime.store.put(ns, "language", {"note": "Prefers Hindi"})16 return {}1718graph = builder.compile(checkpointer=checkpointer, store=store)put(namespace, key, value) writes a JSON document; get reads one by key; search lists or, with an embedding index configured on the store, finds semantically similar items.
Kinds of long-term memory
- Semantic — facts: "has a dual-band router", "account is business tier".
- Episodic — what happened: "technician visit on 12 Sept fixed the issue".
- Procedural — how to behave: "this customer wants short answers" — often stored as instructions loaded into the system prompt.
Deciding what to write
Write a memory when a fact is stable, useful later, and confirmed. Writing every sentence creates noise and conflicting facts.
A real-life example
A broadband provider's support bot serves 2 million customers. Each support case is a new thread: thread_id=f"{customer_id}:{case_id}", in PostgresSaver. At the end of a case, a save_memory node writes two or three confirmed facts to PostgresStore under ("customers", customer_id). When customer 4471 opens a new chat a week later, load_profile searches that namespace and finds "prefers Hindi", "router: dual-band, model XR-200" and "last issue: fibre cut, fixed 12 Sept". The bot greets them in Hindi and skips four diagnostic questions. Average handling time drops from 7 minutes to about 4.
Follow-up questions to expect
- "Why not keep everything in one long thread?" — The prompt and checkpoint grow without bound, and unrelated old cases pollute new ones.
- "How do you stop wrong memories?" — Write only confirmed facts, include a timestamp, overwrite by key when facts change, and let users see or delete their memories.
- "Is the Store a vector database?" — It is a key-value document store that can add an embedding index for semantic search; for very large corpora a dedicated vector database may still fit better.