Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 6: Persistent Memory Requirement
What you need to know
The scenario: users expect the assistant to remember the conversation after a page refresh, a deploy, or a week away.
Two kinds of memory
Short-term: thread state
- The messages and state of one conversation
- Saved by the checkpointer after each step
- Keyed by
thread_id - Grows fast; trim or summarise it
Long-term: a store
- Facts about a user across conversations
- Saved in a store (LangGraph's Store API or your own table)
- Keyed by
user_id, often searched by embedding - Small and curated: preferences, decisions
Mixing them is how you end up with a 40,000-token prompt that still forgets the user's name.
Wiring it up
1from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver2from langgraph.store.postgres.aio import AsyncPostgresStore34async with AsyncPostgresSaver.from_conn_string(DB_URI) as saver, \5 AsyncPostgresStore.from_conn_string(DB_URI) as store:6 await saver.setup(); await store.setup() # create tables once7 graph = builder.compile(checkpointer=saver, store=store)8 cfg = {"configurable": {"thread_id": f"conv:{conversation_id}", "user_id": user_id}}9 await graph.ainvoke({"messages": [user_msg]}, config=cfg)Nodes can read and write the store (for example store.aput(("users", user_id), "prefs", {...})) to save long-term facts, and search it at the start of each run.
Operational rules
- Bound the context — a trim or summarise node runs when history passes a token threshold.
- Curate long-term facts — write only stable facts ("prefers Hindi", "account type: business"), with a timestamp, and let newer facts replace older ones.
- Set retention — a TTL for old threads, matching your privacy policy.
- Delete for real — "delete my data" removes checkpoints and store entries for that user; conversation state is personal data.
- Protect secrets — never put tokens or raw PII in state unencrypted; the checkpointer writes it to disk.
A real-life example
Scenario, numbers made up. A mutual-fund advisory chatbot runs on three pods with MemorySaver. Every deploy wipes conversations mid-flow, and returning users must repeat their risk profile each time.
The team switches to a Postgres checkpointer keyed by conversation and a Postgres store keyed by user. A trim node summarises history above 6,000 tokens. Long-term memory stores six fields such as risk profile, goals and preferred language. Conversations now survive deploys, prompt size per turn falls by about 60% for long chats, and returning users are greeted with their saved goals. A deletion job clears checkpoints and store entries within 24 hours of a request.
Follow-up questions to expect
- "Why not keep the whole history in the prompt?" — Cost and quality: long histories are expensive and bury the important facts. Summaries plus curated facts work better.
- "What goes into long-term memory, and who decides?" — A node or background job extracts candidate facts; keep only stable, useful ones, and let users see and delete them.
- "How do you debug a bad answer from last week?" — Load that thread's checkpoint history and replay from the step before the mistake.