Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 6: Persistent Memory Requirement


What you need to know

The scenario: users expect the assistant to remember the conversation after a page refresh, a deploy, or a week away.

Two kinds of memory

Short-term: thread state

  • The messages and state of one conversation
  • Saved by the checkpointer after each step
  • Keyed by thread_id
  • Grows fast; trim or summarise it

Long-term: a store

  • Facts about a user across conversations
  • Saved in a store (LangGraph's Store API or your own table)
  • Keyed by user_id, often searched by embedding
  • Small and curated: preferences, decisions

Mixing them is how you end up with a 40,000-token prompt that still forgets the user's name.

Wiring it up

Python
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaverfrom langgraph.store.postgres.aio import AsyncPostgresStoreasync with AsyncPostgresSaver.from_conn_string(DB_URI) as saver, \           AsyncPostgresStore.from_conn_string(DB_URI) as store:    await saver.setup(); await store.setup()             # create tables once    graph = builder.compile(checkpointer=saver, store=store)    cfg = {"configurable": {"thread_id": f"conv:{conversation_id}", "user_id": user_id}}    await graph.ainvoke({"messages": [user_msg]}, config=cfg)

Nodes can read and write the store (for example store.aput(("users", user_id), "prefs", {...})) to save long-term facts, and search it at the start of each run.

Operational rules

  1. Bound the context — a trim or summarise node runs when history passes a token threshold.
  2. Curate long-term facts — write only stable facts ("prefers Hindi", "account type: business"), with a timestamp, and let newer facts replace older ones.
  3. Set retention — a TTL for old threads, matching your privacy policy.
  4. Delete for real — "delete my data" removes checkpoints and store entries for that user; conversation state is personal data.
  5. Protect secrets — never put tokens or raw PII in state unencrypted; the checkpointer writes it to disk.

A real-life example

Scenario, numbers made up. A mutual-fund advisory chatbot runs on three pods with MemorySaver. Every deploy wipes conversations mid-flow, and returning users must repeat their risk profile each time.

The team switches to a Postgres checkpointer keyed by conversation and a Postgres store keyed by user. A trim node summarises history above 6,000 tokens. Long-term memory stores six fields such as risk profile, goals and preferred language. Conversations now survive deploys, prompt size per turn falls by about 60% for long chats, and returning users are greeted with their saved goals. A deletion job clears checkpoints and store entries within 24 hours of a request.

Follow-up questions to expect

  • "Why not keep the whole history in the prompt?" — Cost and quality: long histories are expensive and bury the important facts. Summaries plus curated facts work better.
  • "What goes into long-term memory, and who decides?" — A node or background job extracts candidate facts; keep only stable, useful ones, and let users see and delete them.
  • "How do you debug a bad answer from last week?" — Load that thread's checkpoint history and replay from the step before the mistake.