Course Content
LangChain Mastery
7 sections · 109 lessons
How do you implement memory with vector stores in LangChain?
What you need to know
How it works
- Write — after each turn (or each extracted fact), embed the text and store it with metadata:
user_id,thread_id, timestamp. - Search — on a new message, embed it and find the top k similar memories for this user only.
- Inject — add those memories to the prompt as a labelled block, plus the last 4 to 6 turns verbatim.
- Answer — the model sees relevant old context and recent context, at a bounded size.
Option A: LangGraph store with semantic search
1from langchain.embeddings import init_embeddings2from langgraph.store.postgres import PostgresStore34with PostgresStore.from_conn_string(DB_URI, index={5 "embed": init_embeddings(settings.embedding_model), # e.g. a 1536-dim model6 "dims": settings.embedding_dims, "fields": ["text"]}) as store:7 store.setup()89 def save_memory(user_id: str, text: str, mem_id: str):10 store.put(("memories", user_id), mem_id, {"text": text})1112 def recall(user_id: str, query: str, k: int = 4) -> list[str]:13 hits = store.search(("memories", user_id), query=query, limit=k)14 return [h.value["text"] for h in hits]The namespace ("memories", user_id) does the per-user filtering: a search in one user's namespace cannot return another user's items. Inside an agent, tools reach the same store through runtime.store, and middleware through request.runtime.store.
Option B: a vector store retriever
retriever = vectorstore.as_retriever(search_kwargs={"k": 4, "filter": {"user_id": uid}})docs = retriever.invoke(question)The exact filter syntax depends on the vector store (Chroma, pgvector, Pinecone, Qdrant each differ slightly). Without it, the search covers every user's history.
What to store
- Whole turns — easy, but noisy: "ok thanks" becomes a memory.
- Extracted facts or short summaries — better: "Customer's X200 laptop was replaced under warranty on 3 March." One sentence, high value.
The recency problem
Similarity search ignores time. Ask "what did I just say?" and it may return a turn from last month that uses similar words, and miss the message right before. That is why vector memory is combined with the last few turns kept verbatim, and why a timestamp in each memory helps the model judge freshness.
A real-life example
An electronics store's support assistant serves repeat customers. A customer writes: "The replacement charger you sent is also getting hot." The current chat has no mention of an earlier charger.
Vector memory search, filtered to that customer, returns a fact saved 5 weeks earlier: "Customer reported X200 charger overheating; replacement shipped 12 Aug, order ORD-7713." The assistant replies with that context, and escalates as a repeat fault instead of starting from scratch. Prompt size stays at about 3,000 tokens however many past chats the customer has had.
During testing, one early version forgot the user_id filter, and a search for "charger overheating" returned another customer's order number. The team added a test that creates two users with similar memories and asserts that neither can retrieve the other's.
Follow-up questions to expect
- "Why not keep the whole history instead?" — Cost and context size grow without limit; vector search keeps the prompt fixed while still reaching old facts.
- "What are the weaknesses?" — It misses recent context, can return plausible but irrelevant memories, and needs strict per-user filtering.
- "How do you delete a user's memories?" — Delete their namespace in the store, or delete by
user_idmetadata in the vector store.