LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you implement memory with vector stores in LangChain?


Recalling an old fact without replaying old chatsSave fact: chargeroverheated, ORD-7713Embed intonamespacememories, userNew message:replacementis hot tooSearch only thisuser's namespaceInject top 4plus thelast 6 turnsThe fact was 5 weeks old; the prompt stayed near 3,000 tokens.
Semantic search reaches far back at a fixed prompt size, but it needs the recent turns alongside it because it ignores time.

What you need to know

How it works

  1. Write — after each turn (or each extracted fact), embed the text and store it with metadata: user_id, thread_id, timestamp.
  2. Search — on a new message, embed it and find the top k similar memories for this user only.
  3. Inject — add those memories to the prompt as a labelled block, plus the last 4 to 6 turns verbatim.
  4. Answer — the model sees relevant old context and recent context, at a bounded size.

Option A: LangGraph store with semantic search

Python
from langchain.embeddings import init_embeddingsfrom langgraph.store.postgres import PostgresStorewith PostgresStore.from_conn_string(DB_URI, index={        "embed": init_embeddings(settings.embedding_model),   # e.g. a 1536-dim model        "dims": settings.embedding_dims, "fields": ["text"]}) as store:    store.setup()    def save_memory(user_id: str, text: str, mem_id: str):        store.put(("memories", user_id), mem_id, {"text": text})    def recall(user_id: str, query: str, k: int = 4) -> list[str]:        hits = store.search(("memories", user_id), query=query, limit=k)        return [h.value["text"] for h in hits]

The namespace ("memories", user_id) does the per-user filtering: a search in one user's namespace cannot return another user's items. Inside an agent, tools reach the same store through runtime.store, and middleware through request.runtime.store.

Option B: a vector store retriever

Python
retriever = vectorstore.as_retriever(search_kwargs={"k": 4, "filter": {"user_id": uid}})docs = retriever.invoke(question)

The exact filter syntax depends on the vector store (Chroma, pgvector, Pinecone, Qdrant each differ slightly). Without it, the search covers every user's history.

What to store

  • Whole turns — easy, but noisy: "ok thanks" becomes a memory.
  • Extracted facts or short summaries — better: "Customer's X200 laptop was replaced under warranty on 3 March." One sentence, high value.

The recency problem

Similarity search ignores time. Ask "what did I just say?" and it may return a turn from last month that uses similar words, and miss the message right before. That is why vector memory is combined with the last few turns kept verbatim, and why a timestamp in each memory helps the model judge freshness.

A real-life example

An electronics store's support assistant serves repeat customers. A customer writes: "The replacement charger you sent is also getting hot." The current chat has no mention of an earlier charger.

Vector memory search, filtered to that customer, returns a fact saved 5 weeks earlier: "Customer reported X200 charger overheating; replacement shipped 12 Aug, order ORD-7713." The assistant replies with that context, and escalates as a repeat fault instead of starting from scratch. Prompt size stays at about 3,000 tokens however many past chats the customer has had.

During testing, one early version forgot the user_id filter, and a search for "charger overheating" returned another customer's order number. The team added a test that creates two users with similar memories and asserts that neither can retrieve the other's.

Follow-up questions to expect

  • "Why not keep the whole history instead?" — Cost and context size grow without limit; vector search keeps the prompt fixed while still reaching old facts.
  • "What are the weaknesses?" — It misses recent context, can return plausible but irrelevant memories, and needs strict per-user filtering.
  • "How do you delete a user's memories?" — Delete their namespace in the store, or delete by user_id metadata in the vector store.