LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you use LangChain to implement entity-based memory?


Facts stored per entity in the user's namespacepincode41101412 AugmodelX200,bought March3 MarlanguageHindi12 AugAttributeValueLast updatedcustomerdevicecustomerWritten by a save_fact tool that reads user_id from runtime context, never from the model.
A fact stored as a field survives any trim or summary, and the user can see and correct it.

What you need to know

Why not just keep the transcript

A transcript answers "what was said"; entity memory answers "what do we know about X". After a summary or trim, the fact that the customer's delivery pincode is 411014 may be gone from the transcript. Stored as {"customer": {"pincode": "411014"}}, it is kept until it changes, and it costs a few tokens to include.

Pattern 1: extract after each turn

Python
from pydantic import BaseModel, Fieldclass Fact(BaseModel):    entity: str = Field(description="Who or what, e.g. 'customer' or 'order ORD-1042'")    attribute: str = Field(description="e.g. 'pincode', 'preferred_language'")    value: strclass Facts(BaseModel):    facts: list[Fact]extractor = small_model.with_structured_output(Facts)def remember(store, user_id: str, user_text: str) -> None:    found = extractor.invoke("Extract stable facts the user states about themselves, "                             "their orders or devices. Ignore guesses.\n\n" + user_text)    for f in found.facts:        ns = ("users", user_id, "entities")        current = store.get(ns, f.entity)        data = current.value if current else {}        data[f.attribute] = f.value               # newer fact overwrites older        store.put(ns, f.entity, data)

Run it after the turn, ideally in the background, so it does not add latency to the reply.

Pattern 2: let the agent write memories with a tool

Python
from dataclasses import dataclassfrom langchain.tools import tool, ToolRuntime@dataclassclass Ctx:    user_id: str@tooldef save_fact(entity: str, attribute: str, value: str, runtime: ToolRuntime[Ctx]) -> str:    """Save a stable fact the user told you, e.g. their pincode or router model."""    ns = ("users", runtime.context.user_id, "entities")    item = runtime.store.get(ns, entity)    runtime.store.put(ns, entity, {**(item.value if item else {}), attribute: value})    return "Saved."

Pass store= and context_schema=Ctx to create_agent, and invoke with context=Ctx(user_id=...). The user_id comes from your backend, not from the model, so the agent cannot write into another user's memory.

Reading it back

Before the model call, load the user's entities (for example in a dynamic_prompt middleware) and add them to the system prompt as a short block: Known facts: customer.pincode=411014; router=TP-200. For many entities, create the store with an embedding index and use store.search(ns, query=question, limit=5) to include only relevant ones.

Risks

  • Wrong extractions persist. Let later facts overwrite earlier ones, store a timestamp, and let users see and correct what is remembered.
  • Sensitive data. Decide which attributes you are allowed to keep; do not store card numbers or health details just because the user typed them.

A real-life example

An electronics store's support assistant kept being asked the same things: "Which model do you have?", "What's your pincode?". Customers came back days later in new chats, and a new thread started with no knowledge of them.

The team added save_fact plus a dynamic_prompt that loads the user's saved entities. A returning customer who says "the laptop's fan is loud again" now gets "Is this the X200 you bought in March, delivered to 411014?" In a month, repeat-question complaints fell by about 40%. They also added a "What we remember about you" page with a delete button, which about 2% of users used.

Follow-up questions to expect

  • "How is this different from summary memory?" — A summary is free text about the conversation; entity memory is structured key-value facts per thing, which you can query, update and delete individually.
  • "What if two facts conflict?" — Keep the newest with a timestamp, or store both and ask the user to confirm.
  • "Is there a library for this?" — LangChain's LangMem library builds memory extraction and management on top of the LangGraph store; you can also write it yourself as above.