Context Management and Memory

Vector-Based Memory — Recall Beyond the Window


A personal assistant has been in daily use for eight months. Roughly 240 sessions, about 30 exchanges each — call it 7,200 exchanges and 3.5 million tokens of transcript.

In March the user mentioned, once, in passing, that they had relocated to Berlin. In November they type: "what's a good weekend trip from here?"

The assistant has no idea where "here" is. The March session ended long ago. A sliding window covers the last twenty minutes. A running summary covers the whole eight months in 600 tokens, which means it says something like "the user has discussed travel, work projects and cooking" — a compression ratio of nearly 6,000 to 1, and at that ratio nothing specific survives.

Here is the important observation. The Berlin fact is roughly ten tokens long. The problem was never that it was too big to keep. The problem was that it was buried in 3.5 million tokens and there was no way to find it. Compression was the wrong tool entirely. What was needed was search.

Recall by selection, not compressionSplit thesessioninto exchangesEmbed each oneStorevector plusoriginal textEmbed thenew questionInject the topfew, labelled3.5 million tokens of transcript never enter the prompt; only the handful that match the question do.
Summarisation shrinks everything a little, retrieval keeps a few things whole — only the second survives eight months of use.

Compression versus selection

Every short-term strategy is a compression strategy: take everything, make it smaller, send the smaller thing. Compression scales badly because the amount you must throw away grows with the history.

Compression (window / summary)Selection (vector memory)
What is storedOnly what fits in the promptEverything, in an external store
What is sentA shrunken version of everythingA full-fidelity version of a few relevant things
Cost per turnGrows with conversation lengthRoughly constant — top-k results
Fails whenHistory is long; old facts matterRetrieval misses; the query is vague
Survives a new session?No — state lives in the processYes — state lives in a database

Put numbers on the second column. Suppose you extract about one durable fact per eight exchanges, giving 900 memories over those eight months, each roughly 40 tokens. Retrieve the best 5 for a given question and you inject 200 tokens — against 3.5 million tokens of raw transcript, a selection ratio of about 17,500 to 1, with the retrieved 200 tokens exact rather than blurred.

Summarisation makes the whole past smaller. Retrieval makes the relevant past findable. Long-lived assistants need the second, because the past keeps growing and the prompt does not.

Why keyword search is not enough

The obvious way to search is by keyword. Store the messages, run a text index, look for matches. It works right up until the user's words differ from the stored words, which is most of the time.

Text
Stored memory:  "User relocated to Berlin in March for a new job."Query:          "Where do I live?"Shared words:   none that matterKeyword score:  ~0

"Relocated", "Berlin", "live", "where" — no lexical overlap at all, yet the memory is exactly the right answer. Keyword search matches strings. You need something that matches meaning.

Embeddings: just enough to use them

An embedding is a list of numbers that represents a piece of text's meaning. An embedding model takes text in and returns a fixed-length vector out — commonly 384, 768, 1024 or 1536 numbers — and it is trained so that texts with similar meanings produce vectors that point in similar directions.

"Direction" is the operative word. The measure is cosine similarity: the cosine of the angle between two vectors, which ignores their length and looks only at orientation.

cos⁡(θ)=a⋅b∥a∥ ∥b∥=∑iaibi∑iai2 ∑ibi2\cos(\theta) = \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\|\,\|\mathbf{b}\|} = \frac{\sum_i a_i b_i}{\sqrt{\sum_i a_i^2}\,\sqrt{\sum_i b_i^2}}

It ranges from −1 (opposite) through 0 (unrelated) to 1 (identical direction). Real embeddings live in hundreds of dimensions, but the arithmetic is the same in three, so let us do it in three where you can watch it happen.

Worked example

Imagine three interpretable dimensions: location-ness, food-ness, time-ness.

Text
query   q = "Where do I live?"            [0.9, 0.1, 0.2]memory  A = "Relocated to Berlin in March" [0.8, 0.0, 0.4]memory  B = "Prefers oat milk in coffee"   [0.1, 0.9, 0.1]

Dot products first:

  • q⋅A=(0.9)(0.8)+(0.1)(0.0)+(0.2)(0.4)=0.72+0+0.08=0.80\mathbf{q}\cdot\mathbf{A} = (0.9)(0.8) + (0.1)(0.0) + (0.2)(0.4) = 0.72 + 0 + 0.08 = 0.80
  • q⋅B=(0.9)(0.1)+(0.1)(0.9)+(0.2)(0.1)=0.09+0.09+0.02=0.20\mathbf{q}\cdot\mathbf{B} = (0.9)(0.1) + (0.1)(0.9) + (0.2)(0.1) = 0.09 + 0.09 + 0.02 = 0.20

Then the magnitudes:

  • ∥q∥=0.81+0.01+0.04=0.86=0.9274\|\mathbf{q}\| = \sqrt{0.81 + 0.01 + 0.04} = \sqrt{0.86} = 0.9274
  • ∥A∥=0.64+0.00+0.16=0.80=0.8944\|\mathbf{A}\| = \sqrt{0.64 + 0.00 + 0.16} = \sqrt{0.80} = 0.8944
  • ∥B∥=0.01+0.81+0.01=0.83=0.9110\|\mathbf{B}\| = \sqrt{0.01 + 0.81 + 0.01} = \sqrt{0.83} = 0.9110

And the similarities:

  • cos⁡(q,A)=0.80/(0.9274×0.8944)=0.80/0.8295=0.964\cos(\mathbf{q},\mathbf{A}) = 0.80 / (0.9274 \times 0.8944) = 0.80 / 0.8295 = \mathbf{0.964}
  • cos⁡(q,B)=0.20/(0.9274×0.9110)=0.20/0.8449=0.237\cos(\mathbf{q},\mathbf{B}) = 0.20 / (0.9274 \times 0.9110) = 0.20 / 0.8449 = \mathbf{0.237}

The Berlin memory scores 0.964; the oat milk memory scores 0.237. Note that not a single word is shared between the query and the winning memory. That is the whole point — the match happened in meaning space, not in string space.

Choosing an embedding model

Embeddings come from a dedicated model, not from the chat model. The main axes are dimensionality, quality, cost and whether it runs locally.

OptionTypical dimsRuns locally?Use when
Small open model (MiniLM class)384Yes, on CPUPrototypes, privacy constraints, very high volume, cost near zero
Mid open model (BGE / E5 class)768–1024Yes, GPU preferredGood quality without an external API call; self-hosted deployments
Hosted general-purpose1024–1536NoDefault choice; strong quality, priced in cents per million tokens
Hosted domain-tuned (code, legal, finance)1024–1536NoYour corpus is specialised and general models confuse near-synonyms

Two rules matter more than the choice itself.

Query and document must use the same model. Vectors from different models are not comparable — they live in different spaces and cosine similarity between them is meaningless noise that still returns confident-looking numbers.

Changing the model means re-embedding everything. There is no migration path. Store the model name and version alongside every vector so that when you do switch, you can tell which rows are stale.

What the dimensions cost you

A 1536-dimension vector stored as 32-bit floats is 1536×4=6,1441536 \times 4 = 6{,}144 bytes, about 6 KB. For one user with 900 memories that is 5.4 MB — nothing. For 10,000 users it is 9 million vectors and about 54 GB, which is a real infrastructure decision.

ConfigurationBytes per vector9M vectors
1536 dims, float326,14454.0 GB
768 dims, float323,07227.0 GB
1536 dims, int8 quantised1,53613.8 GB
768 dims, int8 quantised7686.9 GB

Halving dimensions or quantising to 8-bit integers typically costs a small amount of retrieval accuracy for a four-fold reduction in storage and a matching speed-up in search. Measure before you decide; on memory records — short, distinctive, well-separated texts — the accuracy loss is usually negligible.

The architecture

Vector memory has two paths that run at different times and can be tuned independently.

Text
WRITE PATH  (after each exchange, off the critical path)  conversation turn        |        v  extract durable facts  -->  nothing worth keeping?  -->  stop        |        v  check for near-duplicate / contradiction        |        v  embed  -->  store {text, vector, user_id, type, timestamp, confidence}READ PATH  (before each model call, on the critical path)  user message        |        v  embed the query        |        v  search: filter by user_id, then top-k by cosine similarity        |        v  drop results below a score floor        |        v  inject survivors into the prompt as a labelled block

The read path is in the user's latency budget, so it must be fast: one embedding call (tens of milliseconds) plus one index query (single-digit milliseconds). The write path is not, so it can afford an extra model call to extract facts properly.

Implementation

Embedding and storing

Python
import timefrom uuid import uuid4import chromadbfrom chromadb.utils import embedding_functionsclient = chromadb.PersistentClient(path="./memory_store")embedder = embedding_functions.SentenceTransformerEmbeddingFunction(    model_name="all-MiniLM-L6-v2"   # 384 dims, runs on CPU)memories = client.get_or_create_collection(    name="user_memories",    embedding_function=embedder,    configuration={"hnsw": {"space": "cosine"}},   # Chroma's default is l2)def remember(user_id: str, text: str, kind: str, confidence: float = 1.0):    memories.add(        documents=[text],        ids=[f"{user_id}:{uuid4().hex}"],        metadatas=[{            "user_id": user_id,          # non-negotiable, see below            "kind": kind,                # fact | preference | decision | event            "created_at": int(time.time()),            "confidence": confidence,        }],    )

Retrieving

Python
def recall(user_id: str, query: str, k: int = 5, floor: float = 0.35):    res = memories.query(        query_texts=[query],        n_results=k,        where={"user_id": user_id},      # ALWAYS filter, never post-filter    )    out = []    for text, meta, dist in zip(res["documents"][0],                                res["metadatas"][0],                                res["distances"][0]):        similarity = 1.0 - dist          # cosine distance -> similarity        if similarity >= floor:            out.append({"text": text, "meta": meta, "score": similarity})    return out

The user_id filter belongs inside the query, not in a loop that runs afterwards. Filtering after retrieval means asking for the top 5 across all users and hoping they belong to the right one — and when they do not, one user reads another user's private memories.

Injecting into the prompt

Python
def build_prompt(user_id, user_message, recent_turns):    hits = recall(user_id, user_message, k=5)    if hits:        block = "\n".join(f"- {h['text']}" for h in hits)        memory_note = (            "Things you know about this user from earlier conversations:\n"            f"{block}\n"            "Use these only where relevant. If a memory conflicts with what "            "the user says now, the user is right."        )    else:        memory_note = ""    return SYSTEM_PROMPT + ("\n\n" + memory_note if memory_note else ""), recent_turns

Three details in that block are doing real work. Memories are labelled as memories, so the model does not treat them as things the user just said. The model is told they may be irrelevant, which stops it forcing them into every reply. And the model is told the live conversation wins over stored memory, which is how you avoid an assistant arguing that the user still lives in Munich.

What to store, and at what size

The single biggest quality lever is granularity. Store the wrong unit and no amount of retrieval tuning saves you.

Unit storedTokens eachRetrieval behaviourVerdict
Whole session transcript~15,000Every session looks vaguely similar; injecting one blows the budgetNever
Raw message~150Retrieves "sure, that works" and "let me check" — conversational noise with no contentPoor
Exchange (user + assistant)~480Workable; carries context but is diluted by fillerAcceptable
Extracted fact~40One idea per vector; scores are sharp and injection is cheapBest

Extraction costs an extra model call per exchange, but it is the difference between a memory store and a transcript dump. Run it on a small fast model with a tight schema:

Python
EXTRACT_PROMPT = """From the exchange below, extract durable facts worthremembering for months. A durable fact is stable, specific and about the user.Extract: identity, location, role, relationships, hard constraints,preferences, decisions taken, ongoing goals.Do NOT extract: pleasantries, questions, anything true only today,anything the assistant said about itself.Return a JSON list of objects with keys: text, kind, confidence.Return [] if there is nothing durable. Prefer [] over guessing.EXCHANGE:{exchange}"""

Where your provider supports it, pass the list shape as a JSON schema through its structured-output mode instead of asking for JSON in prose; the reply is then guaranteed to parse, and your code only has to judge the content.

"Prefer [] over guessing" is not decoration. Without it, extractors invent a fact from every exchange, and a store full of low-value assertions is worse than an empty one — it crowds real memories out of the top 5.

Keeping the store healthy

A memory store degrades on its own. Three maintenance problems appear in every deployment.

Duplicates, and retrieval collapse

A user who mentions being vegetarian in twelve separate sessions produces twelve near-identical vectors. All twelve score around 0.95 against a food question, so your top-5 returns five copies of the same sentence and the budget-limit fact, the allergy and the schedule constraint never appear.

This is retrieval collapse: a duplicated fact monopolises the result set. Fix it on write.

Python
DUPLICATE_THRESHOLD = 0.93def remember_deduped(user_id, text, kind, confidence=1.0):    similar = recall(user_id, text, k=1, floor=DUPLICATE_THRESHOLD)    if similar:        # Same fact restated: refresh recency and confidence, do not add a row.        bump(similar[0], confidence)        return "merged"    remember(user_id, text, kind, confidence)    return "added"

Pick the threshold empirically. Around 0.93–0.96 catches restatements; below about 0.90 you start merging genuinely different facts, which is a worse error because it silently deletes information.

Staleness, and time-aware decay

"User is training for the Berlin marathon" was true last March and is misleading now. Decay old memories by multiplying their similarity by an exponential factor:

score=similarity×e−λ⋅age_days,λ=ln⁡2t1/2\text{score} = \text{similarity} \times e^{-\lambda \cdot \text{age\_days}}, \qquad \lambda = \frac{\ln 2}{t_{1/2}}

With a half-life t1/2t_{1/2} of 90 days, λ=0.6931/90=0.0077\lambda = 0.6931/90 = 0.0077. A memory 180 days old — two half-lives — is multiplied by e−0.0077×180=e−1.386=0.25e^{-0.0077 \times 180} = e^{-1.386} = 0.25. So a stale memory scoring 0.90 on similarity ends up at 0.90×0.25=0.2250.90 \times 0.25 = 0.225 and falls below a 0.35 floor. A fresh memory scoring 0.70 beats it easily.

The crucial refinement: not everything should decay. Apply decay by kind.

KindHalf-lifeReasoning
Identity (name, language, timezone)Never decaysStable for years
Hard constraint (allergy, accessibility need)Never decaysForgetting this causes harm
Preference (likes dark mode, prefers brevity)~365 daysDrifts slowly
Project or goal~90 daysProjects end
Transient state (currently travelling)~7 daysTrue for days, not months

Decaying an allergy on the same schedule as a dark-mode preference is how a memory system becomes a safety problem.

Metadata, and the filters you will wish you had

Every memory should carry, at minimum: user_id, kind, created_at, last_seen_at, confidence, source_session_id and the embedding_model version. Retrofitting these onto a populated store is painful because the information is gone — you cannot reconstruct which session a memory came from after the fact.

FieldWhat it unlocks
user_idIsolation. Without it you have a data breach, not a feature.
kindPer-kind decay rates and targeted queries ("all hard constraints")
created_at / last_seen_atDecay, and distinguishing "stated once" from "reconfirmed nine times"
confidenceRanking, and a floor below which a memory is never injected
source_session_idDeletion requests, and auditing where a wrong memory came from
embedding_modelKnowing which rows to re-embed after a model change

Failure modes

SymptomCauseFix
Assistant references facts about someone elseMissing or post-applied user_id filterFilter inside the query; add an integration test that asserts isolation
Same fact retrieved five timesNo write-time deduplicationMerge near-duplicates above ~0.93
Assistant insists on outdated informationNo decay, no contradiction handlingPer-kind half-lives; supersede on conflict; tell the model the live turn wins
Irrelevant memories shoehorned into every replyNo score floor, or a prompt that implies memories must be usedSet a floor around 0.35–0.4 and say explicitly that memories may be ignored
Retrieval returns "okay, sounds good"Storing raw messages rather than extracted factsExtract; return an empty list when nothing is durable
Every query returns 0.99 similarityQuery and documents embedded with different models, or distance metric mismatched to the indexPin one model; set the index metric to cosine explicitly
Store works in test, useless in productionTested with 20 memories; at 2,000 the good ones are outrankedEvaluate on a realistic store size with a labelled query set

Building this for real

Vector memory is a database, and it deserves the discipline you would give any other database.

Instrument retrieval from day one. Log, for every turn: the query, the top-k results with their scores, and which ones cleared the floor. Without that log, "the bot forgot my address" is unanswerable — you cannot tell whether the memory was never written, was written and not retrieved, was retrieved and filtered out, or was retrieved and ignored by the model. Those are four different bugs with four different fixes.

Make deletion work. Users will ask you to forget things, and in many jurisdictions they have a legal right to. That means every memory must be traceable to a user and a source, and "delete everything for user X" must be a single indexed operation rather than a scan.

Budget the injection. Five memories at 40 tokens is 200 tokens per turn — about 0.06 cents at 3 dollars per million. Ten memories at 400 tokens each is 4,000 tokens per turn, which is twenty times the cost and measurably worse answers, because the model now has to find the relevant fact inside a wall of marginal ones. Fewer, sharper memories beat more, longer ones almost every time.

Know what your framework already gives you. If you build on LangGraph, its long-term memory store (InMemoryStore for development, PostgresStore in production) implements the pieces above: memories live under a namespace such as (user_id, "memories"), store.put writes them, and store.search(namespace, query=...) does the semantic lookup when the store is created with an embedding index. The namespace does the job of the user_id filter. Extraction, deduplication, decay and deletion are still yours to design.

Decide what you refuse to store. A memory system that silently records everything a user says is a liability. Payment details, health information, credentials and third-party personal data should be excluded at the extraction step, by an explicit rule, not by hoping the model chooses well. Write the exclusion list before you write the extractor.