Course Content
Context Management and Memory
3 sections · 6 lessons
Vector-Based Memory — Recall Beyond the Window
A personal assistant has been in daily use for eight months. Roughly 240 sessions, about 30 exchanges each — call it 7,200 exchanges and 3.5 million tokens of transcript.
In March the user mentioned, once, in passing, that they had relocated to Berlin. In November they type: "what's a good weekend trip from here?"
The assistant has no idea where "here" is. The March session ended long ago. A sliding window covers the last twenty minutes. A running summary covers the whole eight months in 600 tokens, which means it says something like "the user has discussed travel, work projects and cooking" — a compression ratio of nearly 6,000 to 1, and at that ratio nothing specific survives.
Here is the important observation. The Berlin fact is roughly ten tokens long. The problem was never that it was too big to keep. The problem was that it was buried in 3.5 million tokens and there was no way to find it. Compression was the wrong tool entirely. What was needed was search.
Compression versus selection
Every short-term strategy is a compression strategy: take everything, make it smaller, send the smaller thing. Compression scales badly because the amount you must throw away grows with the history.
| Compression (window / summary) | Selection (vector memory) | |
|---|---|---|
| What is stored | Only what fits in the prompt | Everything, in an external store |
| What is sent | A shrunken version of everything | A full-fidelity version of a few relevant things |
| Cost per turn | Grows with conversation length | Roughly constant — top-k results |
| Fails when | History is long; old facts matter | Retrieval misses; the query is vague |
| Survives a new session? | No — state lives in the process | Yes — state lives in a database |
Put numbers on the second column. Suppose you extract about one durable fact per eight exchanges, giving 900 memories over those eight months, each roughly 40 tokens. Retrieve the best 5 for a given question and you inject 200 tokens — against 3.5 million tokens of raw transcript, a selection ratio of about 17,500 to 1, with the retrieved 200 tokens exact rather than blurred.
Summarisation makes the whole past smaller. Retrieval makes the relevant past findable. Long-lived assistants need the second, because the past keeps growing and the prompt does not.
Why keyword search is not enough
The obvious way to search is by keyword. Store the messages, run a text index, look for matches. It works right up until the user's words differ from the stored words, which is most of the time.
Stored memory: "User relocated to Berlin in March for a new job."Query: "Where do I live?"Shared words: none that matterKeyword score: ~0"Relocated", "Berlin", "live", "where" — no lexical overlap at all, yet the memory is exactly the right answer. Keyword search matches strings. You need something that matches meaning.
Embeddings: just enough to use them
An embedding is a list of numbers that represents a piece of text's meaning. An embedding model takes text in and returns a fixed-length vector out — commonly 384, 768, 1024 or 1536 numbers — and it is trained so that texts with similar meanings produce vectors that point in similar directions.
"Direction" is the operative word. The measure is cosine similarity: the cosine of the angle between two vectors, which ignores their length and looks only at orientation.
It ranges from −1 (opposite) through 0 (unrelated) to 1 (identical direction). Real embeddings live in hundreds of dimensions, but the arithmetic is the same in three, so let us do it in three where you can watch it happen.
Worked example
Imagine three interpretable dimensions: location-ness, food-ness, time-ness.
query q = "Where do I live?" [0.9, 0.1, 0.2]memory A = "Relocated to Berlin in March" [0.8, 0.0, 0.4]memory B = "Prefers oat milk in coffee" [0.1, 0.9, 0.1]Dot products first:
- q⋅A=(0.9)(0.8)+(0.1)(0.0)+(0.2)(0.4)=0.72+0+0.08=0.80
- q⋅B=(0.9)(0.1)+(0.1)(0.9)+(0.2)(0.1)=0.09+0.09+0.02=0.20
Then the magnitudes:
- ∥q∥=0.81+0.01+0.04=0.86=0.9274
- ∥A∥=0.64+0.00+0.16=0.80=0.8944
- ∥B∥=0.01+0.81+0.01=0.83=0.9110
And the similarities:
- cos(q,A)=0.80/(0.9274×0.8944)=0.80/0.8295=0.964
- cos(q,B)=0.20/(0.9274×0.9110)=0.20/0.8449=0.237
The Berlin memory scores 0.964; the oat milk memory scores 0.237. Note that not a single word is shared between the query and the winning memory. That is the whole point — the match happened in meaning space, not in string space.
Choosing an embedding model
Embeddings come from a dedicated model, not from the chat model. The main axes are dimensionality, quality, cost and whether it runs locally.
| Option | Typical dims | Runs locally? | Use when |
|---|---|---|---|
| Small open model (MiniLM class) | 384 | Yes, on CPU | Prototypes, privacy constraints, very high volume, cost near zero |
| Mid open model (BGE / E5 class) | 768–1024 | Yes, GPU preferred | Good quality without an external API call; self-hosted deployments |
| Hosted general-purpose | 1024–1536 | No | Default choice; strong quality, priced in cents per million tokens |
| Hosted domain-tuned (code, legal, finance) | 1024–1536 | No | Your corpus is specialised and general models confuse near-synonyms |
Two rules matter more than the choice itself.
Query and document must use the same model. Vectors from different models are not comparable — they live in different spaces and cosine similarity between them is meaningless noise that still returns confident-looking numbers.
Changing the model means re-embedding everything. There is no migration path. Store the model name and version alongside every vector so that when you do switch, you can tell which rows are stale.
What the dimensions cost you
A 1536-dimension vector stored as 32-bit floats is 1536×4=6,144 bytes, about 6 KB. For one user with 900 memories that is 5.4 MB — nothing. For 10,000 users it is 9 million vectors and about 54 GB, which is a real infrastructure decision.
| Configuration | Bytes per vector | 9M vectors |
|---|---|---|
| 1536 dims, float32 | 6,144 | 54.0 GB |
| 768 dims, float32 | 3,072 | 27.0 GB |
| 1536 dims, int8 quantised | 1,536 | 13.8 GB |
| 768 dims, int8 quantised | 768 | 6.9 GB |
Halving dimensions or quantising to 8-bit integers typically costs a small amount of retrieval accuracy for a four-fold reduction in storage and a matching speed-up in search. Measure before you decide; on memory records — short, distinctive, well-separated texts — the accuracy loss is usually negligible.
The architecture
Vector memory has two paths that run at different times and can be tuned independently.
WRITE PATH (after each exchange, off the critical path) conversation turn | v extract durable facts --> nothing worth keeping? --> stop | v check for near-duplicate / contradiction | v embed --> store {text, vector, user_id, type, timestamp, confidence}READ PATH (before each model call, on the critical path) user message | v embed the query | v search: filter by user_id, then top-k by cosine similarity | v drop results below a score floor | v inject survivors into the prompt as a labelled blockThe read path is in the user's latency budget, so it must be fast: one embedding call (tens of milliseconds) plus one index query (single-digit milliseconds). The write path is not, so it can afford an extra model call to extract facts properly.
Implementation
Embedding and storing
1import time2from uuid import uuid434import chromadb5from chromadb.utils import embedding_functions67client = chromadb.PersistentClient(path="./memory_store")89embedder = embedding_functions.SentenceTransformerEmbeddingFunction(10 model_name="all-MiniLM-L6-v2" # 384 dims, runs on CPU11)1213memories = client.get_or_create_collection(14 name="user_memories",15 embedding_function=embedder,16 configuration={"hnsw": {"space": "cosine"}}, # Chroma's default is l217)1819def remember(user_id: str, text: str, kind: str, confidence: float = 1.0):20 memories.add(21 documents=[text],22 ids=[f"{user_id}:{uuid4().hex}"],23 metadatas=[{24 "user_id": user_id, # non-negotiable, see below25 "kind": kind, # fact | preference | decision | event26 "created_at": int(time.time()),27 "confidence": confidence,28 }],29 )Retrieving
1def recall(user_id: str, query: str, k: int = 5, floor: float = 0.35):2 res = memories.query(3 query_texts=[query],4 n_results=k,5 where={"user_id": user_id}, # ALWAYS filter, never post-filter6 )7 out = []8 for text, meta, dist in zip(res["documents"][0],9 res["metadatas"][0],10 res["distances"][0]):11 similarity = 1.0 - dist # cosine distance -> similarity12 if similarity >= floor:13 out.append({"text": text, "meta": meta, "score": similarity})14 return outThe
user_idfilter belongs inside the query, not in a loop that runs afterwards. Filtering after retrieval means asking for the top 5 across all users and hoping they belong to the right one — and when they do not, one user reads another user's private memories.
Injecting into the prompt
1def build_prompt(user_id, user_message, recent_turns):2 hits = recall(user_id, user_message, k=5)34 if hits:5 block = "\n".join(f"- {h['text']}" for h in hits)6 memory_note = (7 "Things you know about this user from earlier conversations:\n"8 f"{block}\n"9 "Use these only where relevant. If a memory conflicts with what "10 "the user says now, the user is right."11 )12 else:13 memory_note = ""1415 return SYSTEM_PROMPT + ("\n\n" + memory_note if memory_note else ""), recent_turnsThree details in that block are doing real work. Memories are labelled as memories, so the model does not treat them as things the user just said. The model is told they may be irrelevant, which stops it forcing them into every reply. And the model is told the live conversation wins over stored memory, which is how you avoid an assistant arguing that the user still lives in Munich.
What to store, and at what size
The single biggest quality lever is granularity. Store the wrong unit and no amount of retrieval tuning saves you.
| Unit stored | Tokens each | Retrieval behaviour | Verdict |
|---|---|---|---|
| Whole session transcript | ~15,000 | Every session looks vaguely similar; injecting one blows the budget | Never |
| Raw message | ~150 | Retrieves "sure, that works" and "let me check" — conversational noise with no content | Poor |
| Exchange (user + assistant) | ~480 | Workable; carries context but is diluted by filler | Acceptable |
| Extracted fact | ~40 | One idea per vector; scores are sharp and injection is cheap | Best |
Extraction costs an extra model call per exchange, but it is the difference between a memory store and a transcript dump. Run it on a small fast model with a tight schema:
1EXTRACT_PROMPT = """From the exchange below, extract durable facts worth2remembering for months. A durable fact is stable, specific and about the user.34Extract: identity, location, role, relationships, hard constraints,5preferences, decisions taken, ongoing goals.6Do NOT extract: pleasantries, questions, anything true only today,7anything the assistant said about itself.89Return a JSON list of objects with keys: text, kind, confidence.10Return [] if there is nothing durable. Prefer [] over guessing.1112EXCHANGE:13{exchange}"""Where your provider supports it, pass the list shape as a JSON schema through its structured-output mode instead of asking for JSON in prose; the reply is then guaranteed to parse, and your code only has to judge the content.
"Prefer [] over guessing" is not decoration. Without it, extractors invent a fact from every exchange, and a store full of low-value assertions is worse than an empty one — it crowds real memories out of the top 5.
Keeping the store healthy
A memory store degrades on its own. Three maintenance problems appear in every deployment.
Duplicates, and retrieval collapse
A user who mentions being vegetarian in twelve separate sessions produces twelve near-identical vectors. All twelve score around 0.95 against a food question, so your top-5 returns five copies of the same sentence and the budget-limit fact, the allergy and the schedule constraint never appear.
This is retrieval collapse: a duplicated fact monopolises the result set. Fix it on write.
1DUPLICATE_THRESHOLD = 0.9323def remember_deduped(user_id, text, kind, confidence=1.0):4 similar = recall(user_id, text, k=1, floor=DUPLICATE_THRESHOLD)5 if similar:6 # Same fact restated: refresh recency and confidence, do not add a row.7 bump(similar[0], confidence)8 return "merged"9 remember(user_id, text, kind, confidence)10 return "added"Pick the threshold empirically. Around 0.93–0.96 catches restatements; below about 0.90 you start merging genuinely different facts, which is a worse error because it silently deletes information.
Staleness, and time-aware decay
"User is training for the Berlin marathon" was true last March and is misleading now. Decay old memories by multiplying their similarity by an exponential factor:
With a half-life t1/2 of 90 days, λ=0.6931/90=0.0077. A memory 180 days old — two half-lives — is multiplied by e−0.0077×180=e−1.386=0.25. So a stale memory scoring 0.90 on similarity ends up at 0.90×0.25=0.225 and falls below a 0.35 floor. A fresh memory scoring 0.70 beats it easily.
The crucial refinement: not everything should decay. Apply decay by kind.
| Kind | Half-life | Reasoning |
|---|---|---|
| Identity (name, language, timezone) | Never decays | Stable for years |
| Hard constraint (allergy, accessibility need) | Never decays | Forgetting this causes harm |
| Preference (likes dark mode, prefers brevity) | ~365 days | Drifts slowly |
| Project or goal | ~90 days | Projects end |
| Transient state (currently travelling) | ~7 days | True for days, not months |
Decaying an allergy on the same schedule as a dark-mode preference is how a memory system becomes a safety problem.
Metadata, and the filters you will wish you had
Every memory should carry, at minimum: user_id, kind, created_at, last_seen_at, confidence, source_session_id and the embedding_model version. Retrofitting these onto a populated store is painful because the information is gone — you cannot reconstruct which session a memory came from after the fact.
| Field | What it unlocks |
|---|---|
user_id | Isolation. Without it you have a data breach, not a feature. |
kind | Per-kind decay rates and targeted queries ("all hard constraints") |
created_at / last_seen_at | Decay, and distinguishing "stated once" from "reconfirmed nine times" |
confidence | Ranking, and a floor below which a memory is never injected |
source_session_id | Deletion requests, and auditing where a wrong memory came from |
embedding_model | Knowing which rows to re-embed after a model change |
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Assistant references facts about someone else | Missing or post-applied user_id filter | Filter inside the query; add an integration test that asserts isolation |
| Same fact retrieved five times | No write-time deduplication | Merge near-duplicates above ~0.93 |
| Assistant insists on outdated information | No decay, no contradiction handling | Per-kind half-lives; supersede on conflict; tell the model the live turn wins |
| Irrelevant memories shoehorned into every reply | No score floor, or a prompt that implies memories must be used | Set a floor around 0.35–0.4 and say explicitly that memories may be ignored |
| Retrieval returns "okay, sounds good" | Storing raw messages rather than extracted facts | Extract; return an empty list when nothing is durable |
| Every query returns 0.99 similarity | Query and documents embedded with different models, or distance metric mismatched to the index | Pin one model; set the index metric to cosine explicitly |
| Store works in test, useless in production | Tested with 20 memories; at 2,000 the good ones are outranked | Evaluate on a realistic store size with a labelled query set |
Building this for real
Vector memory is a database, and it deserves the discipline you would give any other database.
Instrument retrieval from day one. Log, for every turn: the query, the top-k results with their scores, and which ones cleared the floor. Without that log, "the bot forgot my address" is unanswerable — you cannot tell whether the memory was never written, was written and not retrieved, was retrieved and filtered out, or was retrieved and ignored by the model. Those are four different bugs with four different fixes.
Make deletion work. Users will ask you to forget things, and in many jurisdictions they have a legal right to. That means every memory must be traceable to a user and a source, and "delete everything for user X" must be a single indexed operation rather than a scan.
Budget the injection. Five memories at 40 tokens is 200 tokens per turn — about 0.06 cents at 3 dollars per million. Ten memories at 400 tokens each is 4,000 tokens per turn, which is twenty times the cost and measurably worse answers, because the model now has to find the relevant fact inside a wall of marginal ones. Fewer, sharper memories beat more, longer ones almost every time.
Know what your framework already gives you. If you build on LangGraph, its long-term memory store (InMemoryStore for development, PostgresStore in production) implements the pieces above: memories live under a namespace such as (user_id, "memories"), store.put writes them, and store.search(namespace, query=...) does the semantic lookup when the store is created with an embedding index. The namespace does the job of the user_id filter. Extraction, deduplication, decay and deletion are still yours to design.
Decide what you refuse to store. A memory system that silently records everything a user says is a liability. Payment details, health information, credentials and third-party personal data should be excluded at the extraction step, by an explicit rule, not by hoping the model chooses well. Write the exclusion list before you write the extractor.