Course Content
Capstone Project: Multimodal Assistant
1 sections · 6 lessons
Memory & Semantic Search
On Monday a user tells your assistant: "I can't eat peanuts — they make my throat swell." On Friday they ask for a snack suggestion and the assistant recommends peanut butter on toast.
You investigate. The Monday conversation is on disk. So you add a lookup: before answering, search the stored history for anything relevant.
>>> [t for t in history if "allergy" in t.content.lower()][ConversationTurn(role='user', content='Allergy season in Boston is brutal.')]The search found the one line that contains the word "allergy" and it is about pollen. The line that actually matters — the one about peanuts — never uses the word. Keyword matching cannot connect "can't eat peanuts" to "allergy", because the connection is in the meaning, not the characters.
This stage builds the machinery that makes an assistant remember properly: conversation history that survives a restart, a store for stable facts, and search that works on meaning rather than spelling.
Why the context window is not memory
The tempting shortcut is to send the entire conversation with every request and call it memory. In a narrow sense it works — the model can refer to anything currently in front of it. It breaks in three specific ways.
It does not survive a restart. Close the process and the model knows nothing. Unless something wrote those turns to disk and reloads them, Monday is gone.
It does not scale, and the cost is easy to quantify. Take an average turn of 120 tokens. Fifty turns of history is 6,000 tokens prepended to every single request. At 100 requests a day that is 600,000 tokens per day spent re-sending things the user said weeks ago. Retrieve instead — pull the three most relevant past turns — and you send 3×120=360 tokens. That is a 16.7× reduction in context, with better focus, because three relevant turns beat fifty mostly-irrelevant ones for steering the answer.
It cannot be searched by meaning. The peanut example above. A conversation of 2,000 turns cannot be pasted into a prompt, and scanning it for literal words finds the wrong lines.
Memory for an assistant is not one mechanism. Recency, persistence, and meaning-based retrieval are three different problems with three different data structures, and trying to solve retrieval by simply keeping more history does not work.
Three kinds of memory
| Kind | Question it answers | Structure | Lifetime | Retrieval |
|---|---|---|---|---|
| Conversational | "What did we just say?" | Ordered list of recent turns | This session, capped | Take the last k |
| Factual | "What is the user's name / diet / timezone?" | Key-value store | Indefinite, overwritten on change | Exact key lookup |
| Semantic | "What have we discussed related to X?" | Vector index over text | Indefinite, append-only | Nearest neighbours by embedding |
Keep them as separate classes. A bug in fact storage should never be confusable with a bug in history, and you should be able to replace pickle-on-disk with a real database without touching the retrieval code.
Conversation memory
1# src/memory_manager.py2import json3from dataclasses import dataclass, asdict, field4from datetime import datetime, timedelta5from pathlib import Path67from config.settings import settings8from src.exceptions import MemoryError_9from src.utils import logger101112@dataclass13class Turn:14 role: str # "user" or "assistant"15 content: str16 timestamp: str = field(default_factory=lambda: datetime.now().isoformat())17 image_path: str | None = None181920class ConversationMemory:21 """Recent turns, capped and persisted."""2223 def __init__(self, max_turns: int = 50, retention_days: int = 30):24 self.max_turns = max_turns25 self.retention_days = retention_days26 self.turns: list[Turn] = []2728 def add(self, role: str, content: str, image_path: str | None = None) -> None:29 self.turns.append(Turn(role, content, image_path=image_path))30 if len(self.turns) > self.max_turns:31 dropped = len(self.turns) - self.max_turns32 self.turns = self.turns[-self.max_turns:]33 logger.debug("Dropped %d turn(s) past the %d cap",34 dropped, self.max_turns)3536 def recent(self, k: int = 5) -> str:37 """The last k turns, formatted for prepending to a prompt."""38 return "\n".join(39 f"{t.role.upper()}: {t.content}" for t in self.turns[-k:]40 )4142 def prune(self) -> int:43 cutoff = datetime.now() - timedelta(days=self.retention_days)44 before = len(self.turns)45 self.turns = [46 t for t in self.turns47 if datetime.fromisoformat(t.timestamp) >= cutoff48 ]49 return before - len(self.turns)5051 def save(self, path: Path) -> None:52 try:53 path.parent.mkdir(parents=True, exist_ok=True)54 path.write_text(55 json.dumps([asdict(t) for t in self.turns], indent=2),56 encoding="utf-8",57 )58 except OSError as e:59 raise MemoryError_(f"could not save session to {path}: {e}") from e6061 def load(self, path: Path) -> None:62 if not path.exists():63 logger.info("No session at %s; starting fresh", path)64 return65 try:66 raw = json.loads(path.read_text(encoding="utf-8"))67 except (OSError, json.JSONDecodeError) as e:68 raise MemoryError_(f"corrupt session file {path}: {e}") from e69 self.turns = [Turn(**item) for item in raw]70 logger.info("Loaded %d turns from %s", len(self.turns), path)Two decisions deserve explanation. The max_turns cap bounds both memory usage and prompt size: without it an assistant left running for a week accumulates turns until recent() starts returning more text than the model's context window can hold, and every request gets slower and more expensive as it grows. Dropping the oldest turns is the crude fix; summarising them into a compact paragraph is the better one, and it is the natural extension once the basics work.
The second is JSON rather than pickle. Pickle is faster to write and can serialise arbitrary Python objects, but unpickling a file executes code inside it, which makes a session file a remote code execution vector the moment it comes from anywhere you do not fully control. JSON is inspectable, diffable, and inert.
Stable facts
1class FactMemory:2 """Small, stable, exact-lookup facts. Not a search index."""34 def __init__(self):5 self.facts: dict[str, str] = {}67 def remember(self, key: str, value: str) -> None:8 key = key.strip().lower()9 if key in self.facts and self.facts[key] != value:10 logger.info("Updating fact %r: %r -> %r",11 key, self.facts[key], value)12 self.facts[key] = value1314 def recall(self, key: str) -> str | None:15 return self.facts.get(key.strip().lower())1617 def as_context(self) -> str:18 if not self.facts:19 return ""20 return "Known facts about the user:\n" + "\n".join(21 f"- {k}: {v}" for k, v in sorted(self.facts.items())22 )Facts are deliberately overwritten rather than appended. If the user moves from Boston to Berlin, you want one current timezone, not two contradictory ones that the model has to arbitrate between — and it will arbitrate badly, usually by picking whichever appeared last in the prompt.
Embeddings: what "similar" actually means
An embedding model maps a piece of text to a fixed-length vector of numbers, trained so that texts with similar meaning end up pointing in similar directions. all-MiniLM-L6-v2 produces 384 numbers per sentence and runs comfortably on a CPU.
Similarity is the cosine of the angle between two vectors:
For unit-length vectors — which most embedding libraries produce, or can be asked to produce — this connects directly to Euclidean distance, and that connection matters because vector databases usually report distance, not similarity:
So a squared Euclidean distance of 0.78 means a cosine similarity of 1−0.39=0.61. Chroma reports squared L2 by default; configure the collection with "hnsw:space": "cosine" and it reports cosine distance instead, which is simply 1−cos — so the same pair comes back as 0.39 rather than 0.78. The conversion depends on which space you configured, and applying the wrong one silently shifts every threshold you set.
| Configured space | What distances holds | Similarity from distance | Unrelated pair looks like |
|---|---|---|---|
l2 (default) | Squared Euclidean distance | 1−d/2 | d≈2.0 |
cosine | Cosine distance, 1−cos | 1−d | d≈1.0 |
ip | Negative inner product | −d for unit vectors | d≈0 |
Whichever you pick, distance ascends — smaller is closer — while similarity descends. Reversing that sort is one of the most common bugs in retrieval code, and it fails in the nastiest possible way: it returns results reliably, quickly, and always the least relevant documents you have.
Where keyword and semantic search disagree
Return to the opening failure with real numbers. The user asks "which foods am I allergic to?" and the store holds three sentences. The similarities below were measured with all-MiniLM-L6-v2; other models give different values, but the ranking is the point.
| Stored text | Contains "allergy"? | Keyword rank | Cosine similarity | Semantic rank |
|---|---|---|---|---|
| "I can't eat peanuts — they make my throat swell." | No | not returned | 0.51 | 1 |
| "Allergy season in Boston is brutal in April." | Yes | 1 | 0.36 | 2 |
| "My favourite coffee is a flat white." | No | not returned | 0.17 | 3 |
Keyword search for "allergy" returns exactly one result and it is the wrong one. Semantic search ranks the peanut sentence first despite sharing not a single content word with the query, because "can't eat X, throat swells" and "allergic to foods" occupy nearby regions of the embedding space. Notice the margins, too: 0.51 against 0.36 is a clear win, not a landslide. Phrase the query as "what did I tell you about my allergy?" and the same model puts the pollen sentence first, because that query is about the word as much as the meaning. That is why you test retrieval against real queries, and why the threshold below matters.
The reverse case exists too, and it is why serious systems eventually run both. Search for an order number, a function name, or a rare proper noun and semantic search will happily return things that are about the same topic while missing the exact string. Embeddings are good at meaning and mediocre at identifiers.
Semantic search finds text that means the same thing; keyword search finds text that says the same thing. Assistants need the first far more often, and product codes are the case where you still need the second.
The vector store
1import chromadb2from chromadb.utils import embedding_functions345class SemanticSearch:6 """Chroma-backed nearest-neighbour search over remembered text."""78 def __init__(self, collection: str = "assistant_memory"):9 self.client = chromadb.PersistentClient(10 path=str(settings.data_dir / "chroma")11 )12 self.embedder = embedding_functions.SentenceTransformerEmbeddingFunction(13 model_name=settings.embedding_model # all-MiniLM-L6-v2, 384 dims14 )15 # The embedding function is bound to the collection at creation.16 # Changing settings.embedding_model later without rebuilding the17 # collection mixes incompatible vectors -- see the failure table.18 self.collection = self.client.get_or_create_collection(19 name=collection,20 embedding_function=self.embedder,21 metadata={"hnsw:space": "cosine"},22 )23 logger.info("Vector store ready: %d documents",24 self.collection.count())2526 def add(self, text: str, doc_id: str, metadata: dict | None = None) -> None:27 if not text.strip():28 return29 try:30 self.collection.upsert(31 documents=[text],32 ids=[doc_id],33 # current Chroma rejects an empty metadata dict; send None34 metadatas=[metadata] if metadata else None,35 )36 except Exception as e:37 raise MemoryError_(f"indexing failed for {doc_id}: {e}") from e3839 def search(self, query: str, k: int = 3,40 max_distance: float = 0.6) -> list[dict]:41 n = min(k, max(self.collection.count(), 1))42 res = self.collection.query(query_texts=[query], n_results=n)43 if not res["ids"][0]:44 return []4546 hits = []47 for doc, dist, meta in zip(res["documents"][0],48 res["distances"][0],49 res["metadatas"][0]):50 # Cosine space: distance = 1 - cosine similarity, ascending.51 if dist > max_distance: # 0.6 keeps similarity >= 0.452 continue53 hits.append({"text": doc, "distance": dist,54 "similarity": 1 - dist, "metadata": meta})55 logger.debug("Query %r returned %d hit(s) within %.2f",56 query, len(hits), max_distance)57 return hitsThree details are load-bearing. upsert rather than add makes re-indexing idempotent — run your ingestion script twice with add and you get duplicate-ID errors or duplicated results. Clamping n_results to the collection size keeps behaviour the same across Chroma versions — older releases warned or raised when you asked for ten neighbours from a store holding four. And the distance filter matters because nearest-neighbour search always returns neighbours: ask an unrelated question of a store containing only recipes and you get the three closest recipes, at cosine distance 0.88, which the model will then dutifully weave into an answer. A threshold turns "here is the least-bad match" into "nothing relevant found".
Sizing it
Storage is straightforward: 384 dimensions at 4 bytes each is 1,536 bytes per vector. Ten thousand remembered turns is 10,000×384×4=15,360,000 bytes, about 15 MB, plus the original text. Switch to a 1,536-dimension model and the same corpus needs 10,000×1536×4=61,440,000 bytes — four times the storage and four times the arithmetic per comparison.
Search cost is why the index exists. A brute-force scan of 10,000 vectors at 384 dimensions is 3.84 million multiply-adds, which takes single-digit milliseconds and is entirely fine. At 10 million documents it is 3.84 billion per query, which is not. Chroma's HNSW index trades exactness for approximate results in logarithmic time, and that trade is invisible at small scale and essential at large.
Putting it together
1class MemoryManager:2 """One object the rest of the assistant talks to."""34 def __init__(self, session_id: str = "default"):5 self.session_path = settings.data_dir / f"session_{session_id}.json"6 self.conversation = ConversationMemory()7 self.facts = FactMemory()8 self.semantic = SemanticSearch()9 self.conversation.load(self.session_path)1011 def record(self, role: str, content: str,12 image_path: str | None = None) -> None:13 self.conversation.add(role, content, image_path)14 # Index user statements only. Indexing the assistant's own replies15 # means later searches retrieve things the assistant made up.16 if role == "user" and len(content.split()) >= 4:17 doc_id = f"{self.session_path.stem}-{len(self.conversation.turns)}"18 self.semantic.add(content, doc_id, {"role": role})1920 def context_for(self, query: str, k_recent: int = 5,21 k_semantic: int = 3) -> str:22 parts = []23 if facts := self.facts.as_context():24 parts.append(facts)2526 hits = self.semantic.search(query, k=k_semantic)27 if hits:28 parts.append("Relevant things the user said previously:\n" + "\n".join(29 f"- {h['text']} (similarity {h['similarity']:.2f})" for h in hits30 ))3132 if recent := self.conversation.recent(k_recent):33 parts.append("Recent conversation:\n" + recent)3435 return "\n\n".join(parts)3637 def checkpoint(self) -> None:38 self.conversation.save(self.session_path)The rule in record() is worth dwelling on. Index the user's statements, not the assistant's replies. If you index everything, a hallucinated answer from Tuesday gets retrieved on Thursday as "relevant context", the model treats it as established fact, and the error compounds. The minimum word count is the same instinct: indexing "ok" and "thanks" fills the store with vectors that match everything weakly and nothing usefully.
Tests
1# tests/test_memory.py2import pytest3from src.memory_manager import ConversationMemory, FactMemory456def test_cap_keeps_the_newest_turns():7 m = ConversationMemory(max_turns=3)8 for i in range(5):9 m.add("user", f"message {i}")10 assert len(m.turns) == 311 assert m.turns[0].content == "message 2"12 assert m.turns[-1].content == "message 4"131415def test_round_trip_survives_a_restart(tmp_path):16 a = ConversationMemory()17 a.add("user", "I can't eat peanuts")18 a.save(tmp_path / "s.json")1920 b = ConversationMemory()21 b.load(tmp_path / "s.json")22 assert b.turns[0].content == "I can't eat peanuts"232425def test_corrupt_session_raises_named_error(tmp_path):26 from src.exceptions import MemoryError_27 p = tmp_path / "s.json"28 p.write_text("{not json", encoding="utf-8")29 with pytest.raises(MemoryError_):30 ConversationMemory().load(p)313233def test_facts_are_updated_not_duplicated():34 f = FactMemory()35 f.remember("City", "Boston")36 f.remember("city", "Berlin")37 assert f.recall("CITY") == "Berlin"38 assert len(f.facts) == 1394041@pytest.mark.slow42def test_semantic_beats_keyword(tmp_path, monkeypatch):43 from src.memory_manager import SemanticSearch44 monkeypatch.setattr("config.settings.settings.data_dir", tmp_path)45 s = SemanticSearch(collection="test_sem")46 s.add("I can't eat peanuts, they make my throat swell", "d1")47 s.add("Allergy season in Boston is brutal in April", "d2")48 s.add("My favourite coffee is a flat white", "d3")4950 top = s.search("which foods am I allergic to?", k=3)[0]51 assert "peanuts" in top["text"] # the doc WITHOUT the word "allergy"Mark the embedding test slow and give it its own collection in tmp_path. Two habits it enforces: never let a test write into your real vector store — a test that adds three documents to your production collection pollutes every subsequent search — and never assert on an exact similarity value, because it shifts between model versions. Assert on ordering, which is the property you actually depend on.
When things go wrong here
| Symptom | Cause | Fix |
|---|---|---|
| Results are consistently the least relevant documents | Sorted distances descending, or treated distance as similarity | Distance ascends, similarity descends; convert once, using the formula for the space you configured, and never mix units |
| Irrelevant results returned for every query | No distance threshold — k-NN always returns k things | Filter on max_distance; return an empty list and let the caller say "nothing found" |
A dimension-mismatch error (InvalidDimensionException) on add or query | Collection built with a 384-dim model, queried after switching to 1,536-dim | One embedding model per collection; changing it means deleting and re-indexing |
| Silently poor results after a model change with the same dimension | Old and new vectors coexist in one incompatible space | Version the collection name, e.g. memory_minilm_v1, so a change forces a rebuild |
| Empty results from a non-empty store | n_results exceeds the document count, or the threshold is too tight | Clamp n to collection.count(); log distances at DEBUG before filtering |
| The assistant "remembers" things it invented | Assistant replies were indexed alongside user statements | Index user turns only; tag metadata with role and filter on it |
| Duplicate hits after re-running ingestion | add used where upsert was needed | Deterministic IDs plus upsert |
| Memory forgotten after every restart | checkpoint() never called, or called only on the clean-exit path | Save in a finally block so a crash or interrupt still persists |
| Prompts creeping past the context limit | Retrieved context appended without any budget | Cap k_semantic, cap k_recent, and log the token count of assembled context |
| Search returns nothing for an exact order number | Embeddings are weak on identifiers | Keep a keyword index alongside and merge results for identifier-shaped queries |
Acceptance criteria for this stage
- Recording 60 turns into a memory capped at 50 leaves exactly 50, with the oldest ten gone and the newest intact.
- A session saved, the process killed, and a new process started reloads the same turn count and the same first message.
- Deliberately corrupting the session JSON raises
MemoryError_with the file path in the message — not a bareJSONDecodeError. - Querying "which foods am I allergic to?" against the three-sentence store ranks the peanut sentence first, above the sentence that literally contains "allergy".
- Querying something unrelated to everything stored returns an empty list, not three distant matches.
- Asking for 10 neighbours from a store holding 3 documents returns 3 results and raises nothing.
- Re-running the ingestion script twice leaves
collection.count()unchanged. context_for()on a store of 500 turns produces context under 800 tokens — verify by measuring, not by assuming.
What this buys you when the assistant is running
The visible payoff is the Friday snack question answering correctly. The structural payoff is that context_for() returns a plain string, which drops straight into the context slot your prompt builder already accepts. Nothing upstream needs to know whether that string came from the last five turns, a stored fact, or a vector search across six months of conversation. Retrieval strategy becomes something you can change — reweight the mix, add reranking, summarise old turns instead of dropping them — without touching a single prompt or a single line of reasoning code.
The discipline that matters most is the distance threshold. Everything else in this stage fails loudly: a corrupt session raises, a dimension mismatch raises, a missing file logs. Retrieval failure is silent. It returns three plausible-looking documents at cosine distance 0.9, the model incorporates them, and the user gets a confidently wrong answer built on genuinely irrelevant context. Log the distances of everything you retrieve, set a threshold you have actually calibrated against your own data, and treat "no relevant memory" as a normal, expected outcome rather than something to paper over.