Capstone Project: Multimodal Assistant

Memory & Semantic Search


On Monday a user tells your assistant: "I can't eat peanuts — they make my throat swell." On Friday they ask for a snack suggestion and the assistant recommends peanut butter on toast.

You investigate. The Monday conversation is on disk. So you add a lookup: before answering, search the stored history for anything relevant.

Python
>>> [t for t in history if "allergy" in t.content.lower()][ConversationTurn(role='user', content='Allergy season in Boston is brutal.')]

The search found the one line that contains the word "allergy" and it is about pollen. The line that actually matters — the one about peanuts — never uses the word. Keyword matching cannot connect "can't eat peanuts" to "allergy", because the connection is in the meaning, not the characters.

This stage builds the machinery that makes an assistant remember properly: conversation history that survives a restart, a store for stable facts, and search that works on meaning rather than spelling.

Why Friday forgot Monday's peanut allergyKeyword search finds nothing• Query says "snack",the note says "peanuts"• No shared token, no match• Exact words only, no paraphrase• Fails silently — returns an empty listSemantic search finds it• Both become vectors in one space• Cosine similarity, not string overlap• "Allergy" retrieves "throat swell"• Threshold decides what counts as related
The context window is a buffer, not memory: anything that must survive Monday to Friday has to be embedded and stored.

Why the context window is not memory

The tempting shortcut is to send the entire conversation with every request and call it memory. In a narrow sense it works — the model can refer to anything currently in front of it. It breaks in three specific ways.

It does not survive a restart. Close the process and the model knows nothing. Unless something wrote those turns to disk and reloads them, Monday is gone.

It does not scale, and the cost is easy to quantify. Take an average turn of 120 tokens. Fifty turns of history is 6,000 tokens prepended to every single request. At 100 requests a day that is 600,000 tokens per day spent re-sending things the user said weeks ago. Retrieve instead — pull the three most relevant past turns — and you send 3×120=3603 \times 120 = 360 tokens. That is a 16.7× reduction in context, with better focus, because three relevant turns beat fifty mostly-irrelevant ones for steering the answer.

It cannot be searched by meaning. The peanut example above. A conversation of 2,000 turns cannot be pasted into a prompt, and scanning it for literal words finds the wrong lines.

Memory for an assistant is not one mechanism. Recency, persistence, and meaning-based retrieval are three different problems with three different data structures, and trying to solve retrieval by simply keeping more history does not work.

Three kinds of memory

KindQuestion it answersStructureLifetimeRetrieval
Conversational"What did we just say?"Ordered list of recent turnsThis session, cappedTake the last k
Factual"What is the user's name / diet / timezone?"Key-value storeIndefinite, overwritten on changeExact key lookup
Semantic"What have we discussed related to X?"Vector index over textIndefinite, append-onlyNearest neighbours by embedding

Keep them as separate classes. A bug in fact storage should never be confusable with a bug in history, and you should be able to replace pickle-on-disk with a real database without touching the retrieval code.

Conversation memory

Python
# src/memory_manager.pyimport jsonfrom dataclasses import dataclass, asdict, fieldfrom datetime import datetime, timedeltafrom pathlib import Pathfrom config.settings import settingsfrom src.exceptions import MemoryError_from src.utils import logger@dataclassclass Turn:    role: str                      # "user" or "assistant"    content: str    timestamp: str = field(default_factory=lambda: datetime.now().isoformat())    image_path: str | None = Noneclass ConversationMemory:    """Recent turns, capped and persisted."""    def __init__(self, max_turns: int = 50, retention_days: int = 30):        self.max_turns = max_turns        self.retention_days = retention_days        self.turns: list[Turn] = []    def add(self, role: str, content: str, image_path: str | None = None) -> None:        self.turns.append(Turn(role, content, image_path=image_path))        if len(self.turns) > self.max_turns:            dropped = len(self.turns) - self.max_turns            self.turns = self.turns[-self.max_turns:]            logger.debug("Dropped %d turn(s) past the %d cap",                         dropped, self.max_turns)    def recent(self, k: int = 5) -> str:        """The last k turns, formatted for prepending to a prompt."""        return "\n".join(            f"{t.role.upper()}: {t.content}" for t in self.turns[-k:]        )    def prune(self) -> int:        cutoff = datetime.now() - timedelta(days=self.retention_days)        before = len(self.turns)        self.turns = [            t for t in self.turns            if datetime.fromisoformat(t.timestamp) >= cutoff        ]        return before - len(self.turns)    def save(self, path: Path) -> None:        try:            path.parent.mkdir(parents=True, exist_ok=True)            path.write_text(                json.dumps([asdict(t) for t in self.turns], indent=2),                encoding="utf-8",            )        except OSError as e:            raise MemoryError_(f"could not save session to {path}: {e}") from e    def load(self, path: Path) -> None:        if not path.exists():            logger.info("No session at %s; starting fresh", path)            return        try:            raw = json.loads(path.read_text(encoding="utf-8"))        except (OSError, json.JSONDecodeError) as e:            raise MemoryError_(f"corrupt session file {path}: {e}") from e        self.turns = [Turn(**item) for item in raw]        logger.info("Loaded %d turns from %s", len(self.turns), path)

Two decisions deserve explanation. The max_turns cap bounds both memory usage and prompt size: without it an assistant left running for a week accumulates turns until recent() starts returning more text than the model's context window can hold, and every request gets slower and more expensive as it grows. Dropping the oldest turns is the crude fix; summarising them into a compact paragraph is the better one, and it is the natural extension once the basics work.

The second is JSON rather than pickle. Pickle is faster to write and can serialise arbitrary Python objects, but unpickling a file executes code inside it, which makes a session file a remote code execution vector the moment it comes from anywhere you do not fully control. JSON is inspectable, diffable, and inert.

Stable facts

Python
class FactMemory:    """Small, stable, exact-lookup facts. Not a search index."""    def __init__(self):        self.facts: dict[str, str] = {}    def remember(self, key: str, value: str) -> None:        key = key.strip().lower()        if key in self.facts and self.facts[key] != value:            logger.info("Updating fact %r: %r -> %r",                        key, self.facts[key], value)        self.facts[key] = value    def recall(self, key: str) -> str | None:        return self.facts.get(key.strip().lower())    def as_context(self) -> str:        if not self.facts:            return ""        return "Known facts about the user:\n" + "\n".join(            f"- {k}: {v}" for k, v in sorted(self.facts.items())        )

Facts are deliberately overwritten rather than appended. If the user moves from Boston to Berlin, you want one current timezone, not two contradictory ones that the model has to arbitrate between — and it will arbitrate badly, usually by picking whichever appeared last in the prompt.

Embeddings: what "similar" actually means

An embedding model maps a piece of text to a fixed-length vector of numbers, trained so that texts with similar meaning end up pointing in similar directions. all-MiniLM-L6-v2 produces 384 numbers per sentence and runs comfortably on a CPU.

Similarity is the cosine of the angle between two vectors:

cos⁡(a,b)=a⋅b∥a∥ ∥b∥\cos(a,b) = \frac{a \cdot b}{\|a\|\,\|b\|}

For unit-length vectors — which most embedding libraries produce, or can be asked to produce — this connects directly to Euclidean distance, and that connection matters because vector databases usually report distance, not similarity:

∥a−b∥2=∥a∥2+∥b∥2−2 a ⁣⋅ ⁣b=2−2cos⁡(a,b)⟹cos⁡(a,b)=1−∥a−b∥22\|a-b\|^2 = \|a\|^2 + \|b\|^2 - 2\,a\!\cdot\!b = 2 - 2\cos(a,b) \quad\Longrightarrow\quad \cos(a,b) = 1 - \frac{\|a-b\|^2}{2}

So a squared Euclidean distance of 0.78 means a cosine similarity of 1−0.39=0.611 - 0.39 = 0.61. Chroma reports squared L2 by default; configure the collection with "hnsw:space": "cosine" and it reports cosine distance instead, which is simply 1−cos⁡1 - \cos — so the same pair comes back as 0.39 rather than 0.78. The conversion depends on which space you configured, and applying the wrong one silently shifts every threshold you set.

Configured spaceWhat distances holdsSimilarity from distanceUnrelated pair looks like
l2 (default)Squared Euclidean distance1−d/21 - d/2d≈2.0d \approx 2.0
cosineCosine distance, 1−cos⁡1-\cos1−d1 - dd≈1.0d \approx 1.0
ipNegative inner product−d-d for unit vectorsd≈0d \approx 0

Whichever you pick, distance ascends — smaller is closer — while similarity descends. Reversing that sort is one of the most common bugs in retrieval code, and it fails in the nastiest possible way: it returns results reliably, quickly, and always the least relevant documents you have.

Where keyword and semantic search disagree

Return to the opening failure with real numbers. The user asks "which foods am I allergic to?" and the store holds three sentences. The similarities below were measured with all-MiniLM-L6-v2; other models give different values, but the ranking is the point.

Stored textContains "allergy"?Keyword rankCosine similaritySemantic rank
"I can't eat peanuts — they make my throat swell."Nonot returned0.511
"Allergy season in Boston is brutal in April."Yes10.362
"My favourite coffee is a flat white."Nonot returned0.173

Keyword search for "allergy" returns exactly one result and it is the wrong one. Semantic search ranks the peanut sentence first despite sharing not a single content word with the query, because "can't eat X, throat swells" and "allergic to foods" occupy nearby regions of the embedding space. Notice the margins, too: 0.51 against 0.36 is a clear win, not a landslide. Phrase the query as "what did I tell you about my allergy?" and the same model puts the pollen sentence first, because that query is about the word as much as the meaning. That is why you test retrieval against real queries, and why the threshold below matters.

The reverse case exists too, and it is why serious systems eventually run both. Search for an order number, a function name, or a rare proper noun and semantic search will happily return things that are about the same topic while missing the exact string. Embeddings are good at meaning and mediocre at identifiers.

Semantic search finds text that means the same thing; keyword search finds text that says the same thing. Assistants need the first far more often, and product codes are the case where you still need the second.

The vector store

Python
import chromadbfrom chromadb.utils import embedding_functionsclass SemanticSearch:    """Chroma-backed nearest-neighbour search over remembered text."""    def __init__(self, collection: str = "assistant_memory"):        self.client = chromadb.PersistentClient(            path=str(settings.data_dir / "chroma")        )        self.embedder = embedding_functions.SentenceTransformerEmbeddingFunction(            model_name=settings.embedding_model      # all-MiniLM-L6-v2, 384 dims        )        # The embedding function is bound to the collection at creation.        # Changing settings.embedding_model later without rebuilding the        # collection mixes incompatible vectors -- see the failure table.        self.collection = self.client.get_or_create_collection(            name=collection,            embedding_function=self.embedder,            metadata={"hnsw:space": "cosine"},        )        logger.info("Vector store ready: %d documents",                    self.collection.count())    def add(self, text: str, doc_id: str, metadata: dict | None = None) -> None:        if not text.strip():            return        try:            self.collection.upsert(                documents=[text],                ids=[doc_id],                # current Chroma rejects an empty metadata dict; send None                metadatas=[metadata] if metadata else None,            )        except Exception as e:            raise MemoryError_(f"indexing failed for {doc_id}: {e}") from e    def search(self, query: str, k: int = 3,               max_distance: float = 0.6) -> list[dict]:        n = min(k, max(self.collection.count(), 1))        res = self.collection.query(query_texts=[query], n_results=n)        if not res["ids"][0]:            return []        hits = []        for doc, dist, meta in zip(res["documents"][0],                                   res["distances"][0],                                   res["metadatas"][0]):            # Cosine space: distance = 1 - cosine similarity, ascending.            if dist > max_distance:      # 0.6 keeps similarity >= 0.4                continue            hits.append({"text": doc, "distance": dist,                         "similarity": 1 - dist, "metadata": meta})        logger.debug("Query %r returned %d hit(s) within %.2f",                     query, len(hits), max_distance)        return hits

Three details are load-bearing. upsert rather than add makes re-indexing idempotent — run your ingestion script twice with add and you get duplicate-ID errors or duplicated results. Clamping n_results to the collection size keeps behaviour the same across Chroma versions — older releases warned or raised when you asked for ten neighbours from a store holding four. And the distance filter matters because nearest-neighbour search always returns neighbours: ask an unrelated question of a store containing only recipes and you get the three closest recipes, at cosine distance 0.88, which the model will then dutifully weave into an answer. A threshold turns "here is the least-bad match" into "nothing relevant found".

Sizing it

Storage is straightforward: 384 dimensions at 4 bytes each is 1,536 bytes per vector. Ten thousand remembered turns is 10,000×384×4=15,360,00010{,}000 \times 384 \times 4 = 15{,}360{,}000 bytes, about 15 MB, plus the original text. Switch to a 1,536-dimension model and the same corpus needs 10,000×1536×4=61,440,00010{,}000 \times 1536 \times 4 = 61{,}440{,}000 bytes — four times the storage and four times the arithmetic per comparison.

Search cost is why the index exists. A brute-force scan of 10,000 vectors at 384 dimensions is 3.84 million multiply-adds, which takes single-digit milliseconds and is entirely fine. At 10 million documents it is 3.84 billion per query, which is not. Chroma's HNSW index trades exactness for approximate results in logarithmic time, and that trade is invisible at small scale and essential at large.

Putting it together

Python
class MemoryManager:    """One object the rest of the assistant talks to."""    def __init__(self, session_id: str = "default"):        self.session_path = settings.data_dir / f"session_{session_id}.json"        self.conversation = ConversationMemory()        self.facts = FactMemory()        self.semantic = SemanticSearch()        self.conversation.load(self.session_path)    def record(self, role: str, content: str,               image_path: str | None = None) -> None:        self.conversation.add(role, content, image_path)        # Index user statements only. Indexing the assistant's own replies        # means later searches retrieve things the assistant made up.        if role == "user" and len(content.split()) >= 4:            doc_id = f"{self.session_path.stem}-{len(self.conversation.turns)}"            self.semantic.add(content, doc_id, {"role": role})    def context_for(self, query: str, k_recent: int = 5,                    k_semantic: int = 3) -> str:        parts = []        if facts := self.facts.as_context():            parts.append(facts)        hits = self.semantic.search(query, k=k_semantic)        if hits:            parts.append("Relevant things the user said previously:\n" + "\n".join(                f"- {h['text']} (similarity {h['similarity']:.2f})" for h in hits            ))        if recent := self.conversation.recent(k_recent):            parts.append("Recent conversation:\n" + recent)        return "\n\n".join(parts)    def checkpoint(self) -> None:        self.conversation.save(self.session_path)

The rule in record() is worth dwelling on. Index the user's statements, not the assistant's replies. If you index everything, a hallucinated answer from Tuesday gets retrieved on Thursday as "relevant context", the model treats it as established fact, and the error compounds. The minimum word count is the same instinct: indexing "ok" and "thanks" fills the store with vectors that match everything weakly and nothing usefully.

Tests

Python
# tests/test_memory.pyimport pytestfrom src.memory_manager import ConversationMemory, FactMemorydef test_cap_keeps_the_newest_turns():    m = ConversationMemory(max_turns=3)    for i in range(5):        m.add("user", f"message {i}")    assert len(m.turns) == 3    assert m.turns[0].content == "message 2"    assert m.turns[-1].content == "message 4"def test_round_trip_survives_a_restart(tmp_path):    a = ConversationMemory()    a.add("user", "I can't eat peanuts")    a.save(tmp_path / "s.json")    b = ConversationMemory()    b.load(tmp_path / "s.json")    assert b.turns[0].content == "I can't eat peanuts"def test_corrupt_session_raises_named_error(tmp_path):    from src.exceptions import MemoryError_    p = tmp_path / "s.json"    p.write_text("{not json", encoding="utf-8")    with pytest.raises(MemoryError_):        ConversationMemory().load(p)def test_facts_are_updated_not_duplicated():    f = FactMemory()    f.remember("City", "Boston")    f.remember("city", "Berlin")    assert f.recall("CITY") == "Berlin"    assert len(f.facts) == 1@pytest.mark.slowdef test_semantic_beats_keyword(tmp_path, monkeypatch):    from src.memory_manager import SemanticSearch    monkeypatch.setattr("config.settings.settings.data_dir", tmp_path)    s = SemanticSearch(collection="test_sem")    s.add("I can't eat peanuts, they make my throat swell", "d1")    s.add("Allergy season in Boston is brutal in April", "d2")    s.add("My favourite coffee is a flat white", "d3")    top = s.search("which foods am I allergic to?", k=3)[0]    assert "peanuts" in top["text"]      # the doc WITHOUT the word "allergy"

Mark the embedding test slow and give it its own collection in tmp_path. Two habits it enforces: never let a test write into your real vector store — a test that adds three documents to your production collection pollutes every subsequent search — and never assert on an exact similarity value, because it shifts between model versions. Assert on ordering, which is the property you actually depend on.

When things go wrong here

SymptomCauseFix
Results are consistently the least relevant documentsSorted distances descending, or treated distance as similarityDistance ascends, similarity descends; convert once, using the formula for the space you configured, and never mix units
Irrelevant results returned for every queryNo distance threshold — k-NN always returns k thingsFilter on max_distance; return an empty list and let the caller say "nothing found"
A dimension-mismatch error (InvalidDimensionException) on add or queryCollection built with a 384-dim model, queried after switching to 1,536-dimOne embedding model per collection; changing it means deleting and re-indexing
Silently poor results after a model change with the same dimensionOld and new vectors coexist in one incompatible spaceVersion the collection name, e.g. memory_minilm_v1, so a change forces a rebuild
Empty results from a non-empty storen_results exceeds the document count, or the threshold is too tightClamp n to collection.count(); log distances at DEBUG before filtering
The assistant "remembers" things it inventedAssistant replies were indexed alongside user statementsIndex user turns only; tag metadata with role and filter on it
Duplicate hits after re-running ingestionadd used where upsert was neededDeterministic IDs plus upsert
Memory forgotten after every restartcheckpoint() never called, or called only on the clean-exit pathSave in a finally block so a crash or interrupt still persists
Prompts creeping past the context limitRetrieved context appended without any budgetCap k_semantic, cap k_recent, and log the token count of assembled context
Search returns nothing for an exact order numberEmbeddings are weak on identifiersKeep a keyword index alongside and merge results for identifier-shaped queries

Acceptance criteria for this stage

  1. Recording 60 turns into a memory capped at 50 leaves exactly 50, with the oldest ten gone and the newest intact.
  2. A session saved, the process killed, and a new process started reloads the same turn count and the same first message.
  3. Deliberately corrupting the session JSON raises MemoryError_ with the file path in the message — not a bare JSONDecodeError.
  4. Querying "which foods am I allergic to?" against the three-sentence store ranks the peanut sentence first, above the sentence that literally contains "allergy".
  5. Querying something unrelated to everything stored returns an empty list, not three distant matches.
  6. Asking for 10 neighbours from a store holding 3 documents returns 3 results and raises nothing.
  7. Re-running the ingestion script twice leaves collection.count() unchanged.
  8. context_for() on a store of 500 turns produces context under 800 tokens — verify by measuring, not by assuming.

What this buys you when the assistant is running

The visible payoff is the Friday snack question answering correctly. The structural payoff is that context_for() returns a plain string, which drops straight into the context slot your prompt builder already accepts. Nothing upstream needs to know whether that string came from the last five turns, a stored fact, or a vector search across six months of conversation. Retrieval strategy becomes something you can change — reweight the mix, add reranking, summarise old turns instead of dropping them — without touching a single prompt or a single line of reasoning code.

The discipline that matters most is the distance threshold. Everything else in this stage fails loudly: a corrupt session raises, a dimension mismatch raises, a missing file logs. Retrieval failure is silent. It returns three plausible-looking documents at cosine distance 0.9, the model incorporates them, and the user gets a confidently wrong answer built on genuinely irrelevant context. Log the distances of everything you retrieve, set a threshold you have actually calibrated against your own data, and treat "no relevant memory" as a normal, expected outcome rather than something to paper over.