Course Content
Context Management and Memory
3 sections · 6 lessons
Mini-Project: A Chatbot That Remembers
Here is the only thing this project has to do. You run it, tell it something, kill the process, run it again, and it still knows.
1$ python -m mnemo --user alice2you> I'm vegetarian, and I'm allergic to peanuts.3bot> Noted. I'll keep both in mind.4you> ^C56$ python -m mnemo --user alice # new process, new session, hours later7you> suggest a quick lunch8bot> A chickpea and roasted-red-pepper wrap with tahini takes about ten9 minutes. Tahini is sesame, not peanut, so you're fine there.Two details in that output are the entire assignment. The bot applied a constraint from a session that no longer exists in any process memory. And it went out of its way to confirm that tahini is not peanut — which means the allergy was not a vague impression, it was a retrievable, specific fact.
Now the harder test, the one that separates a working memory system from a store of stale assertions:
1$ python -m mnemo --user alice2you> actually I've started eating fish again3bot> Got it, I'll update that.4you> ^C56$ python -m mnemo --user alice7you> suggest a quick lunch8bot> Pan-fried salmon with lemon and a green salad, about fifteen minutes.9 Still no peanuts.The vegetarian fact was superseded, not appended. The allergy was not touched. Most first attempts get this wrong in one of two ways: they keep both dietary facts and the bot contradicts itself, or they overwrite the whole dietary record and the peanut allergy quietly disappears. The second one is the dangerous bug, and it is the one a naive implementation ships.
What you are building
A command-line assistant called mnemo with three storage layers, each doing a job the others cannot.
| Layer | Technology | Holds | Why not one of the others |
|---|---|---|---|
| Working state | Redis | Active session, recent turns, cached retrievals | Needs TTLs and atomic counters; must be shared across processes |
| Semantic memory | Chroma | Embedded text of episodes and facts | Only a vector index answers "what do I know that relates to this?" |
| Structured facts | SQLite | Versioned (subject, predicate, object) triples | Only a relational key detects that "Munich" and "Berlin" conflict |
Three stores is not over-engineering. A vector index cannot detect contradiction, a relational table cannot answer a fuzzy question, and neither should be on the hot path for "what did the user just say".
mnemo/ __main__.py CLI loop, argument parsing config.py budgets, thresholds, model names -- all constants here chat.py prompt assembly + model call budget.py token counting and trimming memory/ vectors.py Chroma: embed, write, recall facts.py SQLite: assert_fact, current_facts, supersede extract.py turn -> candidate facts (small model) profile.py current_facts -> rendered profile block session.py Redis: start, append turn, end, consolidate cache.py Redis: embedding + retrieval caches, version keystests/ test_persistence.py test_isolation.py test_conflicts.py test_budget.py test_degradation.pyEverything tunable lives in config.py. Thresholds scattered through the codebase are how a retrieval system becomes impossible to evaluate — you cannot sweep a parameter that is written in four files.
The token budget, decided first
Write this table before writing code. It is the contract every other component has to respect.
| Component | Budget | Source |
|---|---|---|
| System prompt | 350 | Static, cacheable |
| Rendered profile | 250 | SQLite → deterministic render |
| Retrieved memories (4 × 40) | 160 | Chroma, floored at 0.45 |
| Recent turns (8 exchanges × 300) | 2,400 | Redis working set |
| Current user message | 80 | — |
| Total input | 3,240 | |
| Reserved output | 800 | max_tokens |
On a small fast model priced at 1 dollar per million input tokens and 5 per million output, one turn costs 3,240×1/106=0.324 cents of input plus roughly 300×5/106=0.15 cents of output — about 0.47 cents per turn. Five hundred turns of development and testing comes to roughly 2.40 dollars. Embeddings run locally on a MiniLM-class model, so they cost nothing.
Decide the token budget before you write the first line of the bot. Every component underneath it is a negotiation over a fixed number, and a negotiation with no number produces a project you stop running because it costs too much to test.
Compare that with the version of this project that skips the budget and just appends every turn: by turn 200 a single request is over 60,000 tokens, one turn costs 6 cents, and the same 500 turns of testing cost about 20 dollars. The budget is not an optimisation you add later; it is the difference between a project you can iterate on and one you stop running.
Build order
Build in stages that each end with something runnable. Resist the urge to write all six layers before running anything — memory bugs are only visible in a working loop.
Stage 1 — a bot with no memory at all
A CLI loop that sends the system prompt and the current message, prints the reply, and exits. No history. Ten minutes of work, and it gives you a baseline that unambiguously fails the acceptance test — which is exactly what you want to watch improve.
Stage 2 — the working set and the budget
Keep recent turns in Redis under a per-session key, and trim by token count before every call.
1# budget.py2from anthropic import Anthropic3client = Anthropic()45def count_request(system, messages, model) -> int:6 """Count the real request, including scaffolding -- not just the text."""7 return client.messages.count_tokens(8 model=model, system=system, messages=messages).input_tokens910def trim(messages, budget, model, system):11 """Drop whole exchanges from the front until the request fits."""12 msgs = list(messages)13 while msgs and count_request(system, msgs, model) > budget:14 del msgs[0:2] # a user+assistant pair, never a half pair15 while msgs and msgs[0]["role"] != "user":16 del msgs[0] # first message must be from the user17 return msgsDeleting in pairs is not cosmetic. Cutting between an assistant turn and the user turn that answers it can orphan a tool result and produce a 400 from the API, and it can leave an assistant message first, which most chat APIs reject.
One practical cost: count_request is a network call, and this loop makes one per deleted exchange. That is fine for a project, but in a service record each exchange's token count when you store it, count the full request once, and subtract the stored counts as you drop exchanges.
Trim in chunks rather than one message per turn. If you evict continuously, the prompt prefix changes on every single request and prompt caching never hits. Let the working set grow to the budget, then cut it back to half.
Stage 3 — vector memory
Now the acceptance test becomes achievable. Write each exchange's extracted content to Chroma; retrieve the top few before each call.
1# memory/vectors.py2def recall(user_id: str, query: str, k: int = 4, floor: float = 0.45):3 res = collection.query(4 query_texts=[query],5 n_results=k * 4, # over-fetch, then re-rank6 where={"user_id": user_id}, # in the query, never after it7 include=["documents", "metadatas", "distances"],8 )9 scored = [10 {"text": d, "meta": m, "score": 1.0 - dist}11 for d, m, dist in zip(res["documents"][0], res["metadatas"][0],12 res["distances"][0])13 ]14 scored = [s for s in scored if s["score"] >= floor]15 return mmr_select(query, scored, k=k, lam=0.7)Two things there will save you a day each. The where clause is inside the query, so another user's rows never leave the database. And the over-fetch-then-MMR step stops four near-identical restatements of the same fact from consuming all four slots.
Stage 4 — structured facts and conflict handling
This is the stage that makes the salmon test pass. Extraction produces triples; assertion applies cardinality rules.
1# memory/extract.py2EXTRACT_PROMPT = """Extract durable facts about the user from this exchange.34Return JSON: [{"predicate": str, "object": str, "confidence": 0-1,5 "supersedes_previous": bool}]67Predicates you may use:8 single-valued: city, country, timezone, employer, job_title, preferred_name9 multi-valued: allergy, dietary_restriction, language, skill, current_project1011Rules:12- Only facts about the USER, stated by the user.13- "actually", "not any more", "I've started" => supersedes_previous: true14- Nothing durable in this exchange => return []15- Prefer [] over guessing.1617EXCHANGE:18{exchange}"""Send that list shape as a JSON schema through your provider's structured-output mode rather than relying on the prompt alone, so a malformed reply never reaches assert_fact.
1# memory/facts.py2SINGLE_VALUED = {"city", "country", "timezone", "employer",3 "job_title", "preferred_name"}45def assert_fact(user_id, predicate, obj, confidence, session_id,6 supersedes=False):7 now = utcnow()8 single = predicate in SINGLE_VALUED910 if single or supersedes:11 rows = db.query(12 "SELECT * FROM facts WHERE user_id=? AND predicate=? "13 "AND valid_to IS NULL", (user_id, predicate))14 for row in rows:15 if row["object"] == obj:16 db.execute("UPDATE facts SET confidence=?, last_seen=? "17 "WHERE fact_id=?",18 (max(row["confidence"], confidence), now,19 row["fact_id"]))20 return "reconfirmed"21 db.execute("UPDATE facts SET valid_to=? WHERE fact_id=?",22 (now, row["fact_id"]))23 vectors.delete(ids=[row["fact_id"]]) # <-- the easy one to miss2425 fid = new_id()26 db.execute("INSERT INTO facts VALUES (?,?,?,?,?,NULL,?,?,?)",27 (fid, user_id, "user", predicate, obj, now, confidence,28 session_id))29 vectors.add(ids=[fid], documents=[f"{predicate}: {obj}"],30 metadatas=[{"user_id": user_id, "fact_id": fid,31 "predicate": predicate}])32 return "asserted"The commented line is the single most commonly missed step in this whole project. Closing the SQLite row without deleting the Chroma vector leaves the old text retrievable, and the bot keeps recommending vegetarian lunches to someone who told it three sessions ago that they eat fish. The relational table looks perfectly correct while you debug, which is why it takes people so long to find.
Stage 5 — session lifecycle and caching
Sessions start on first message and end after 30 minutes of silence. The end is where consolidation runs.
1# session.py2IDLE_SECONDS = 180034def end_session(session_id):5 transcript = redis.lrange(f"turns:{session_id}", 0, -1)6 facts = consolidate(transcript) # one call over the whole session7 for f in facts:8 assert_fact(**f)9 profile.rerender(user_id_of(session_id))10 cache.bump_version(user_id_of(session_id)) # invalidates every cached recall11 archive(transcript)12 redis.delete(f"turns:{session_id}")The version bump is the cheap trick worth internalising. Cached retrieval results are keyed as ret:{user}:v{n}:{hash}. Incrementing n makes every previous key unreachable in one atomic operation, so new facts show up immediately and the stale entries expire on their own. The alternative — scanning Redis for keys to delete — is slow, racy, and one forgotten pattern away from serving stale memories forever.
Stage 6 — evaluation
Assemble 30 to 50 labelled queries against a seeded store: the query, and the fact IDs that should come back. Then measure recall@k, precision@k and MRR after every change. Without this, tuning the floor and λ is guesswork, and the effects you are hunting are small enough to be invisible by feel.
Tests that actually catch the bugs
The tests below are ordered by how much pain each one prevents. Every single one corresponds to a bug that ships regularly.
| Test | What it does | Bug it catches |
|---|---|---|
test_persists_across_processes | Write a fact, tear down every object, rebuild from config, recall | State living in a Python object rather than a store — passes in one process, fails in production |
test_user_isolation | Seed distinctive facts for two users; assert A's query never returns B's row | Post-filtering instead of pre-filtering. This is a data breach, not a bug |
test_supersession | Assert city=Munich, then city=Berlin; assert exactly one live row and that Chroma no longer returns Munich | The forgotten vector deletion |
test_multivalued_append | Assert two allergies; assert both are live | Cardinality treated as single-valued — erases a safety-critical fact |
test_dedup | Assert the same fact five times; assert one row and a raised confidence | Retrieval collapse: five copies fill every slot |
test_budget_never_exceeded | Replay a synthetic 200-turn conversation; assert every request is under budget | Message-count trimming defeated by one large turn |
test_degrades_without_redis | Point Redis at a dead port; assert the bot still replies | Memory as a hard dependency — a cache blip becomes an outage |
test_deletion_cascades | Delete one fact; assert it is gone from SQLite, Chroma, the profile and the cache | Facts that resurface after a user asked you to forget them |
1def test_supersession(store):2 facts.assert_fact("u1", "city", "Munich", 0.9, "s1")3 facts.assert_fact("u1", "city", "Berlin", 0.9, "s2")45 live = facts.current("u1", "city")6 assert len(live) == 1 and live[0]["object"] == "Berlin"78 # the half everyone forgets9 hits = vectors.recall("u1", "where does the user live?", k=5, floor=0.0)10 assert not any("Munich" in h["text"] for h in hits)A memory system that has never been run from a cold start in a second process has not been tested at all — almost every bug in this project hides behind a Python object that happened to still be in scope.
How to know it is finished
| Criterion | Weight | Passing looks like |
|---|---|---|
| Cross-session recall | 25% | Both opening acceptance tests pass from a cold start |
| Conflict handling | 20% | Supersession works in both stores; multi-valued predicates append |
| Budget discipline | 15% | No request exceeds budget across a 200-turn replay; trimming is chunked |
| Retrieval quality | 15% | Measured recall@4 and precision@4 on a labelled set; MMR and a floor in place |
| Isolation and deletion | 15% | Pre-filtered queries; deletion cascades across all four places |
| Degradation | 10% | Bot answers with Redis and Chroma both unavailable; degraded rate is logged |
Where it will break
| Symptom | Likely cause | First thing to check |
|---|---|---|
| Nothing is ever recalled | Score floor set against the wrong metric | Print raw distances for a known-good pair; confirm the distance-to-similarity conversion |
| Everything scores 0.99 | Query and documents embedded by different models | Pin one model name in config.py; re-embed the store |
| Bot recites old facts confidently | Vector not deleted on supersession | Query Chroma directly for the superseded text |
| Same fact four times in the prompt | No write-time dedup, or MMR disabled | Log the four retrieved texts and their pairwise similarity |
| Costs far above the estimate | Prompt cache never hits | cache_read_input_tokens on a real response; look for a timestamp in the prefix |
| Bot volunteers memories nobody asked about | No floor, or a prompt implying memories must be used | Raise the floor; add "use these only where relevant" |
| Works on 20 memories, useless at 2,000 | Only ever tested on a toy store | Seed 2,000 synthetic memories and re-run the labelled set |
That last row deserves emphasis. Retrieval quality is a function of store size, and almost every memory project is developed against a store small enough that any retrieval strategy looks fine. Seed a realistic store early — synthetic memories generated from templates are perfectly adequate — because a system tuned on twenty memories will be re-tuned from scratch at two thousand.
Once the acceptance tests pass
Four extensions, each of which teaches something the base project does not.
Add time-aware decay by predicate kind. Multiply similarity by e−λ⋅age with a per-kind half-life: never for allergies and identity, 90 days for projects, 7 days for transient state. Then confirm with a test that an allergy survives a simulated two-year gap. Uniform decay is a common and genuinely dangerous shortcut.
Add a /memories command. Let the user list, edit and delete what the system has stored. It takes an hour, it exercises the deletion cascade properly, and it turns an opaque system into one people will trust. It also, reliably, reveals extraction bugs you had no idea were there — reading your own extractor's output is a humbling exercise.
Add bitemporal columns and a point-in-time query. Separate valid_from (when it became true) from asserted_at (when you learned it), then answer "what did the bot believe about this user on 1 May?" That query is what turns a bug report about a wrong answer from a mystery into a lookup.
Run the whole thing on a labelled evaluation set and write the numbers down. Recall@4, precision@4, MRR, mean tokens per request, mean cost per turn, and the percentage of turns served in degraded mode. Six numbers. Every serious decision you make afterwards — a different embedding model, a different floor, a bigger k — becomes a measurement against those six rather than an opinion. Most memory systems in production have never had these numbers computed even once, which is precisely why so many of them quietly serve stale facts to users who stopped mentioning it.