Context Management and Memory

Mini-Project: A Chatbot That Remembers


Here is the only thing this project has to do. You run it, tell it something, kill the process, run it again, and it still knows.

Bash
$ python -m mnemo --user aliceyou> I'm vegetarian, and I'm allergic to peanuts.bot> Noted. I'll keep both in mind.you> ^C$ python -m mnemo --user alice          # new process, new session, hours lateryou> suggest a quick lunchbot> A chickpea and roasted-red-pepper wrap with tahini takes about ten     minutes. Tahini is sesame, not peanut, so you're fine there.

Two details in that output are the entire assignment. The bot applied a constraint from a session that no longer exists in any process memory. And it went out of its way to confirm that tahini is not peanut — which means the allergy was not a vague impression, it was a retrievable, specific fact.

Now the harder test, the one that separates a working memory system from a store of stale assertions:

Bash
$ python -m mnemo --user aliceyou> actually I've started eating fish againbot> Got it, I'll update that.you> ^C$ python -m mnemo --user aliceyou> suggest a quick lunchbot> Pan-fried salmon with lemon and a green salad, about fifteen minutes.     Still no peanuts.

The vegetarian fact was superseded, not appended. The allergy was not touched. Most first attempts get this wrong in one of two ways: they keep both dietary facts and the bot contradicts itself, or they overwrite the whole dietary record and the peanut allergy quietly disappears. The second one is the dangerous bug, and it is the one a naive implementation ships.

Kill the process, run it again, it still knowsMemory thatsurvives a restartWorking setinside a token budgetVector recall ofolder sessionsStructured facts,conflicts resolvedConsolidation whenthe session endsCached prefix forthe fixed partsTests thatrestart the process
Each stage is only finished when a test proves the fact survives something — an overflow, a contradiction, or a restart.

What you are building

A command-line assistant called mnemo with three storage layers, each doing a job the others cannot.

LayerTechnologyHoldsWhy not one of the others
Working stateRedisActive session, recent turns, cached retrievalsNeeds TTLs and atomic counters; must be shared across processes
Semantic memoryChromaEmbedded text of episodes and factsOnly a vector index answers "what do I know that relates to this?"
Structured factsSQLiteVersioned (subject, predicate, object) triplesOnly a relational key detects that "Munich" and "Berlin" conflict

Three stores is not over-engineering. A vector index cannot detect contradiction, a relational table cannot answer a fuzzy question, and neither should be on the hot path for "what did the user just say".

Text
mnemo/  __main__.py        CLI loop, argument parsing  config.py          budgets, thresholds, model names -- all constants here  chat.py            prompt assembly + model call  budget.py          token counting and trimming  memory/    vectors.py       Chroma: embed, write, recall    facts.py         SQLite: assert_fact, current_facts, supersede    extract.py       turn -> candidate facts (small model)    profile.py       current_facts -> rendered profile block  session.py         Redis: start, append turn, end, consolidate  cache.py           Redis: embedding + retrieval caches, version keystests/  test_persistence.py  test_isolation.py  test_conflicts.py  test_budget.py       test_degradation.py

Everything tunable lives in config.py. Thresholds scattered through the codebase are how a retrieval system becomes impossible to evaluate — you cannot sweep a parameter that is written in four files.

The token budget, decided first

Write this table before writing code. It is the contract every other component has to respect.

ComponentBudgetSource
System prompt350Static, cacheable
Rendered profile250SQLite → deterministic render
Retrieved memories (4 × 40)160Chroma, floored at 0.45
Recent turns (8 exchanges × 300)2,400Redis working set
Current user message80—
Total input3,240
Reserved output800max_tokens

On a small fast model priced at 1 dollar per million input tokens and 5 per million output, one turn costs 3,240×1/106=0.3243{,}240 \times 1/10^6 = 0.324 cents of input plus roughly 300×5/106=0.15300 \times 5/10^6 = 0.15 cents of output — about 0.47 cents per turn. Five hundred turns of development and testing comes to roughly 2.40 dollars. Embeddings run locally on a MiniLM-class model, so they cost nothing.

Decide the token budget before you write the first line of the bot. Every component underneath it is a negotiation over a fixed number, and a negotiation with no number produces a project you stop running because it costs too much to test.

Compare that with the version of this project that skips the budget and just appends every turn: by turn 200 a single request is over 60,000 tokens, one turn costs 6 cents, and the same 500 turns of testing cost about 20 dollars. The budget is not an optimisation you add later; it is the difference between a project you can iterate on and one you stop running.

Build order

Build in stages that each end with something runnable. Resist the urge to write all six layers before running anything — memory bugs are only visible in a working loop.

Stage 1 — a bot with no memory at all

A CLI loop that sends the system prompt and the current message, prints the reply, and exits. No history. Ten minutes of work, and it gives you a baseline that unambiguously fails the acceptance test — which is exactly what you want to watch improve.

Stage 2 — the working set and the budget

Keep recent turns in Redis under a per-session key, and trim by token count before every call.

Python
# budget.pyfrom anthropic import Anthropicclient = Anthropic()def count_request(system, messages, model) -> int:    """Count the real request, including scaffolding -- not just the text."""    return client.messages.count_tokens(        model=model, system=system, messages=messages).input_tokensdef trim(messages, budget, model, system):    """Drop whole exchanges from the front until the request fits."""    msgs = list(messages)    while msgs and count_request(system, msgs, model) > budget:        del msgs[0:2]              # a user+assistant pair, never a half pair    while msgs and msgs[0]["role"] != "user":        del msgs[0]                # first message must be from the user    return msgs

Deleting in pairs is not cosmetic. Cutting between an assistant turn and the user turn that answers it can orphan a tool result and produce a 400 from the API, and it can leave an assistant message first, which most chat APIs reject.

One practical cost: count_request is a network call, and this loop makes one per deleted exchange. That is fine for a project, but in a service record each exchange's token count when you store it, count the full request once, and subtract the stored counts as you drop exchanges.

Trim in chunks rather than one message per turn. If you evict continuously, the prompt prefix changes on every single request and prompt caching never hits. Let the working set grow to the budget, then cut it back to half.

Stage 3 — vector memory

Now the acceptance test becomes achievable. Write each exchange's extracted content to Chroma; retrieve the top few before each call.

Python
# memory/vectors.pydef recall(user_id: str, query: str, k: int = 4, floor: float = 0.45):    res = collection.query(        query_texts=[query],        n_results=k * 4,                    # over-fetch, then re-rank        where={"user_id": user_id},         # in the query, never after it        include=["documents", "metadatas", "distances"],    )    scored = [        {"text": d, "meta": m, "score": 1.0 - dist}        for d, m, dist in zip(res["documents"][0], res["metadatas"][0],                              res["distances"][0])    ]    scored = [s for s in scored if s["score"] >= floor]    return mmr_select(query, scored, k=k, lam=0.7)

Two things there will save you a day each. The where clause is inside the query, so another user's rows never leave the database. And the over-fetch-then-MMR step stops four near-identical restatements of the same fact from consuming all four slots.

Stage 4 — structured facts and conflict handling

This is the stage that makes the salmon test pass. Extraction produces triples; assertion applies cardinality rules.

Python
# memory/extract.pyEXTRACT_PROMPT = """Extract durable facts about the user from this exchange.Return JSON: [{"predicate": str, "object": str, "confidence": 0-1,               "supersedes_previous": bool}]Predicates you may use:  single-valued: city, country, timezone, employer, job_title, preferred_name  multi-valued:  allergy, dietary_restriction, language, skill, current_projectRules:- Only facts about the USER, stated by the user.- "actually", "not any more", "I've started" => supersedes_previous: true- Nothing durable in this exchange => return []- Prefer [] over guessing.EXCHANGE:{exchange}"""

Send that list shape as a JSON schema through your provider's structured-output mode rather than relying on the prompt alone, so a malformed reply never reaches assert_fact.

Python
# memory/facts.pySINGLE_VALUED = {"city", "country", "timezone", "employer",                 "job_title", "preferred_name"}def assert_fact(user_id, predicate, obj, confidence, session_id,                supersedes=False):    now = utcnow()    single = predicate in SINGLE_VALUED    if single or supersedes:        rows = db.query(            "SELECT * FROM facts WHERE user_id=? AND predicate=? "            "AND valid_to IS NULL", (user_id, predicate))        for row in rows:            if row["object"] == obj:                db.execute("UPDATE facts SET confidence=?, last_seen=? "                           "WHERE fact_id=?",                           (max(row["confidence"], confidence), now,                            row["fact_id"]))                return "reconfirmed"            db.execute("UPDATE facts SET valid_to=? WHERE fact_id=?",                       (now, row["fact_id"]))            vectors.delete(ids=[row["fact_id"]])     # <-- the easy one to miss    fid = new_id()    db.execute("INSERT INTO facts VALUES (?,?,?,?,?,NULL,?,?,?)",               (fid, user_id, "user", predicate, obj, now, confidence,                session_id))    vectors.add(ids=[fid], documents=[f"{predicate}: {obj}"],                metadatas=[{"user_id": user_id, "fact_id": fid,                            "predicate": predicate}])    return "asserted"

The commented line is the single most commonly missed step in this whole project. Closing the SQLite row without deleting the Chroma vector leaves the old text retrievable, and the bot keeps recommending vegetarian lunches to someone who told it three sessions ago that they eat fish. The relational table looks perfectly correct while you debug, which is why it takes people so long to find.

Stage 5 — session lifecycle and caching

Sessions start on first message and end after 30 minutes of silence. The end is where consolidation runs.

Python
# session.pyIDLE_SECONDS = 1800def end_session(session_id):    transcript = redis.lrange(f"turns:{session_id}", 0, -1)    facts = consolidate(transcript)            # one call over the whole session    for f in facts:        assert_fact(**f)    profile.rerender(user_id_of(session_id))    cache.bump_version(user_id_of(session_id)) # invalidates every cached recall    archive(transcript)    redis.delete(f"turns:{session_id}")

The version bump is the cheap trick worth internalising. Cached retrieval results are keyed as ret:{user}:v{n}:{hash}. Incrementing n makes every previous key unreachable in one atomic operation, so new facts show up immediately and the stale entries expire on their own. The alternative — scanning Redis for keys to delete — is slow, racy, and one forgotten pattern away from serving stale memories forever.

Stage 6 — evaluation

Assemble 30 to 50 labelled queries against a seeded store: the query, and the fact IDs that should come back. Then measure recall@k, precision@k and MRR after every change. Without this, tuning the floor and λ is guesswork, and the effects you are hunting are small enough to be invisible by feel.

Tests that actually catch the bugs

The tests below are ordered by how much pain each one prevents. Every single one corresponds to a bug that ships regularly.

TestWhat it doesBug it catches
test_persists_across_processesWrite a fact, tear down every object, rebuild from config, recallState living in a Python object rather than a store — passes in one process, fails in production
test_user_isolationSeed distinctive facts for two users; assert A's query never returns B's rowPost-filtering instead of pre-filtering. This is a data breach, not a bug
test_supersessionAssert city=Munich, then city=Berlin; assert exactly one live row and that Chroma no longer returns MunichThe forgotten vector deletion
test_multivalued_appendAssert two allergies; assert both are liveCardinality treated as single-valued — erases a safety-critical fact
test_dedupAssert the same fact five times; assert one row and a raised confidenceRetrieval collapse: five copies fill every slot
test_budget_never_exceededReplay a synthetic 200-turn conversation; assert every request is under budgetMessage-count trimming defeated by one large turn
test_degrades_without_redisPoint Redis at a dead port; assert the bot still repliesMemory as a hard dependency — a cache blip becomes an outage
test_deletion_cascadesDelete one fact; assert it is gone from SQLite, Chroma, the profile and the cacheFacts that resurface after a user asked you to forget them
Python
def test_supersession(store):    facts.assert_fact("u1", "city", "Munich", 0.9, "s1")    facts.assert_fact("u1", "city", "Berlin", 0.9, "s2")    live = facts.current("u1", "city")    assert len(live) == 1 and live[0]["object"] == "Berlin"    # the half everyone forgets    hits = vectors.recall("u1", "where does the user live?", k=5, floor=0.0)    assert not any("Munich" in h["text"] for h in hits)

A memory system that has never been run from a cold start in a second process has not been tested at all — almost every bug in this project hides behind a Python object that happened to still be in scope.

How to know it is finished

CriterionWeightPassing looks like
Cross-session recall25%Both opening acceptance tests pass from a cold start
Conflict handling20%Supersession works in both stores; multi-valued predicates append
Budget discipline15%No request exceeds budget across a 200-turn replay; trimming is chunked
Retrieval quality15%Measured recall@4 and precision@4 on a labelled set; MMR and a floor in place
Isolation and deletion15%Pre-filtered queries; deletion cascades across all four places
Degradation10%Bot answers with Redis and Chroma both unavailable; degraded rate is logged

Where it will break

SymptomLikely causeFirst thing to check
Nothing is ever recalledScore floor set against the wrong metricPrint raw distances for a known-good pair; confirm the distance-to-similarity conversion
Everything scores 0.99Query and documents embedded by different modelsPin one model name in config.py; re-embed the store
Bot recites old facts confidentlyVector not deleted on supersessionQuery Chroma directly for the superseded text
Same fact four times in the promptNo write-time dedup, or MMR disabledLog the four retrieved texts and their pairwise similarity
Costs far above the estimatePrompt cache never hitscache_read_input_tokens on a real response; look for a timestamp in the prefix
Bot volunteers memories nobody asked aboutNo floor, or a prompt implying memories must be usedRaise the floor; add "use these only where relevant"
Works on 20 memories, useless at 2,000Only ever tested on a toy storeSeed 2,000 synthetic memories and re-run the labelled set

That last row deserves emphasis. Retrieval quality is a function of store size, and almost every memory project is developed against a store small enough that any retrieval strategy looks fine. Seed a realistic store early — synthetic memories generated from templates are perfectly adequate — because a system tuned on twenty memories will be re-tuned from scratch at two thousand.

Once the acceptance tests pass

Four extensions, each of which teaches something the base project does not.

Add time-aware decay by predicate kind. Multiply similarity by e−λ⋅agee^{-\lambda \cdot \text{age}} with a per-kind half-life: never for allergies and identity, 90 days for projects, 7 days for transient state. Then confirm with a test that an allergy survives a simulated two-year gap. Uniform decay is a common and genuinely dangerous shortcut.

Add a /memories command. Let the user list, edit and delete what the system has stored. It takes an hour, it exercises the deletion cascade properly, and it turns an opaque system into one people will trust. It also, reliably, reveals extraction bugs you had no idea were there — reading your own extractor's output is a humbling exercise.

Add bitemporal columns and a point-in-time query. Separate valid_from (when it became true) from asserted_at (when you learned it), then answer "what did the bot believe about this user on 1 May?" That query is what turns a bug report about a wrong answer from a mystery into a lookup.

Run the whole thing on a labelled evaluation set and write the numbers down. Recall@4, precision@4, MRR, mean tokens per request, mean cost per turn, and the percentage of turns served in degraded mode. Six numbers. Every serious decision you make afterwards — a different embedding model, a different floor, a bigger k — becomes a measurement against those six rather than an opinion. Most memory systems in production have never had these numbers computed even once, which is precisely why so many of them quietly serve stale facts to users who stopped mentioning it.