Course Content
Context Management and Memory
3 sections · 6 lessons
Multi-Session Continuity — Memory Across Logins
On Monday a user tells the assistant: "I've just moved from Munich to Berlin." The extractor does its job, writes the memory, everything works.
On Wednesday, in a brand new session, they ask for a dentist recommendation. The assistant suggests one in Munich.
The retrieval log shows both memories were found:
rank score created memory 1 0.91 2025-03-02 "User lives in Munich, in the Schwabing district." 2 0.88 2026-08-17 "User has just moved from Munich to Berlin."Nothing malfunctioned. The eighteen-month-old memory is phrased more like the query — a plain statement of residence — so it scores higher. The vector index has no notion of truth and no notion of time. It ranks by geometry, and geometry does not know that people move.
Both memories were injected. The model saw a contradiction, had no basis for choosing, and went with the one listed first. A perfectly implemented retrieval system produced a confidently wrong answer, and this failure gets worse the longer the system runs, because the store fills up with an ever-longer trail of things that used to be true.
What changes when memory outlives the session
Within a single session, contradiction is rarely a problem — the transcript is ordered, and later statements visibly supersede earlier ones. Across sessions, that ordering is gone. Everything is a flat pile of assertions with timestamps nobody consults.
| Within one session | Across many sessions | |
|---|---|---|
| Ordering | Explicit in the message list | Only a timestamp field, if you stored one |
| Contradictions | Visible; later wins naturally | Invisible; both surface with similar scores |
| State lives in | Process memory | A database that must be designed |
| Failure mode | Forgetting | Confidently remembering something false |
| What the user expects | "It followed the conversation" | "It knows me, and it keeps up" |
Short-term memory fails by forgetting, which users notice and forgive. Long-term memory fails by remembering something that stopped being true, which users notice and do not forgive.
The session lifecycle
A session is one continuous stretch of interaction. Three moments matter, and most implementations only handle the middle one.
SESSION START load profile (compact, always injected) compute gap since last session decide greeting posture | vPER TURN retrieve relevant memories (top-k, filtered, floored) build prompt: system + profile + memories + recent turns generate reply lightweight fact extraction (fast model, cheap) | vSESSION END (or idle timeout) consolidate: extract facts from the full transcript detect and resolve conflicts against existing facts deduplicate, update confidences re-render the profile bump the cache version archive the raw transcriptThe end-of-session step is where nearly all the value lives, and it is the step people skip because sessions do not have a tidy "end" event. Users close a tab. Define the end yourself: an idle timeout of, say, 30 minutes, plus a sweep job that consolidates any session with no activity.
Why consolidate at the end rather than per turn
Both, actually — but for different reasons, and the costs differ enough to be worth working out. Take a typical session of 30 exchanges at 480 tokens each.
| Per-turn extraction | End-of-session consolidation | |
|---|---|---|
| Calls per session | 30 | 1 |
| Tokens per call | ~960 in, 60 out | ~14,400 in, 300 out |
| Cost per session | 30 × 0.378¢ = 11.3¢ | 4.32¢ + 0.45¢ = 4.8¢ |
| Sees the whole conversation? | No — one exchange at a time | Yes |
| Catches later corrections? | No — writes the wrong fact, then the right one | Yes — records only the final value |
| Available within the session? | Yes, immediately | No |
| On the user's critical path? | Yes, unless made async | No |
End-of-session consolidation is both cheaper (4.8 cents against 11.3) and more accurate, because it can see that the user corrected themselves at turn 22. Its only weakness is that a fact stated at turn 3 is not retrievable at turn 25 of the same session — which does not matter, because turn 3 is still in the verbatim window anyway.
The sensible arrangement is a cheap per-turn extractor writing to a short-lived working set for use inside the session, and a proper consolidation pass at the end that decides what becomes permanent.
The data model
Five tables, and the shape of the fact table is what makes conflict handling possible at all.
| Table | Holds | Retention |
|---|---|---|
users | Identity, preferences, consent flags | Life of the account |
sessions | id, user_id, started_at, ended_at, turn_count, summary | 1–2 years |
messages | Raw transcript, session_id, role, content, ts | 30–90 days, then archive or drop |
facts | Structured, versioned assertions (below) | Until superseded or deleted |
memories | Embedded text + vector for semantic recall | Mirrors facts, plus episodic entries |
The critical choice is storing facts as triples — subject, predicate, object — rather than as free text. Free text cannot be checked for conflict, because "lives in Munich" and "moved to Berlin" have no field in common. Triples collide on a key.
1CREATE TABLE facts (2 fact_id TEXT PRIMARY KEY,3 user_id TEXT NOT NULL,4 subject TEXT NOT NULL, -- almost always 'user'5 predicate TEXT NOT NULL, -- 'city', 'employer', 'allergy'6 object TEXT NOT NULL, -- 'Berlin', 'Acme Ltd', 'peanuts'7 valid_from TIMESTAMPTZ NOT NULL, -- when it became true in the world8 valid_to TIMESTAMPTZ, -- NULL = still true9 asserted_at TIMESTAMPTZ NOT NULL, -- when we learned it10 source TEXT NOT NULL, -- 'stated' | 'inferred' | 'imported'11 confidence REAL NOT NULL,12 session_id TEXT NOT NULL13);14CREATE INDEX ON facts (user_id, predicate, valid_to);Two separate time columns is not over-engineering. valid_from is when the fact became true in the world; asserted_at is when your system found out. They differ constantly — a user says in November that they moved in August — and keeping both lets you answer two very different questions: "where did they live in September?" and "what did we believe in September?" The second one is what you need when a user complains about a wrong answer from three weeks ago.
Conflict resolution
With triples, detecting conflict is a lookup on (user_id, predicate) where valid_to IS NULL. What to do about it depends entirely on the predicate's cardinality.
| Predicate | Cardinality | A new value means |
|---|---|---|
city, employer, job_title, timezone | One | Supersede — close the old row's valid_to |
allergy, language, skill, owns_device | Many | Append — both are true at once |
dietary_restriction | Many, revocable | Append; remove only on an explicit statement |
current_project | Few | Append, with an expiry; close on completion |
Getting cardinality wrong produces both classic bugs. Treating allergy as single-valued means learning about a shellfish allergy erases the peanut allergy — a safety failure. Treating city as multi-valued means the user lives in Munich and Berlin simultaneously, which is the failure that opened this lesson.
1SINGLE_VALUED = {"city", "country", "timezone", "employer",2 "job_title", "relationship_status", "preferred_name"}34def assert_fact(user_id, predicate, obj, source, confidence, session_id,5 valid_from=None):6 now = utcnow()7 valid_from = valid_from or now89 if predicate in SINGLE_VALUED:10 current = db.fetchone(11 "SELECT * FROM facts WHERE user_id=%s AND predicate=%s "12 "AND valid_to IS NULL", (user_id, predicate))1314 if current:15 if current["object"] == obj:16 bump_confidence(current["fact_id"], confidence)17 return "reconfirmed"1819 if not supersedes(new_source=source, new_conf=confidence,20 old=current):21 return "rejected"2223 db.execute("UPDATE facts SET valid_to=%s WHERE fact_id=%s",24 (valid_from, current["fact_id"]))25 retire_memory_vector(current["fact_id"])2627 insert_fact(user_id, predicate, obj, valid_from, now,28 source, confidence, session_id)29 return "asserted"The supersedes check is where judgement lives. A reasonable policy:
| Situation | Decision | Why |
|---|---|---|
| User states it directly, contradicting an inference | Supersede | Stated always beats inferred |
| User explicitly corrects ("no, actually...") | Supersede, confidence 1.0 | The clearest signal you will ever get |
| New inference contradicts an older stated fact | Reject; flag for confirmation | A guess must not overwrite a statement |
| Both stated, new one is newer | Supersede | Reality changes; recency is the tiebreak |
| Conflicts with a hard constraint (allergy, accessibility) | Never auto-resolve; ask | The downside of being wrong is harm, not annoyance |
| Both low confidence, no clear winner | Keep both, mark uncertain | Better to hedge in the reply than to guess silently |
Note retire_memory_vector in the code above. Superseding a fact in the relational table is only half the job — if the old text is still sitting in the vector store, retrieval will keep surfacing it and the Munich dentist comes straight back. Every supersession must delete or re-tag the corresponding vector.
Answering a question about the past
Because rows are closed rather than deleted, point-in-time queries are free:
1-- What did we hold true about this user on 1 May 2026?2SELECT predicate, object FROM facts3WHERE user_id = %s4 AND valid_from <= '2026-05-01'5 AND (valid_to IS NULL OR valid_to > '2026-05-01');Keeping the history costs almost nothing. A user who changes 30 single-valued facts over five years accumulates perhaps 90 rows — kilobytes — and in exchange you can explain any past answer the system gave.
The rendered profile
Retrieval handles the specific and the occasional. But some facts belong in every prompt: name, location, language, hard constraints, current projects. Retrieving those is wasteful and unreliable — you do not want the assistant's knowledge of a peanut allergy to depend on whether the user's question happened to embed near it.
So render them into a compact block, deterministically, from the current facts:
1PROFILE_SECTIONS = [2 ("Identity", ["preferred_name", "pronouns", "language", "timezone"]),3 ("Location", ["city", "country"]),4 ("Work", ["employer", "job_title"]),5 ("Constraints", ["allergy", "dietary_restriction", "accessibility_need"]),6 ("Active", ["current_project", "current_goal"]),7]89def render_profile(user_id) -> str:10 facts = current_facts(user_id) # valid_to IS NULL11 lines = []12 for heading, predicates in PROFILE_SECTIONS:13 vals = [f"{p}: {', '.join(facts[p])}" for p in predicates if facts.get(p)]14 if vals:15 lines.append(f"{heading}: " + "; ".join(vals))16 return "\n".join(lines)The economics are stark. Injecting every memory at session start — 900 memories at 40 tokens — is 36,000 tokens, about 10.8 cents per session at 3 dollars per million, and it consumes 28% of a 128,000-token window before the user says anything. Across 240 sessions that is 25.92 dollars for one user. A rendered profile of 400 tokens costs 0.12 cents per session, about 29 cents across the same 240 sessions, and leaves the window free.
Profile for what is always relevant; retrieval for what is sometimes relevant. Trying to do either job with the other tool is how memory systems become expensive and vague at the same time.
Render the profile only when facts change, cache it, and place it in the stable part of the prompt so it participates in prompt caching. A profile that changes once a fortnight costs you one cache miss a fortnight.
Session-aware behaviour
Users read continuity from small cues, and the strongest one is how the assistant handles the gap since last time. Treat the gap as a first-class input.
| Gap | Posture | What to do about volatile facts |
|---|---|---|
| < 30 min | Same session — just continue | Trust everything |
| 30 min – 24 h | Brief resumption, no fanfare | Trust everything |
| 1–30 days | "Welcome back" plus the last open thread | Trust; mention what was in progress |
| > 30 days | Light reintroduction | Confirm volatile facts before acting on them |
| > 6 months | Treat volatile facts as expired | Re-ask; keep only identity and hard constraints |
1def opening_context(user_id):2 last = last_session(user_id)3 if not last:4 return ""5 gap_days = (utcnow() - last["ended_at"]).days6 if gap_days < 1:7 return ""8 if gap_days <= 30:9 return (f"Last conversation was {gap_days} day(s) ago. "10 f"Open thread: {last['summary']}. "11 f"Reference it only if the user's message relates to it.")12 return ("It has been a while since the last conversation. "13 "Treat time-sensitive details as possibly outdated and "14 "confirm before relying on them.")The instruction "reference it only if the user's message relates to it" is doing important work. Without it, an assistant opens every session by reciting last week's topic, which reads as clumsy rather than attentive.
Graceful degradation
Memory is an enhancement. The assistant must work without it. The most common architectural mistake in long-term memory systems is making the memory store a hard dependency of the request path, which turns a Redis blip into a total outage.
| Failure | Naive behaviour | Correct behaviour |
|---|---|---|
| Vector store unreachable | Exception propagates; user sees an error | Reply without retrieved memories; log; keep the profile if cached |
| Embedding provider slow | Request hangs for 30 s | Hard timeout at ~300 ms, skip retrieval, continue |
| Profile store unavailable | Error, or a blank identity | Serve the last cached render, even if slightly stale |
| Consolidation fails at session end | Facts lost silently | Queue the transcript for retry; never block session close |
| Conflict cannot be resolved | Pick one arbitrarily | Include both, tell the model to ask |
1def safe_recall(user_id, query, timeout_ms=300):2 if breaker.is_open("memory"):3 return []4 try:5 with deadline(timeout_ms):6 hits = recall(user_id, query, k=5, floor=0.45)7 breaker.record_success("memory")8 return hits9 except (TimeoutError, StoreUnavailable) as e:10 breaker.record_failure("memory")11 log.warning("memory degraded", user_id=user_id, error=str(e))12 return [] # degraded, not brokenTrack the degraded rate as a first-class metric. An assistant silently running without memory for 8% of requests looks, from the outside, exactly like an assistant with a flaky personality — and no one will diagnose it as an infrastructure problem unless the number is on a dashboard.
Deletion, and what you are obliged to support
A user who says "forget that I told you about my divorce" is making a request you must be able to honour, and in many jurisdictions is exercising a legal right. That requirement reaches backwards into the data model, which is why it belongs here rather than in a compliance appendix.
Deletion has to cascade across every place the fact landed:
- The row in
facts— hard-deleted, not closed with avalid_to, because a closed row is still a record of it. - The vector in the memory store — matched by
fact_id, which is why the metadata field exists. - The rendered profile — re-render and overwrite.
- Any cached retrieval results — handled free by bumping the per-user cache version.
- Raw transcripts containing the statement — the reason
messageshas a retention policy rather than living forever.
If you cannot execute all five from a single user_id plus fact_id, the deletion is incomplete and the fact will resurface. Design for that on the first day; retrofitting traceability onto a populated store is not possible, because the links you need were never recorded.
Making this work in production
Store facts as triples, not as sentences. Everything else in this lesson depends on it. Conflict detection, cardinality rules, point-in-time queries, profile rendering, targeted deletion — all of them need a structured key, and none of them are possible over free text. If you start with free text you will rewrite the whole store later, and you will lose the history in the migration.
Every supersession must touch two stores. Close the row and retire the vector. A superseded fact left in the vector index is the exact bug this lesson opened with, and it is easy to ship because the relational side looks perfectly correct in the database console.
Define session end explicitly and consolidate there. A 30-minute idle timeout plus a sweeper. Consolidation at the end is cheaper than per-turn extraction (4.8 cents against 11.3 for a 30-exchange session) and strictly more accurate, because it sees corrections the user made later in the conversation.
Make memory optional in the request path. Timeout, circuit-break, degrade, log. A user with a slightly forgetful assistant is a user with a working assistant; a user staring at a 500 page is not.
Show your work to the user. An assistant that says "I have you in Berlin — is that still right?" after a three-month gap earns far more trust than one that either silently uses stale data or silently forgets. Let people see and correct what you have stored; a memory system users cannot inspect is one they will not rely on.