Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Add memory (short-term + long-term) to an agent.
What you need to know
A language model has no memory between calls; "memory" is whatever you put back into the prompt. There are two different questions to answer:
Short-term memory
- "What did we just say?"
- Last N turns, verbatim, in order
- Bounded window, oldest evicted first
- Lives for one conversation
Long-term memory
- "What do I know about this user?"
- Extracted facts, retrieved by relevance
- Grows over time; needs deletion and updates
- Survives across conversations
Evict in turns, not messages. A turn is a user message plus the assistant's reply (plus any tool calls in between). Dropping only the user half leaves an answer to a question that is gone, and dropping a tool_use without its tool_result makes the API reject the request.
Writing to long-term memory needs a policy. If you store every message, retrieval returns chit-chat. Usually a cheap model call after each turn extracts stable facts ("vegetarian", "lives in Bengaluru") and skips transient ones ("is hungry right now").
1import time, uuid2from collections import deque3from collections.abc import Callable4import numpy as np56class ShortTermMemory:7 """The last max_turns (user, assistant) pairs, evicted a whole turn at a time."""89 def __init__(self, max_turns: int = 10) -> None:10 self.turns: deque[tuple[str, str]] = deque(maxlen=max_turns)1112 def add_turn(self, user: str, assistant: str) -> None:13 self.turns.append((user, assistant))1415 def messages(self) -> list[dict]:16 return [m for u, a in self.turns for m in17 ({"role": "user", "content": u}, {"role": "assistant", "content": a})]1819class LongTermMemory:20 """Durable facts, recalled by cosine similarity to the current query."""2122 def __init__(self, embed_fn: Callable) -> None:23 self.embed_fn = embed_fn24 self.facts: list[dict] = []25 self.matrix = np.zeros((0, 0), dtype=np.float32)2627 def _unit(self, text: str) -> np.ndarray:28 v = np.asarray(self.embed_fn([text])[0], dtype=np.float32)29 return v / (np.linalg.norm(v) + 1e-10)3031 def remember(self, text: str, **metadata) -> None:32 v = self._unit(text)33 self.matrix = v[None, :] if not self.facts else np.vstack([self.matrix, v])34 self.facts.append({"id": str(uuid.uuid4()), "text": text, "ts": time.time(), **metadata})3536 def recall(self, query: str, k: int = 3, min_score: float = 0.3) -> list[dict]:37 if not self.facts:38 return []39 scores = self.matrix @ self._unit(query)40 return [self.facts[i] for i in np.argsort(-scores, kind="stable")[:k]41 if scores[i] >= min_score]4243def build_messages(query: str, stm: ShortTermMemory, ltm: LongTermMemory) -> tuple[str, list]:44 """Return (system_text, messages) for the next model call."""45 facts = "\n".join(f"- {f['text']}" for f in ltm.recall(query))46 system = "You are a helpful assistant." + (f"\nKnown facts about the user:\n{facts}" if facts else "")47 return system, stm.messages() + [{"role": "user", "content": query}]The tricky parts:
deque(maxlen=...)holding tuples evicts a whole turn at once, in O(1). A deque of single messages with an odd count would cut a pair in half.- Facts go into the system text, clearly labelled, not disguised as conversation turns.
min_scorekeeps unrelated facts out of the prompt. Without it, every call includes the top 3 facts even when none apply.
Complexity: short-term append and evict are O(1); messages() is O(window). remember is one embedding call plus an O(n·d) copy from vstack (preallocate for large stores). recall is one embedding call, O(n·d) scoring and an O(n log n) sort. Memory is O(n·d) for n facts.
A real-life example
1VOCAB = ["veg", "bengaluru", "spicy"]2toy_embed = lambda ts: [[float(t.lower().count(w)) for w in VOCAB] for t in ts]3ltm = LongTermMemory(toy_embed)4ltm.remember("User is veg and never eats egg", source="turn 3")5ltm.remember("User lives in Bengaluru", source="turn 9")6stm = ShortTermMemory(max_turns=2)7for u, a in [("hi", "hello!"), ("any offers?", "20% off today"), ("thanks", "welcome")]:8 stm.add_turn(u, a)910system, msgs = build_messages("suggest a veg dinner, not spicy", stm, ltm)11print(system)12# You are a helpful assistant.13# Known facts about the user:14# - User is veg and never eats egg15print([m["content"] for m in msgs])16# ['any offers?', '20% off today', 'thanks', 'welcome', 'suggest a veg dinner, not spicy']Step by step: the window holds 2 turns, so adding the third turn evicted ("hi", "hello!") as a pair. The query embeds to [1, 0, 1] ("veg" and "spicy"). The veg fact [1, 0, 0] scores 0.707 and passes the 0.3 threshold; the Bengaluru fact [0, 1, 0] scores 0 and is left out.
A food-delivery assistant that remembers "vegetarian" across months, while forgetting last week's "I'm in a hurry", is exactly this split.
Follow-up questions to expect
- "The user says they moved to Pune — what happens to the Bengaluru fact?" — Extraction should detect the conflict and update or delete the old fact. At minimum, recall should prefer the newer timestamp when two facts conflict.
- "How do you handle a request to forget everything?" — Delete by user id from the fact store and any caches; it is a legal requirement under GDPR and India's DPDP Act, not only good manners.
- "What about summarising old turns instead of dropping them?" — That is summary memory; combine it with the window so old context survives in compressed form.