Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Add memory (short-term + long-term) to an agent.


What you need to know

A language model has no memory between calls; "memory" is whatever you put back into the prompt. There are two different questions to answer:

Short-term memory

  • "What did we just say?"
  • Last N turns, verbatim, in order
  • Bounded window, oldest evicted first
  • Lives for one conversation

Long-term memory

  • "What do I know about this user?"
  • Extracted facts, retrieved by relevance
  • Grows over time; needs deletion and updates
  • Survives across conversations

Evict in turns, not messages. A turn is a user message plus the assistant's reply (plus any tool calls in between). Dropping only the user half leaves an answer to a question that is gone, and dropping a tool_use without its tool_result makes the API reject the request.

Writing to long-term memory needs a policy. If you store every message, retrieval returns chit-chat. Usually a cheap model call after each turn extracts stable facts ("vegetarian", "lives in Bengaluru") and skips transient ones ("is hungry right now").

Python
import time, uuidfrom collections import dequefrom collections.abc import Callableimport numpy as npclass ShortTermMemory:    """The last max_turns (user, assistant) pairs, evicted a whole turn at a time."""    def __init__(self, max_turns: int = 10) -> None:        self.turns: deque[tuple[str, str]] = deque(maxlen=max_turns)    def add_turn(self, user: str, assistant: str) -> None:        self.turns.append((user, assistant))    def messages(self) -> list[dict]:        return [m for u, a in self.turns for m in                ({"role": "user", "content": u}, {"role": "assistant", "content": a})]class LongTermMemory:    """Durable facts, recalled by cosine similarity to the current query."""    def __init__(self, embed_fn: Callable) -> None:        self.embed_fn = embed_fn        self.facts: list[dict] = []        self.matrix = np.zeros((0, 0), dtype=np.float32)    def _unit(self, text: str) -> np.ndarray:        v = np.asarray(self.embed_fn([text])[0], dtype=np.float32)        return v / (np.linalg.norm(v) + 1e-10)    def remember(self, text: str, **metadata) -> None:        v = self._unit(text)        self.matrix = v[None, :] if not self.facts else np.vstack([self.matrix, v])        self.facts.append({"id": str(uuid.uuid4()), "text": text, "ts": time.time(), **metadata})    def recall(self, query: str, k: int = 3, min_score: float = 0.3) -> list[dict]:        if not self.facts:            return []        scores = self.matrix @ self._unit(query)        return [self.facts[i] for i in np.argsort(-scores, kind="stable")[:k]                if scores[i] >= min_score]def build_messages(query: str, stm: ShortTermMemory, ltm: LongTermMemory) -> tuple[str, list]:    """Return (system_text, messages) for the next model call."""    facts = "\n".join(f"- {f['text']}" for f in ltm.recall(query))    system = "You are a helpful assistant." + (f"\nKnown facts about the user:\n{facts}" if facts else "")    return system, stm.messages() + [{"role": "user", "content": query}]

The tricky parts:

  • deque(maxlen=...) holding tuples evicts a whole turn at once, in O(1). A deque of single messages with an odd count would cut a pair in half.
  • Facts go into the system text, clearly labelled, not disguised as conversation turns.
  • min_score keeps unrelated facts out of the prompt. Without it, every call includes the top 3 facts even when none apply.

Complexity: short-term append and evict are O(1); messages() is O(window). remember is one embedding call plus an O(n·d) copy from vstack (preallocate for large stores). recall is one embedding call, O(n·d) scoring and an O(n log n) sort. Memory is O(n·d) for n facts.

A real-life example

Python
VOCAB = ["veg", "bengaluru", "spicy"]toy_embed = lambda ts: [[float(t.lower().count(w)) for w in VOCAB] for t in ts]ltm = LongTermMemory(toy_embed)ltm.remember("User is veg and never eats egg", source="turn 3")ltm.remember("User lives in Bengaluru", source="turn 9")stm = ShortTermMemory(max_turns=2)for u, a in [("hi", "hello!"), ("any offers?", "20% off today"), ("thanks", "welcome")]:    stm.add_turn(u, a)system, msgs = build_messages("suggest a veg dinner, not spicy", stm, ltm)print(system)# You are a helpful assistant.# Known facts about the user:# - User is veg and never eats eggprint([m["content"] for m in msgs])# ['any offers?', '20% off today', 'thanks', 'welcome', 'suggest a veg dinner, not spicy']

Step by step: the window holds 2 turns, so adding the third turn evicted ("hi", "hello!") as a pair. The query embeds to [1, 0, 1] ("veg" and "spicy"). The veg fact [1, 0, 0] scores 0.707 and passes the 0.3 threshold; the Bengaluru fact [0, 1, 0] scores 0 and is left out.

A food-delivery assistant that remembers "vegetarian" across months, while forgetting last week's "I'm in a hurry", is exactly this split.

Follow-up questions to expect

  • "The user says they moved to Pune — what happens to the Bengaluru fact?" — Extraction should detect the conflict and update or delete the old fact. At minimum, recall should prefer the newer timestamp when two facts conflict.
  • "How do you handle a request to forget everything?" — Delete by user id from the fact store and any caches; it is a legal requirement under GDPR and India's DPDP Act, not only good manners.
  • "What about summarising old turns instead of dropping them?" — That is summary memory; combine it with the window so old context survives in compressed form.