Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Add summarization-based memory to reduce context size.
What you need to know
A sliding window forgets old turns completely. Summary memory compresses them instead:
prompt = [summary of turns 1..k] + [turns k+1..n, verbatim]Two design points decide whether it works:
- What the summary must keep. A generic "summarise this" keeps the gist and drops exactly the details that matter later: the order id, the amount, the constraint stated once ("I'm allergic to peanuts"). The summarisation prompt must list what to preserve.
- Incremental updates. Summarising the whole transcript each time costs more and more. Instead, pass the previous summary plus the newly evicted turns and ask for an updated summary.
Compounding loss. A summary of a summary of a summary drifts, like a message passed along a line of people. Keep the raw transcript in your database, so you can re-summarise from the source when needed.
1from collections.abc import Callable23def approx_tokens(text: str) -> int:4 return max(1, len(text) // 4)56SUMMARY_PROMPT = ("Update the running summary with the new turns. Keep every name, number, "7 "decision, constraint and open question. Drop greetings. Be concise and factual.")89class SummaryMemory:10 """Recent turns verbatim; older turns folded into a running summary."""1112 def __init__(self, llm_fn: Callable[[str], str], max_tokens: int = 3000,13 keep_recent: int = 6, measure: Callable[[str], int] = approx_tokens) -> None:14 self.llm_fn, self.max_tokens = llm_fn, max_tokens15 self.keep_recent, self.measure = keep_recent, measure16 self.summary = ""17 self.turns: list[dict] = []1819 def size(self) -> int:20 return self.measure(self.summary) + sum(self.measure(t["content"]) for t in self.turns)2122 def add(self, role: str, content: str) -> None:23 self.turns.append({"role": role, "content": content})24 if self.size() > self.max_tokens and len(self.turns) > self.keep_recent:25 self._compact()2627 def _compact(self) -> None:28 old, recent = self.turns[:-self.keep_recent], self.turns[-self.keep_recent:]29 transcript = "\n".join(f"{t['role']}: {t['content']}" for t in old)30 new_summary = self.llm_fn(f"{SUMMARY_PROMPT}\n\nCurrent summary:\n"31 f"{self.summary or '(none)'}\n\nNew turns:\n{transcript}").strip()32 self.summary, self.turns = new_summary, recent # only after the call succeeded3334 def context(self) -> str:35 head = f"Summary of earlier conversation:\n{self.summary}\n\n" if self.summary else ""36 return head + "\n".join(f"{t['role']}: {t['content']}" for t in self.turns)The tricky parts:
- State changes only after the model call returns. If the summariser times out, the exception leaves both the old summary and all turns intact; nothing is lost.
len(self.turns) > self.keep_recent— compaction only happens when there is something older than the recent window to fold in.keep_recentshould be even, so the verbatim part starts with a user message.
Complexity: size() is O(number of kept turns), which is bounded. A compaction is one LLM call over (summary + evicted turns). With a budget B and turns of average size t, compaction fires roughly every B/t turns, so the amortised cost per turn is small.
A real-life example
A stub summariser that just lists the facts it is given, a budget of 20 words and keep_recent=2:
1def fake_summariser(prompt: str) -> str:2 new_turns = prompt.split("New turns:\n")[1]3 old = prompt.split("Current summary:\n")[1].split("\n\nNew turns")[0]4 facts = [line.split(": ", 1)[1] for line in new_turns.splitlines()]5 return "; ".join(([] if old == "(none)" else [old]) + facts)67mem = SummaryMemory(fake_summariser, max_tokens=20, keep_recent=2,8 measure=lambda s: len(s.split()))9for role, text in [("user", "I need a laptop under Rs 60000"),10 ("assistant", "For coding or gaming?"),11 ("user", "Coding, and it must weigh under 1.5 kg"),12 ("assistant", "Try the Zenbook 14 at Rs 57990")]:13 mem.add(role, text)14print(mem.context())15# Summary of earlier conversation:16# I need a laptop under Rs 60000; For coding or gaming?17#18# user: Coding, and it must weigh under 1.5 kg19# assistant: Try the Zenbook 14 at Rs 57990| after adding | turns | size (words) | compaction? |
|---|---|---|---|
| message 1 | 1 | 7 | no |
| message 2 | 2 | 11 | no |
| message 3 | 3 | 19 | no — not over 20 |
| message 4 | 4 | 26 | yes: turns 1–2 folded into the summary |
After compaction the context is 11 summary words plus 15 verbatim words. It is still over 20 because the two recent turns are kept whatever their size — keep_recent wins over the budget. The budget and the rupee amount survive in the summary, which is the part a careless summary prompt would lose.
Shopping assistants and long customer-support chats use this pattern so a 60-turn conversation still remembers the budget stated at turn 1.
Follow-up questions to expect
- "Compaction makes one reply slow — how do you fix it?" — Run it in the background after sending the reply, so the next request finds the summary ready.
- "The summary itself grows too big?" — Cap its size in the prompt and, when it passes the cap, re-summarise the summary (or re-summarise from the stored raw transcript).
- "How do you test it?" — Plant a specific fact early ("my order is 88412"), run 50 turns, and check the model can still answer a question about it.