Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Add summarization-based memory to reduce context size.


What you need to know

A sliding window forgets old turns completely. Summary memory compresses them instead:

Text
prompt = [summary of turns 1..k] + [turns k+1..n, verbatim]

Two design points decide whether it works:

  • What the summary must keep. A generic "summarise this" keeps the gist and drops exactly the details that matter later: the order id, the amount, the constraint stated once ("I'm allergic to peanuts"). The summarisation prompt must list what to preserve.
  • Incremental updates. Summarising the whole transcript each time costs more and more. Instead, pass the previous summary plus the newly evicted turns and ask for an updated summary.

Compounding loss. A summary of a summary of a summary drifts, like a message passed along a line of people. Keep the raw transcript in your database, so you can re-summarise from the source when needed.

Python
from collections.abc import Callabledef approx_tokens(text: str) -> int:    return max(1, len(text) // 4)SUMMARY_PROMPT = ("Update the running summary with the new turns. Keep every name, number, "                  "decision, constraint and open question. Drop greetings. Be concise and factual.")class SummaryMemory:    """Recent turns verbatim; older turns folded into a running summary."""    def __init__(self, llm_fn: Callable[[str], str], max_tokens: int = 3000,                 keep_recent: int = 6, measure: Callable[[str], int] = approx_tokens) -> None:        self.llm_fn, self.max_tokens = llm_fn, max_tokens        self.keep_recent, self.measure = keep_recent, measure        self.summary = ""        self.turns: list[dict] = []    def size(self) -> int:        return self.measure(self.summary) + sum(self.measure(t["content"]) for t in self.turns)    def add(self, role: str, content: str) -> None:        self.turns.append({"role": role, "content": content})        if self.size() > self.max_tokens and len(self.turns) > self.keep_recent:            self._compact()    def _compact(self) -> None:        old, recent = self.turns[:-self.keep_recent], self.turns[-self.keep_recent:]        transcript = "\n".join(f"{t['role']}: {t['content']}" for t in old)        new_summary = self.llm_fn(f"{SUMMARY_PROMPT}\n\nCurrent summary:\n"                                  f"{self.summary or '(none)'}\n\nNew turns:\n{transcript}").strip()        self.summary, self.turns = new_summary, recent       # only after the call succeeded    def context(self) -> str:        head = f"Summary of earlier conversation:\n{self.summary}\n\n" if self.summary else ""        return head + "\n".join(f"{t['role']}: {t['content']}" for t in self.turns)

The tricky parts:

  • State changes only after the model call returns. If the summariser times out, the exception leaves both the old summary and all turns intact; nothing is lost.
  • len(self.turns) > self.keep_recent — compaction only happens when there is something older than the recent window to fold in.
  • keep_recent should be even, so the verbatim part starts with a user message.

Complexity: size() is O(number of kept turns), which is bounded. A compaction is one LLM call over (summary + evicted turns). With a budget B and turns of average size t, compaction fires roughly every B/t turns, so the amortised cost per turn is small.

A real-life example

A stub summariser that just lists the facts it is given, a budget of 20 words and keep_recent=2:

Python
def fake_summariser(prompt: str) -> str:    new_turns = prompt.split("New turns:\n")[1]    old = prompt.split("Current summary:\n")[1].split("\n\nNew turns")[0]    facts = [line.split(": ", 1)[1] for line in new_turns.splitlines()]    return "; ".join(([] if old == "(none)" else [old]) + facts)mem = SummaryMemory(fake_summariser, max_tokens=20, keep_recent=2,                    measure=lambda s: len(s.split()))for role, text in [("user", "I need a laptop under Rs 60000"),                   ("assistant", "For coding or gaming?"),                   ("user", "Coding, and it must weigh under 1.5 kg"),                   ("assistant", "Try the Zenbook 14 at Rs 57990")]:    mem.add(role, text)print(mem.context())# Summary of earlier conversation:# I need a laptop under Rs 60000; For coding or gaming?## user: Coding, and it must weigh under 1.5 kg# assistant: Try the Zenbook 14 at Rs 57990
after addingturnssize (words)compaction?
message 117no
message 2211no
message 3319no — not over 20
message 4426yes: turns 1–2 folded into the summary

After compaction the context is 11 summary words plus 15 verbatim words. It is still over 20 because the two recent turns are kept whatever their size — keep_recent wins over the budget. The budget and the rupee amount survive in the summary, which is the part a careless summary prompt would lose.

Shopping assistants and long customer-support chats use this pattern so a 60-turn conversation still remembers the budget stated at turn 1.

Follow-up questions to expect

  • "Compaction makes one reply slow — how do you fix it?" — Run it in the background after sending the reply, so the next request finds the summary ready.
  • "The summary itself grows too big?" — Cap its size in the prompt and, when it passes the cap, re-summarise the summary (or re-summarise from the stored raw transcript).
  • "How do you test it?" — Plant a specific fact early ("my order is 88412"), run 50 turns, and check the model can still answer a question about it.