Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Write token counting and context window management logic.


What you need to know

A token is the unit a model reads: a word, part of a word, or punctuation. The context window is the maximum number of tokens in one request, and it covers the input and the output. If a model has a 200,000-token window and you set max_tokens to 8,000, your input can use at most 192,000.

Counting options:

methodaccuracycostuse it for
Provider endpoint (client.messages.count_tokens)exact for that modela network callenforcing hard limits before sending
Local heuristic (len(text) // 4)rough; worse for code and non-English textfreebudgeting inside loops
Another vendor's tokenizerwrongfreenothing — tokenizers differ between model families

Keeping the newest turns is the usual policy: the latest messages matter most. Two rules keep the conversation valid:

  • keep whole turns — a user message together with the reply to it;
  • the kept history must start with a user message.
Python
from anthropic import Anthropicclient = Anthropic()def count_tokens(messages: list[dict], system: str | None = None,                 model: str = "claude-opus-5") -> int:    """Exact input token count for this model, computed by the API."""    kwargs = {"model": model, "messages": messages}    if system:        kwargs["system"] = system    return client.messages.count_tokens(**kwargs).input_tokensdef estimate_tokens(text: str) -> int:    """Offline heuristic: about 4 characters per token for English prose."""    return max(1, len(text) // 4)

Fitting a conversation into the window:

Python
from collections.abc import Callabledef fit_to_window(system: str, history: list[dict], question: str, window: int,                  reserve_output: int = 2000,                  measure: Callable[[str], int] = estimate_tokens) -> tuple[list[dict], int]:    """Keep the newest whole (user, assistant) turns that fit. Returns (kept, tokens_used)."""    available = window - reserve_output - measure(system) - measure(question)    if available <= 0:        raise ValueError("system prompt and question alone exceed the window")    turns = [history[i:i + 2] for i in range(0, len(history), 2)]   # [user, assistant] pairs    kept: list[list[dict]] = []    used = 0    for turn in reversed(turns):                      # newest first        cost = sum(measure(m["content"]) for m in turn)        if used + cost > available:            break                                     # older turns would leave a gap        kept.append(turn)        used += cost    return [m for turn in reversed(kept) for m in turn], used

The tricky parts:

  • reserve_output comes off the top. A prompt that exactly fills the window leaves zero tokens for the answer.
  • Pairing by history[i:i + 2] assumes history alternates user, assistant, which the API requires anyway. Evicting a pair at a time means the kept part always starts with a user message.
  • break, not continue. Skipping one big turn and keeping older ones would leave a hole in the middle of the conversation, which confuses the model more than dropping the old part.
  • measure is a parameter, so tests can use a simple deterministic counter and production can use the real one.

Complexity: O(n) over the n messages, plus the cost of measure on each. Space O(n) for the kept list.

A real-life example

Using word count as the measure, so the arithmetic is visible, and a window of 30:

Python
words = lambda s: len(s.split())history = [    {"role": "user", "content": "I want to book a train to Jaipur"},          # 8    {"role": "assistant", "content": "Sure, which date and class?"},          # 5    {"role": "user", "content": "Friday, 3AC please"},                        # 3    {"role": "assistant", "content": "Two trains have 3AC seats on Friday."},  # 7]kept, used = fit_to_window("You book trains.", history, "Book the earlier one",                          window=30, reserve_output=10, measure=words)print([m["content"] for m in kept], used)# ['Friday, 3AC please', 'Two trains have 3AC seats on Friday.'] 10
stepvalue
available30 − 10 (output) − 3 (system) − 4 (question) = 13
newest turn (3 + 7)10, fits, used = 10
older turn (8 + 5)13, 10 + 13 = 23 is over 13 → stop

The model loses "Jaipur", which is why long chats combine this window with a running summary or long-term memory.

A travel app's booking assistant uses exactly this budget before every call, and the exact counter before sending anything close to the limit.

Follow-up questions to expect

  • "A single message is bigger than the whole window — now what?" — Dropping turns cannot help. Truncate that message, summarise it, or chunk it and process it in parts.
  • "Why not use tiktoken for Claude?" — tiktoken is OpenAI's tokenizer; token counts differ between model families, sometimes by 20% or more. Use the provider's counting endpoint for the model you call.
  • "The request worked yesterday and is too long today — why?" — Something grew: retrieved context, tool results or the system prompt. Log the token count per component so you can see which one.