Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Write token counting and context window management logic.
What you need to know
A token is the unit a model reads: a word, part of a word, or punctuation. The context window is the maximum number of tokens in one request, and it covers the input and the output. If a model has a 200,000-token window and you set max_tokens to 8,000, your input can use at most 192,000.
Counting options:
| method | accuracy | cost | use it for |
|---|---|---|---|
Provider endpoint (client.messages.count_tokens) | exact for that model | a network call | enforcing hard limits before sending |
Local heuristic (len(text) // 4) | rough; worse for code and non-English text | free | budgeting inside loops |
| Another vendor's tokenizer | wrong | free | nothing — tokenizers differ between model families |
Keeping the newest turns is the usual policy: the latest messages matter most. Two rules keep the conversation valid:
- keep whole turns — a user message together with the reply to it;
- the kept history must start with a user message.
1from anthropic import Anthropic23client = Anthropic()45def count_tokens(messages: list[dict], system: str | None = None,6 model: str = "claude-opus-5") -> int:7 """Exact input token count for this model, computed by the API."""8 kwargs = {"model": model, "messages": messages}9 if system:10 kwargs["system"] = system11 return client.messages.count_tokens(**kwargs).input_tokens1213def estimate_tokens(text: str) -> int:14 """Offline heuristic: about 4 characters per token for English prose."""15 return max(1, len(text) // 4)Fitting a conversation into the window:
1from collections.abc import Callable23def fit_to_window(system: str, history: list[dict], question: str, window: int,4 reserve_output: int = 2000,5 measure: Callable[[str], int] = estimate_tokens) -> tuple[list[dict], int]:6 """Keep the newest whole (user, assistant) turns that fit. Returns (kept, tokens_used)."""7 available = window - reserve_output - measure(system) - measure(question)8 if available <= 0:9 raise ValueError("system prompt and question alone exceed the window")10 turns = [history[i:i + 2] for i in range(0, len(history), 2)] # [user, assistant] pairs11 kept: list[list[dict]] = []12 used = 013 for turn in reversed(turns): # newest first14 cost = sum(measure(m["content"]) for m in turn)15 if used + cost > available:16 break # older turns would leave a gap17 kept.append(turn)18 used += cost19 return [m for turn in reversed(kept) for m in turn], usedThe tricky parts:
reserve_outputcomes off the top. A prompt that exactly fills the window leaves zero tokens for the answer.- Pairing by
history[i:i + 2]assumes history alternates user, assistant, which the API requires anyway. Evicting a pair at a time means the kept part always starts with a user message. break, notcontinue. Skipping one big turn and keeping older ones would leave a hole in the middle of the conversation, which confuses the model more than dropping the old part.measureis a parameter, so tests can use a simple deterministic counter and production can use the real one.
Complexity: O(n) over the n messages, plus the cost of measure on each. Space O(n) for the kept list.
A real-life example
Using word count as the measure, so the arithmetic is visible, and a window of 30:
1words = lambda s: len(s.split())2history = [3 {"role": "user", "content": "I want to book a train to Jaipur"}, # 84 {"role": "assistant", "content": "Sure, which date and class?"}, # 55 {"role": "user", "content": "Friday, 3AC please"}, # 36 {"role": "assistant", "content": "Two trains have 3AC seats on Friday."}, # 77]8kept, used = fit_to_window("You book trains.", history, "Book the earlier one",9 window=30, reserve_output=10, measure=words)10print([m["content"] for m in kept], used)11# ['Friday, 3AC please', 'Two trains have 3AC seats on Friday.'] 10| step | value |
|---|---|
| available | 30 − 10 (output) − 3 (system) − 4 (question) = 13 |
| newest turn (3 + 7) | 10, fits, used = 10 |
| older turn (8 + 5) | 13, 10 + 13 = 23 is over 13 → stop |
The model loses "Jaipur", which is why long chats combine this window with a running summary or long-term memory.
A travel app's booking assistant uses exactly this budget before every call, and the exact counter before sending anything close to the limit.
Follow-up questions to expect
- "A single message is bigger than the whole window — now what?" — Dropping turns cannot help. Truncate that message, summarise it, or chunk it and process it in parts.
- "Why not use tiktoken for Claude?" — tiktoken is OpenAI's tokenizer; token counts differ between model families, sometimes by 20% or more. Use the provider's counting endpoint for the model you call.
- "The request worked yesterday and is too long today — why?" — Something grew: retrieved context, tool results or the system prompt. Log the token count per component so you can see which one.