Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is a context window?
What you need to know
What fills the window
A legal summariser's request might look like this:
System prompt and rules 1,200 tokensRetrieved contract clauses 150,000 tokensEarlier chat with the lawyer 8,000 tokensThe new question 200 tokensReserved for thinking and answer 20,000 tokens-------------------------------------------------Total 179,400 tokensThe model sees only these tokens. Anything not in the window — yesterday's chat, the rest of the contract — does not exist for it.
Why long context costs so much
- Money — input is billed per token. At an assumed $3 per million input tokens, the 179,400-token call above costs about $0.54 before output. Asking 20 questions about the same contract costs about $11 unless you use prompt caching, which bills a repeated prefix at a reduced rate.
- Latency — the whole prompt must be processed before the first output token appears.
- Memory — the KV cache grows with every token. For an 8B model with grouped-query attention (about 128 KB per token), 1 million tokens of context is about 128 GB of cache for a single request.
Big windows versus good use
- Lost in the middle — research in 2023 showed models answer better when the key fact is near the start or end of a long context than in the middle.
- Distraction — irrelevant text can pull the model towards the wrong answer.
- Effective length — quality on hard tasks often drops well before the advertised limit. Test with your own documents.
Managing it
- Retrieve only the relevant chunks, and put the most important ones near the start or end.
- Summarise old conversation turns; keep the last few verbatim.
- Trim boilerplate: repeated disclaimers, HTML, long tool outputs.
- Reserve space for the output and, on reasoning models, for thinking tokens.
A real-life example
A bank's support chat in Hinglish can run for 60 turns while a customer disputes several transactions. The team first sent the whole history each turn. By turn 50 each request was 30,000 tokens, replies took several seconds to start, and the bot began mixing up two disputed amounts (Rs 2,340 and Rs 3,240) mentioned 40 turns apart.
They change the design: keep the last 6 turns verbatim, replace older turns with a running summary that lists each open dispute with its amount and reference number, and fetch transaction details from the database when needed. Requests drop to about 4,000 tokens, the first word appears faster, and the mix-up stops because the key facts are now short, structured and near the end of the prompt.
Follow-up questions to expect
- "Does a 1M-token window remove the need for RAG?" — No. Retrieval is cheaper, faster, easier to keep fresh and often more accurate; long context is useful when the task needs one whole document at once.
- "What happens if you exceed the window?" — The API returns an error, or a framework silently truncates, usually the oldest messages. Silent truncation is the dangerous case.
- "Do output tokens count toward the window?" — Yes; the prompt plus the maximum output (and thinking) must fit.