LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is a context window?


One legal-summariser request, in tokensSystemprompt — 1,200Retrievedclauses — 150,000Earlierchat — 8,000New question — 200Thinking andanswer — 20,000topbottom179,400 tokens of a 200,000-token window.
The answer's space is reserved last and squeezed first — fill the window with clauses and the model has no room to think.

What you need to know

What fills the window

A legal summariser's request might look like this:

Text
System prompt and rules              1,200 tokensRetrieved contract clauses         150,000 tokensEarlier chat with the lawyer         8,000 tokensThe new question                       200 tokensReserved for thinking and answer    20,000 tokens-------------------------------------------------Total                              179,400 tokens

The model sees only these tokens. Anything not in the window — yesterday's chat, the rest of the contract — does not exist for it.

Why long context costs so much

  • Money — input is billed per token. At an assumed $3 per million input tokens, the 179,400-token call above costs about $0.54 before output. Asking 20 questions about the same contract costs about $11 unless you use prompt caching, which bills a repeated prefix at a reduced rate.
  • Latency — the whole prompt must be processed before the first output token appears.
  • Memory — the KV cache grows with every token. For an 8B model with grouped-query attention (about 128 KB per token), 1 million tokens of context is about 128 GB of cache for a single request.

Big windows versus good use

  • Lost in the middle — research in 2023 showed models answer better when the key fact is near the start or end of a long context than in the middle.
  • Distraction — irrelevant text can pull the model towards the wrong answer.
  • Effective length — quality on hard tasks often drops well before the advertised limit. Test with your own documents.

Managing it

  • Retrieve only the relevant chunks, and put the most important ones near the start or end.
  • Summarise old conversation turns; keep the last few verbatim.
  • Trim boilerplate: repeated disclaimers, HTML, long tool outputs.
  • Reserve space for the output and, on reasoning models, for thinking tokens.

A real-life example

A bank's support chat in Hinglish can run for 60 turns while a customer disputes several transactions. The team first sent the whole history each turn. By turn 50 each request was 30,000 tokens, replies took several seconds to start, and the bot began mixing up two disputed amounts (Rs 2,340 and Rs 3,240) mentioned 40 turns apart.

They change the design: keep the last 6 turns verbatim, replace older turns with a running summary that lists each open dispute with its amount and reference number, and fetch transaction details from the database when needed. Requests drop to about 4,000 tokens, the first word appears faster, and the mix-up stops because the key facts are now short, structured and near the end of the prompt.

Follow-up questions to expect

  • "Does a 1M-token window remove the need for RAG?" — No. Retrieval is cheaper, faster, easier to keep fresh and often more accurate; long context is useful when the task needs one whole document at once.
  • "What happens if you exceed the window?" — The API returns an error, or a framework silently truncates, usually the oldest messages. Silent truncation is the dangerous case.
  • "Do output tokens count toward the window?" — Yes; the prompt plus the maximum output (and thinking) must fit.