Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 9: Context Window Limitations
What you need to know
The scenario: long conversations and large documents overflow the context window, and answers get truncated or worse.
Why more context is not the answer
Every token in the prompt costs money and adds prefill time before the first output token. Beyond a point, quality also falls: models tend to use information at the start and end of a long context better than information in the middle, and extra, weaker chunks add distraction. On many RAG tasks, accuracy rises for the first few chunks, then flattens or drops.
Budget and drop order
window 128k (model limit) — but we choose to use about 16k reserve 2k output (max_tokens) fixed 1.5k system prompt + citation rules (never dropped) history up to 4k (recent turns verbatim, older turns summarised) context up to 8k (reranked chunks, best first)drop order when over budget: oldest history -> lowest-ranked chunks -> compress remaining chunksNever drop the system prompt or citation rules; the model then forgets how to behave.
Shrink what must fit
| Technique | What it does |
|---|---|
| Rerank to top 3–5 | Keeps only the best evidence |
| Contextual compression | Keeps only the query-relevant sentences of each chunk |
| Parent-document retrieval | Matches on small chunks, returns only the section needed |
| Rolling summary of history | Keeps the last few turns verbatim, summarises older ones |
| Map-reduce over sections | For whole-document tasks: answer per section, then combine |
Put the most important material at the start of the context (and restate the question at the end), where models use it most reliably.
Find the peak
- Sweep — run the golden set with 2, 4, 6, 8, 12 and 20 chunks.
- Plot — answer quality, cost and p95 latency against chunk count.
- Choose — the smallest count at the quality peak.
- Monitor — p95 tokens per request, so the budget does not creep.
A real-life example
Scenario, numbers made up. An insurance assistant sends 20 chunks plus the full chat history. Long chats hit the limit and answers get cut off; p95 input is 60,000 tokens.
The team sets a 16,000-token budget, summarises history older than six turns, and sweeps chunk counts. Accuracy on a 300-question golden set peaks at 5 chunks (86%) and is slightly lower at 20 (83%). With 5 reranked chunks and compression, p95 input falls to about 9,000 tokens, cost per answer drops by more than 80%, and truncated answers disappear.
Follow-up questions to expect
- "When is a long context the right tool?" — When a task truly needs a whole document at once, such as comparing two contracts clause by clause; even then, test map-reduce against it.
- "How do you count tokens accurately?" — Use the model's tokenizer or the provider's token-counting endpoint; character estimates are wrong near the limit.
- "What about reasoning models?" — Their thinking tokens often count against the output budget, so reserve more output room.