Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 9: Context Window Limitations


A 16k budget inside a 128k windowRetrievedchunks, up to 8kHistory, up to 4kSystem andcitation rules, 1.5kReservedoutput, 2ktopbottomDrop order: oldest history, then lowest-ranked chunks, then compress.
Accuracy peaked at five chunks and was lower at twenty, so the bigger prompt cost more and answered worse.

What you need to know

The scenario: long conversations and large documents overflow the context window, and answers get truncated or worse.

Why more context is not the answer

Every token in the prompt costs money and adds prefill time before the first output token. Beyond a point, quality also falls: models tend to use information at the start and end of a long context better than information in the middle, and extra, weaker chunks add distraction. On many RAG tasks, accuracy rises for the first few chunks, then flattens or drops.

Budget and drop order

Text
window 128k (model limit) — but we choose to use about 16k  reserve   2k   output (max_tokens)  fixed     1.5k system prompt + citation rules   (never dropped)  history   up to 4k  (recent turns verbatim, older turns summarised)  context   up to 8k  (reranked chunks, best first)drop order when over budget: oldest history -> lowest-ranked chunks -> compress remaining chunks

Never drop the system prompt or citation rules; the model then forgets how to behave.

Shrink what must fit

TechniqueWhat it does
Rerank to top 3–5Keeps only the best evidence
Contextual compressionKeeps only the query-relevant sentences of each chunk
Parent-document retrievalMatches on small chunks, returns only the section needed
Rolling summary of historyKeeps the last few turns verbatim, summarises older ones
Map-reduce over sectionsFor whole-document tasks: answer per section, then combine

Put the most important material at the start of the context (and restate the question at the end), where models use it most reliably.

Find the peak

  1. Sweep — run the golden set with 2, 4, 6, 8, 12 and 20 chunks.
  2. Plot — answer quality, cost and p95 latency against chunk count.
  3. Choose — the smallest count at the quality peak.
  4. Monitor — p95 tokens per request, so the budget does not creep.

A real-life example

Scenario, numbers made up. An insurance assistant sends 20 chunks plus the full chat history. Long chats hit the limit and answers get cut off; p95 input is 60,000 tokens.

The team sets a 16,000-token budget, summarises history older than six turns, and sweeps chunk counts. Accuracy on a 300-question golden set peaks at 5 chunks (86%) and is slightly lower at 20 (83%). With 5 reranked chunks and compression, p95 input falls to about 9,000 tokens, cost per answer drops by more than 80%, and truncated answers disappear.

Follow-up questions to expect

  • "When is a long context the right tool?" — When a task truly needs a whole document at once, such as comparing two contracts clause by clause; even then, test map-reduce against it.
  • "How do you count tokens accurately?" — Use the model's tokenizer or the provider's token-counting endpoint; character estimates are wrong near the limit.
  • "What about reasoning models?" — Their thinking tokens often count against the output budget, so reserve more output room.