Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
What happens if output exceeds the model’s context window, functionally and in quality?
What you need to know
The budget
prompt_tokens + generated_tokens ≤ context_windowgenerated_tokens ≤ max_output_tokens (often much smaller than the window)Two different limits catch people. A model may accept a very long prompt but cap its output at a few thousand tokens. Hitting either limit ends generation early.
Functional behaviour — depends on the stack
- Reject up front. The request fails with a validation error because prompt plus
max_tokensis over the limit. - Stop and flag. Generation runs until the limit and returns a finish reason such as
length(OpenAI-style) ormax_tokens(Anthropic-style). If your code ignores it, you ship a half-finished answer that looks complete. - Truncate or slide. Some chat frameworks drop the oldest turns; some local runtimes evict the earliest KV entries and keep going. The model then answers without information it was given earlier.
Quality behaviour — why past-the-end fails
A model has only ever seen positions up to its trained length. With RoPE, larger positions mean rotation angles the model never learned to read. Push past them and attention patterns break: perplexity rises sharply, and the text becomes repetitive or off-topic.
Context can be extended, but only with training. Position Interpolation (Chen et al., 2023) squeezes new positions into the trained range; NTK-aware scaling and YaRN (Peng et al., 2023) rescale RoPE frequencies unevenly. All of them work best with a short phase of fine-tuning on long sequences. Llama 3.1 reached 128K this way, by continued pretraining on longer data.
Inside the window is not uniform either
- Lost in the middle (Liu et al., 2023): on multi-document question answering, accuracy was highest when the answer was near the start or end of the prompt and dropped when it was in the middle.
- Effective length is often shorter than the advertised one. The RULER benchmark (Hsieh et al., 2024) found that many models claiming long windows lose accuracy well before the limit on harder tasks than simple "needle" retrieval.
A real-life example
An Indian-language news app translates English wire stories into Tamil with an 8K-context model. In their own measurements, Tamil output uses about three times as many tokens per word as English. A 2,500-word story is about 3,300 prompt tokens; the Tamil translation needs about 9,000 more. The API stops at 8K, returns finish_reason="length", and the app — which never checked that field — publishes a story that ends mid-paragraph.
The fix has three parts. Split the article by paragraph and translate in chunks of about 800 English words, passing the previous paragraph's translation as context so names stay consistent. Check the finish reason on every call and retry a chunk that was cut. And track tokens per word per language, because the same article costs very different amounts in Tamil, Hindi and English.
Follow-up questions to expect
- "What is the difference between
max_tokensand the context window?" —max_tokenscaps only the output you ask for; the window caps prompt and output together. You must satisfy both. - "How do models extend context length?" — By rescaling RoPE (interpolation, NTK-aware, YaRN) and fine-tuning on long sequences; without the fine-tuning, quality drops.
- "Isn't a sliding window the same as a longer context?" — No. It keeps generating, but tokens that fall out of the window are gone from direct attention, so the model can contradict what it was told earlier.