Course Content
Advanced RAG
3 sections · 38 lessons
Why is context compression important as prompt sizes grow?
What you need to know
Two reasons, not one
- Cost and latency. Providers bill per input token, and the prefill step — reading the prompt — takes longer as the prompt grows. Double the context and you roughly double that part of the bill.
- Quality. More context is not more knowledge. Studies such as "Lost in the Middle" (2023) showed that models use information at the start and end of a long prompt better than information in the middle. Later work on long-context models keeps finding that accuracy drops as the prompt fills with similar-looking but irrelevant passages. Newer models are better at this, but the effect has not disappeared.
The techniques, cheapest first
| Technique | What it does | Keeps exact wording? | Cost |
|---|---|---|---|
| Rerank and truncate | Retrieve 50, keep the best 5 | Yes | One reranker call |
| Deduplicate | Drop near-identical chunks | Yes | Almost free |
| Extractive filtering | Keep sentences that score high against the query | Yes | Small model or embeddings |
| Token pruning (LLMLingua family) | A small model deletes low-information tokens | Mostly | Small model |
| Abstractive summary | An LLM condenses each document for the query | No | An LLM call |
The right order is almost always: rerank first, dedupe, and compress further only if you are still over budget.
Why extractive is usually the sweet spot
Extractive filtering keeps the original sentences, so citations still point at real text, and nothing is paraphrased. Abstractive summarisation gets higher compression but inserts a generative step between the source and the answer — and that is exactly where a qualifier like "except for accounts opened before 2020" disappears.
Measure the right thing
Tokens saved is the easy number. The number that matters is answer quality. On a fixed evaluation set, compare, with and without compression:
- faithfulness — are the answer's claims supported by the context?
- answer accuracy against labelled answers;
- cost and p95 latency per query.
A change that cuts tokens by 60% but lowers faithfulness by several points is usually a bad trade.
A real-life example
An Indian telecom's support bot answers about 200,000 messages a day. Each query retrieves 8 help-article chunks of around 450 tokens — about 3,600 tokens of context per call, on top of the system prompt.
The team tries three changes on a 600-question evaluation set:
- Rerank to top 4 instead of top 8. Context halves. Accuracy is unchanged — the lower four chunks were rarely used.
- Add extractive filtering. Keep sentences with high similarity to the question, plus the chunk's title. Context falls to about 900 tokens. Accuracy is unchanged; faithfulness rises slightly because fewer distractors reach the model.
- Try abstractive summaries of each chunk with a small LLM. Context falls further, but accuracy drops on plan questions: the summaries round "₹299 for 28 days, 1.5 GB/day" to "a ₹299 monthly plan with daily data". A customer's validity question is now answered wrongly.
They ship 1 and 2 and drop 3. The input-token bill for retrieval context falls by roughly three-quarters, and median latency improves because prefill is shorter.
Follow-up questions to expect
- "Doesn't prompt caching make long prompts cheap anyway?" — Only for the fixed prefix. Retrieved context changes every query, so it is rarely cached, and caching does nothing for the quality cost of distractors.
- "Where do you put the most important chunk?" — At the start or end of the context block, not in the middle. Many teams put the best chunk first and repeat the question after the context.
- "When should you not compress?" — When the answer depends on exact wording or complete lists, such as legal clauses, and the reranked context already fits the budget.