RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is context stuffing and why is it dangerous?


Same questions, two values of k2,0000.0066009,5000.02852,850Input tokensUSD per queryUSD per 100k queriesk = 5k = 30Assumed price: 3 USD per million input tokens, 300-token chunks, 500-token prompt.
Stuffing multiplies the bill on every request while pushing the answer-bearing chunk into the middle, where models use it least.

What you need to know

Modern models accept hundreds of thousands of tokens, so it is tempting to skip careful retrieval and send everything. There are four reasons not to.

1. Lost in the middle

The 2023 paper Lost in the Middle (Liu and others) placed one relevant passage at different positions in a long context. Accuracy was highest when the passage was at the start or the end, and dropped when it sat in the middle, a U-shaped curve. Newer models have improved, but position still matters, and long contexts still degrade answers in ways that are hard to predict.

2. Distraction by near-misses

An unrelated chunk is easy to ignore. A chunk about the wrong version of the right topic is not. Stuffing 30 chunks about "home loans" brings in the 2024 rate card, a different product's fees, and a blog post, all of which look relevant.

3. Cost and delay

Input tokens are paid on every request, and reading them (called prefill) takes time before the first output token. Here is the arithmetic with an assumed price of $3 per million input tokens and 300-token chunks:

SettingInput tokens per queryCost per queryCost at 100,000 queries per day
k = 51,500 + 500 prompt = 2,000$0.006$600
k = 309,000 + 500 prompt = 9,500$0.0285$2,850

Same questions, almost 5 times the bill. Prompt caching does not help much here, because the retrieved chunks change on every question.

4. It hides retrieval bugs

With k=50, something relevant is almost always somewhere in the prompt, so "retrieval found the answer" looks fine. But the real ranking might be poor, and you will not notice until you try to cut costs.

Stuffing

  • k = 30 to 50, no reranker
  • High recall, low precision
  • Long prompts, slow first token
  • Answer quality depends on luck of position

Precision

  • Retrieve 30 to 50, rerank, keep 3 to 6
  • High recall and high precision
  • Short prompts, fast and cheap
  • Weak retrieval shows up in metrics

A real-life example

A bank's product-FAQ bot is launched with k=25 "to be safe". Answers are slow (p95 time to first token around 4 seconds) and sometimes quote the wrong fee. A trace shows the question "What is the annual fee for the Platinum card?" pulls in chunks from four card products. The model picks the Gold card fee, which appears in chunk 2, while the Platinum fee is in chunk 14.

The team adds a cross-encoder reranker and keeps the top 4. On their 200-question test set, the right fee now appears in the context for the same share of questions as before, but the wrong-product answers mostly disappear, and input tokens per query fall by about 80%. They also add a metadata filter on product, which removes the other cards before ranking even starts.

Follow-up questions to expect

  • "With 1-million-token models, why not put the whole corpus in?" — For a small, stable corpus you sometimes can, with prompt caching. For a large or changing one, cost, delay, position effects and access control (you cannot show every user every document) still favour retrieval.
  • "How do you choose the final k?" — Measure answer quality and cost at k = 3, 5, 8 on an evaluation set and pick the smallest k before quality stops improving.
  • "What if one question really needs 20 chunks?" — That is a summarise or compare task. Use map-reduce summarisation or an agent that retrieves per sub-question, not one huge prompt.