RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How does poor chunking lead to hallucinations?


How the headphone returns answer went wrongPage introfills the1,000-char chunkRule in chunk1, exceptionsin chunk 2Exceptionschunkranks ninthModel: 10days, forheadphones tooFix: split by heading so the rule and its exceptions stay together.
The model treated a fragment as the whole policy — no prompt change could have added the missing exception.

What you need to know

The model does not know that a chunk is a fragment. It treats whatever it is given as the complete truth and fills gaps with plausible text.

The common failure patterns

Chunking defectWhat the model seesWhat it says
Table split mid-wayRows for 1 to 3 years, not 5 yearsGuesses the 5-year rate from the pattern
Rule without its subject"This must be filed within 30 days"Applies the 30 days to whatever was asked
Heading in previous chunkNumbers without "2023 rates" above themPresents old rates as current
Exception in next chunk"Refunds within 10 days"Omits "except for electronics: 7 days"
Chunk too largeRight chunk ranked 7th, outside top-5Answers from a similar, wrong chunk
Nothing relevant retrievedLoosely related textFalls back on training memory

Why it looks like a model problem

The final answer is wrong, the model wrote it, and it sounds sure. So teams blame the model or the prompt. Printing the retrieved chunks almost always shows the cause: the evidence was incomplete, and the model filled the gap.

Fixes that work

  1. Split on structure first — headings, numbered clauses, FAQ entries, table boundaries.
  2. Keep tables whole, or repeat the header row in each piece.
  3. Carry context into each chunk — prepend the document title and heading path, or use contextual retrieval, where an LLM writes a short context line for each chunk. Anthropic reported that, on their test sets, adding such context to both embeddings and BM25 cut top-20 retrieval failures roughly in half, and adding a reranker on top cut them further.
  4. Let the model refuse — the prompt must allow "the documents do not say".
  5. Check faithfulness — an evaluation step that flags answer claims not supported by the retrieved text.

A real-life example

An e-commerce returns assistant answers "Can I return headphones after 8 days?" with "Yes, returns are accepted within 10 days of delivery." The actual policy page says:

Text
Returns are accepted within 10 days of delivery.Exceptions:- Electronics and accessories: 7 days, only if defective.- Innerwear and personal care: not returnable.

The splitter cut after the first line because the page's 1,000-character chunk filled up with the page intro before it. The "Exceptions" list became the start of the next chunk, which did not mention returns in its first line and ranked ninth.

The fix: split the policy page by its headings so the rule and its exceptions stay together (about 180 tokens), and prepend "Returns policy — time limits and exceptions" to each chunk. The answer becomes "Headphones count as electronics, so they can be returned within 7 days, and only if defective." No prompt change was needed.

Follow-up questions to expect

  • "How do you find these cases at scale?" — Log retrieved chunks with each answer, and run an automated faithfulness check. Low-faithfulness answers point to the chunks to inspect.
  • "Would a bigger k fix it?" — Sometimes, but it adds noise and cost. Fixing the split fixes the cause.
  • "Is contextual retrieval worth the cost?" — For corpora where chunks are ambiguous on their own (contracts, reports, policies), usually yes; ingest cost is paid once per chunk, and caching the document makes each call cheap.