RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is chunk overlap and why is it used?


What you need to know

Why boundaries break meaning

A rule like "Employees in band 4 and above must serve 90 days" might be split as "…Employees in band 4 and above must" | "serve 90 days…". The first half has the subject but no number; the second has the number but no subject. A question about band 4 may retrieve either half and get a wrong or partial answer. Overlap gives the whole sentence a chance to appear in one chunk.

The cost, in real numbers

For a corpus of N characters, the number of chunks is about N / stride:

Text
no overlap:        N / 1000  chunksoverlap 200 (20%): N / 800   chunks  -> 1.25x moreoverlap 500 (50%): N / 500   chunks  -> 2.0x more

That multiplier applies to embedding cost, storage and index memory. It also means neighbouring chunks share text, so for a query that matches the shared part, positions 1, 2 and 3 of the top-k can be three versions of the same passage.

Overlap in LangChain is not exact

In RecursiveCharacterTextSplitter, overlap is built from whole pieces at the current separator level. If two paragraphs are each longer than chunk_overlap, no overlap is added between them at all. So chunk_overlap is an upper limit, not a promise. Print a few neighbouring chunks to see what you actually get.

When overlap matters less

If you split on document structure (numbered clauses, headings, FAQ entries), boundaries fall where meaning already ends, and overlap adds little. Overlap matters most for long unstructured prose, such as transcripts and scanned letters.

Dealing with duplicates

  • Deduplicate retrieved chunks that come from the same source and overlap in start_index range, before building the prompt.
  • Or merge adjacent hits into one passage.
  • Or use MMR (covered in the retrieval section) to pick diverse results.

A real-life example

An HR policy assistant chunks 600 policy pages with 800-token chunks and no overlap. An audit finds 7 of 50 test questions fail because the answering sentence was cut at a chunk boundary, like the band 4 rule above.

The team compares two fixes on their test set:

  1. Add 120 tokens of overlap (15%). Five of the seven failures are fixed. Chunk count rises by about 18%, and the top-5 often contains two neighbouring chunks with the same text, so they add a merge step.
  2. Split on section headings first, then apply the 800-token limit inside long sections, with 60 tokens of overlap. All seven are fixed, because the rules now sit inside one section chunk.

They ship the second option. Overlap stays, but as a small safety margin rather than the main fix.

Follow-up questions to expect

  • "Is more overlap always safer?" — No. Above about 30%, you mostly pay for repeated text, and duplicates start pushing different useful chunks out of the top-k.
  • "Should overlap be measured in tokens?" — Yes, when chunk size is in tokens. Use the same unit for both.
  • "How does overlap interact with citations?" — A fact in the overlap region belongs to two chunks; merge them by start_index so the user sees one citation.