Course Content
RAG Systems
12 sections · 66 lessons
What is chunk overlap and why is it used?
What you need to know
Why boundaries break meaning
A rule like "Employees in band 4 and above must serve 90 days" might be split as "…Employees in band 4 and above must" | "serve 90 days…". The first half has the subject but no number; the second has the number but no subject. A question about band 4 may retrieve either half and get a wrong or partial answer. Overlap gives the whole sentence a chance to appear in one chunk.
The cost, in real numbers
For a corpus of N characters, the number of chunks is about N / stride:
no overlap: N / 1000 chunksoverlap 200 (20%): N / 800 chunks -> 1.25x moreoverlap 500 (50%): N / 500 chunks -> 2.0x moreThat multiplier applies to embedding cost, storage and index memory. It also means neighbouring chunks share text, so for a query that matches the shared part, positions 1, 2 and 3 of the top-k can be three versions of the same passage.
Overlap in LangChain is not exact
In RecursiveCharacterTextSplitter, overlap is built from whole pieces at the current separator level. If two paragraphs are each longer than chunk_overlap, no overlap is added between them at all. So chunk_overlap is an upper limit, not a promise. Print a few neighbouring chunks to see what you actually get.
When overlap matters less
If you split on document structure (numbered clauses, headings, FAQ entries), boundaries fall where meaning already ends, and overlap adds little. Overlap matters most for long unstructured prose, such as transcripts and scanned letters.
Dealing with duplicates
- Deduplicate retrieved chunks that come from the same source and overlap in
start_indexrange, before building the prompt. - Or merge adjacent hits into one passage.
- Or use MMR (covered in the retrieval section) to pick diverse results.
A real-life example
An HR policy assistant chunks 600 policy pages with 800-token chunks and no overlap. An audit finds 7 of 50 test questions fail because the answering sentence was cut at a chunk boundary, like the band 4 rule above.
The team compares two fixes on their test set:
- Add 120 tokens of overlap (15%). Five of the seven failures are fixed. Chunk count rises by about 18%, and the top-5 often contains two neighbouring chunks with the same text, so they add a merge step.
- Split on section headings first, then apply the 800-token limit inside long sections, with 60 tokens of overlap. All seven are fixed, because the rules now sit inside one section chunk.
They ship the second option. Overlap stays, but as a small safety margin rather than the main fix.
Follow-up questions to expect
- "Is more overlap always safer?" — No. Above about 30%, you mostly pay for repeated text, and duplicates start pushing different useful chunks out of the top-k.
- "Should overlap be measured in tokens?" — Yes, when chunk size is in tokens. Use the same unit for both.
- "How does overlap interact with citations?" — A fact in the overlap region belongs to two chunks; merge them by
start_indexso the user sees one citation.