RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What are the roles of RecursiveCharacterTextSplitter?


Four chunks from chunk_size 80, overlap 20Notice period +band rule (74)Buy-out (7)The noticeperiod may be... (75)for theunserved days... (45)nullorphaned headingoverlap:words onlyNumbers are character lengths from a real run.
Overlap is built from whole pieces, so two paragraphs longer than chunk_overlap get no overlap at all.

What you need to know

The algorithm, step by step

  1. Try the biggest separator — split the text on "\n\n" (paragraphs).
  2. Recurse on oversized pieces — any piece still longer than chunk_size is split again on "\n", then " ", then "" (single characters) as a last resort.
  3. Merge — walk the pieces in order and join neighbours until adding the next one would pass chunk_size.
  4. Overlap — when starting a new chunk, carry over trailing pieces from the previous chunk, as long as they fit within chunk_overlap.
  5. Copy metadata — every chunk gets the parent's metadata, plus start_index if you ask for it.

Watch it run

Python
from langchain_text_splitters import RecursiveCharacterTextSplittertext = ("Notice period\n\n"        "Bands 1 to 3 serve 60 days. Band 4 and above serve 90 days.\n\n"        "Buy-out\n\n"        "The notice period may be bought out by paying basic salary "        "for the unserved days, with manager approval.")s = RecursiveCharacterTextSplitter(chunk_size=80, chunk_overlap=20, add_start_index=True)for d in s.create_documents([text]):    print(d.metadata["start_index"], repr(d.page_content))

Output:

Text
0   'Notice period\n\nBands 1 to 3 serve 60 days. Band 4 and above serve 90 days.'76  'Buy-out'85  'The notice period may be bought out by paying basic salary for the unserved'144 'for the unserved days, with manager approval.'

This small run teaches three things interviewers like to probe:

  • Headings can be orphaned. "Buy-out" became a 7-character chunk on its own, because joining it to the next paragraph would pass 80 characters. A chunk that says only "Buy-out" matches the query "buy-out" well and answers nothing.
  • Overlap is made of whole pieces. The last chunk repeats "for the unserved" because that paragraph was cut into words, and word-sized pieces fit inside 20 characters. The earlier chunks have no overlap, because a whole paragraph is bigger than 20 characters. So chunk_overlap is a maximum, not a guarantee.
  • chunk_size is a maximum, not a target. Chunks come out at 7, 74, 75 and 45 characters.

Characters or tokens?

chunk_size=1000 means 1,000 characters by default. With the cl100k_base tokenizer, plain English runs at about 4 to 5 characters per token, so that is roughly 200 to 250 tokens. Code runs closer to 3.5 characters per token, and Hindi in Devanagari script close to 1, so the same 1,000 characters of Hindi can be nearly 1,000 tokens. Embedding models have token limits, so measure in tokens:

Python
splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(    encoding_name="cl100k_base", chunk_size=400, chunk_overlap=50)

The tokenizer here is only an approximation if your embedding model uses a different one, but it is much closer than counting characters.

A real-life example

A bank's product-FAQ bot splits product terms with chunk_size=1000 characters. One chunk ends in the middle of the "Charges" section, and the next begins with "…waived for customers maintaining Rs 25,000". A customer asks "Is the debit card fee waived?" The retriever returns the second chunk, which does not name the card, and the bot says "Yes", which is true for a different card.

The team makes two changes. They split by headings first, so each product's "Charges" section is one unit, and use the recursive splitter only inside sections longer than 400 tokens. They also prepend the product name and section heading to each chunk's text. The wrong "Yes" disappears, because every chunk now says which card it is about.

Follow-up questions to expect

  • "Can you change the separators?" — Yes. For example, add ". " before " " so it prefers to break at sentence ends, or use RecursiveCharacterTextSplitter.from_language(...) for code.
  • "Does it guarantee chunks under chunk_size?" — Almost. With the empty-string separator as a last resort it can always cut, but a custom separator list without "" can leave oversized pieces.
  • "What does keep_separator do?" — It keeps the separator text attached to a piece rather than dropping it, which matters for separators that carry meaning, such as Markdown headings.