Course Content
RAG Systems
12 sections · 66 lessons
What are the roles of RecursiveCharacterTextSplitter?
What you need to know
The algorithm, step by step
- Try the biggest separator — split the text on
"\n\n"(paragraphs). - Recurse on oversized pieces — any piece still longer than
chunk_sizeis split again on"\n", then" ", then""(single characters) as a last resort. - Merge — walk the pieces in order and join neighbours until adding the next one would pass
chunk_size. - Overlap — when starting a new chunk, carry over trailing pieces from the previous chunk, as long as they fit within
chunk_overlap. - Copy metadata — every chunk gets the parent's metadata, plus
start_indexif you ask for it.
Watch it run
1from langchain_text_splitters import RecursiveCharacterTextSplitter23text = ("Notice period\n\n"4 "Bands 1 to 3 serve 60 days. Band 4 and above serve 90 days.\n\n"5 "Buy-out\n\n"6 "The notice period may be bought out by paying basic salary "7 "for the unserved days, with manager approval.")8s = RecursiveCharacterTextSplitter(chunk_size=80, chunk_overlap=20, add_start_index=True)9for d in s.create_documents([text]):10 print(d.metadata["start_index"], repr(d.page_content))Output:
0 'Notice period\n\nBands 1 to 3 serve 60 days. Band 4 and above serve 90 days.'76 'Buy-out'85 'The notice period may be bought out by paying basic salary for the unserved'144 'for the unserved days, with manager approval.'This small run teaches three things interviewers like to probe:
- Headings can be orphaned. "Buy-out" became a 7-character chunk on its own, because joining it to the next paragraph would pass 80 characters. A chunk that says only "Buy-out" matches the query "buy-out" well and answers nothing.
- Overlap is made of whole pieces. The last chunk repeats "for the unserved" because that paragraph was cut into words, and word-sized pieces fit inside 20 characters. The earlier chunks have no overlap, because a whole paragraph is bigger than 20 characters. So
chunk_overlapis a maximum, not a guarantee. chunk_sizeis a maximum, not a target. Chunks come out at 7, 74, 75 and 45 characters.
Characters or tokens?
chunk_size=1000 means 1,000 characters by default. With the cl100k_base tokenizer, plain English runs at about 4 to 5 characters per token, so that is roughly 200 to 250 tokens. Code runs closer to 3.5 characters per token, and Hindi in Devanagari script close to 1, so the same 1,000 characters of Hindi can be nearly 1,000 tokens. Embedding models have token limits, so measure in tokens:
splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder( encoding_name="cl100k_base", chunk_size=400, chunk_overlap=50)The tokenizer here is only an approximation if your embedding model uses a different one, but it is much closer than counting characters.
A real-life example
A bank's product-FAQ bot splits product terms with chunk_size=1000 characters. One chunk ends in the middle of the "Charges" section, and the next begins with "…waived for customers maintaining Rs 25,000". A customer asks "Is the debit card fee waived?" The retriever returns the second chunk, which does not name the card, and the bot says "Yes", which is true for a different card.
The team makes two changes. They split by headings first, so each product's "Charges" section is one unit, and use the recursive splitter only inside sections longer than 400 tokens. They also prepend the product name and section heading to each chunk's text. The wrong "Yes" disappears, because every chunk now says which card it is about.
Follow-up questions to expect
- "Can you change the separators?" — Yes. For example, add
". "before" "so it prefers to break at sentence ends, or useRecursiveCharacterTextSplitter.from_language(...)for code. - "Does it guarantee chunks under
chunk_size?" — Almost. With the empty-string separator as a last resort it can always cut, but a custom separator list without""can leave oversized pieces. - "What does
keep_separatordo?" — It keeps the separator text attached to a piece rather than dropping it, which matters for separators that carry meaning, such as Markdown headings.