Course Content
RAG Systems
12 sections · 66 lessons
Why is RecursiveCharacterTextSplitter recommended for generic text?
What you need to know
The alternatives, and when each wins
| Splitter | How it splits | Use it for |
|---|---|---|
CharacterTextSplitter | One separator only | Very uniform text; rarely the best choice |
RecursiveCharacterTextSplitter | Ordered separators, falling back | Generic prose; the default |
TokenTextSplitter | Fixed token windows | When exact token size matters more than boundaries |
MarkdownHeaderTextSplitter, HTMLHeaderTextSplitter | At headings, recording the heading path | Docs, wikis, help centres |
from_language(Language.PYTHON) and others | At class and function boundaries | Source code |
| Semantic chunking | Embeds sentences, cuts where meaning shifts | Long prose without headings |
| LLM-based chunking | An LLM chooses boundaries | Small, high-value corpora |
Language-aware separators
The recursive splitter can use a language's own structure. For Python, the separators are:
1from langchain_text_splitters import Language, RecursiveCharacterTextSplitter23print(RecursiveCharacterTextSplitter.get_separators_for_language(Language.PYTHON))4# ['\nclass ', '\ndef ', '\n\tdef ', '\n\n', '\n', ' ', '']56code_splitter = RecursiveCharacterTextSplitter.from_language(7 Language.PYTHON, chunk_size=1200, chunk_overlap=0)It tries to cut before class and def first, so functions usually stay whole. For Markdown, the first separators are headings of each level, then code fences, then horizontal rules.
About semantic chunking
Semantic chunking embeds each sentence and starts a new chunk where the similarity between neighbouring sentences drops. It can find good boundaries in prose with no headings. But it costs an embedding pass over the whole corpus at ingest, it behaves oddly on lists and tables, and studies published in 2024 found its gains over simple fixed-size or structure-based splitting to be inconsistent. Measure before adopting it.
The usual production shape
- Split by structure — headings, clauses, FAQ entries, functions.
- Recurse inside long sections — use the recursive splitter to cap size.
- Add context — prepend title and heading path to each chunk.
A real-life example
A developer-tools company builds an assistant over its API documentation (Markdown) and SDK source (Python). The first version runs the default recursive splitter at 1,000 characters on everything.
Two problems show up. Code chunks start in the middle of a function, so "How do I retry a failed upload?" retrieves half of upload_with_retry() without its signature. And Markdown chunks lose their headings, so a parameter table appears without the endpoint it belongs to.
They switch to MarkdownHeaderTextSplitter for docs (with the heading path prepended) and from_language(Language.PYTHON) for code. The recursive splitter still runs inside any section longer than 600 tokens. Both failure types go away, with no model change.
Follow-up questions to expect
- "Why not always use semantic chunking?" — It costs more, is harder to reproduce, and its benefit is not reliable. Structure-aware splitting gets most of the benefit for free when documents have headings.
- "Where is
SemanticChunker?" — Inlangchain_experimental, which signals that its API may change. - "What would you do for PDFs?" — Use a layout-aware parser to recover headings and tables, convert to Markdown, then split by headings.