RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

Why is RecursiveCharacterTextSplitter recommended for generic text?


What you need to know

The alternatives, and when each wins

SplitterHow it splitsUse it for
CharacterTextSplitterOne separator onlyVery uniform text; rarely the best choice
RecursiveCharacterTextSplitterOrdered separators, falling backGeneric prose; the default
TokenTextSplitterFixed token windowsWhen exact token size matters more than boundaries
MarkdownHeaderTextSplitter, HTMLHeaderTextSplitterAt headings, recording the heading pathDocs, wikis, help centres
from_language(Language.PYTHON) and othersAt class and function boundariesSource code
Semantic chunkingEmbeds sentences, cuts where meaning shiftsLong prose without headings
LLM-based chunkingAn LLM chooses boundariesSmall, high-value corpora

Language-aware separators

The recursive splitter can use a language's own structure. For Python, the separators are:

Python
from langchain_text_splitters import Language, RecursiveCharacterTextSplitterprint(RecursiveCharacterTextSplitter.get_separators_for_language(Language.PYTHON))# ['\nclass ', '\ndef ', '\n\tdef ', '\n\n', '\n', ' ', '']code_splitter = RecursiveCharacterTextSplitter.from_language(    Language.PYTHON, chunk_size=1200, chunk_overlap=0)

It tries to cut before class and def first, so functions usually stay whole. For Markdown, the first separators are headings of each level, then code fences, then horizontal rules.

About semantic chunking

Semantic chunking embeds each sentence and starts a new chunk where the similarity between neighbouring sentences drops. It can find good boundaries in prose with no headings. But it costs an embedding pass over the whole corpus at ingest, it behaves oddly on lists and tables, and studies published in 2024 found its gains over simple fixed-size or structure-based splitting to be inconsistent. Measure before adopting it.

The usual production shape

  1. Split by structure — headings, clauses, FAQ entries, functions.
  2. Recurse inside long sections — use the recursive splitter to cap size.
  3. Add context — prepend title and heading path to each chunk.

A real-life example

A developer-tools company builds an assistant over its API documentation (Markdown) and SDK source (Python). The first version runs the default recursive splitter at 1,000 characters on everything.

Two problems show up. Code chunks start in the middle of a function, so "How do I retry a failed upload?" retrieves half of upload_with_retry() without its signature. And Markdown chunks lose their headings, so a parameter table appears without the endpoint it belongs to.

They switch to MarkdownHeaderTextSplitter for docs (with the heading path prepended) and from_language(Language.PYTHON) for code. The recursive splitter still runs inside any section longer than 600 tokens. Both failure types go away, with no model change.

Follow-up questions to expect

  • "Why not always use semantic chunking?" — It costs more, is harder to reproduce, and its benefit is not reliable. Structure-aware splitting gets most of the benefit for free when documents have headings.
  • "Where is SemanticChunker?" — In langchain_experimental, which signals that its API may change.
  • "What would you do for PDFs?" — Use a layout-aware parser to recover headings and tables, convert to Markdown, then split by headings.