RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is the output of the split_documents() method?


What you need to know

The three splitter methods

MethodInputOutputMetadata
split_text(text)strlist[str]None
create_documents(texts, metadatas)list[str] plus optional list[dict]list[Document]From metadatas
split_documents(docs)list[Document]list[Document]Copied from each parent

split_documents is a thin wrapper: it pulls the text and metadata out of each Document and calls create_documents.

What the output looks like

Python
from langchain_core.documents import Documentfrom langchain_text_splitters import RecursiveCharacterTextSplitterdocs = [Document(page_content=open("handbook.txt").read(),                 metadata={"source": "hr/handbook-v7.md", "acl": ["all-staff"]})]splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50,                                          add_start_index=True)chunks = splitter.split_documents(docs)print(type(chunks[0]).__name__, len(chunks))   # Document 7print(chunks[0].metadata)# {'source': 'hr/handbook-v7.md', 'acl': ['all-staff'], 'start_index': 0}lengths = [len(c.page_content) for c in chunks]print(len(chunks), min(lengths), sum(lengths) // len(lengths), max(lengths))# 7 330 400 491

Three things to notice:

  • Metadata is deep-copied. Changing one chunk's acl list does not change the parent or other chunks. That matters when you later add per-chunk fields such as chunk_id or a generated context line.
  • Chunks are shorter than chunk_size. Here the average is 400 characters against a limit of 500, because the splitter will not pass the limit to add the next paragraph.
  • The input is unchanged. You get new objects, so you can keep the parents for small-to-big retrieval.

The sanity check to run every time

The four numbers (count, min, average, max) catch most splitting bugs:

  • Max above chunk_size — a separator never matched, often in tables or minified text.
  • Many very small chunks — orphaned headings or list bullets; consider merging or structure-aware splitting.
  • Count far higher than expected — overlap too large, or a separator that fires constantly.

A real-life example

An intern builds the HR policy assistant index with this code:

Python
texts = splitter.split_text(doc.page_content)store.add_texts(texts)

It works in the demo. In production, every answer's citation reads "source: unknown", the "only current policies" filter returns nothing, and a contractor gets answers from a manager-only document, because the acl field never reached the index.

The fix is one method: split_documents([doc]) and store.add_documents(chunks, ids=...). The lesson the team wrote into their review checklist: any chunk without source and acl in its metadata fails ingestion.

Follow-up questions to expect

  • "How would you add a chunk number to each chunk?" — After splitting, loop over the chunks of each source and set metadata["chunk_no"]; it is safe because metadata was deep-copied.
  • "Is there a streaming version?" — Splitters work on lists, but you can call split_documents per batch of documents from lazy_load() to keep memory flat.
  • "What does transform_documents do?" — It is the same operation under the document-transformer interface, so a splitter can be used wherever a transformer is expected.