Course Content
RAG Systems
12 sections · 66 lessons
What is the output of the split_documents() method?
What you need to know
The three splitter methods
| Method | Input | Output | Metadata |
|---|---|---|---|
split_text(text) | str | list[str] | None |
create_documents(texts, metadatas) | list[str] plus optional list[dict] | list[Document] | From metadatas |
split_documents(docs) | list[Document] | list[Document] | Copied from each parent |
split_documents is a thin wrapper: it pulls the text and metadata out of each Document and calls create_documents.
What the output looks like
Python
1from langchain_core.documents import Document2from langchain_text_splitters import RecursiveCharacterTextSplitter34docs = [Document(page_content=open("handbook.txt").read(),5 metadata={"source": "hr/handbook-v7.md", "acl": ["all-staff"]})]6splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50,7 add_start_index=True)8chunks = splitter.split_documents(docs)910print(type(chunks[0]).__name__, len(chunks)) # Document 711print(chunks[0].metadata)12# {'source': 'hr/handbook-v7.md', 'acl': ['all-staff'], 'start_index': 0}1314lengths = [len(c.page_content) for c in chunks]15print(len(chunks), min(lengths), sum(lengths) // len(lengths), max(lengths))16# 7 330 400 491Three things to notice:
- Metadata is deep-copied. Changing one chunk's
acllist does not change the parent or other chunks. That matters when you later add per-chunk fields such aschunk_idor a generated context line. - Chunks are shorter than
chunk_size. Here the average is 400 characters against a limit of 500, because the splitter will not pass the limit to add the next paragraph. - The input is unchanged. You get new objects, so you can keep the parents for small-to-big retrieval.
The sanity check to run every time
The four numbers (count, min, average, max) catch most splitting bugs:
- Max above
chunk_size— a separator never matched, often in tables or minified text. - Many very small chunks — orphaned headings or list bullets; consider merging or structure-aware splitting.
- Count far higher than expected — overlap too large, or a separator that fires constantly.
A real-life example
An intern builds the HR policy assistant index with this code:
Python
texts = splitter.split_text(doc.page_content)store.add_texts(texts)It works in the demo. In production, every answer's citation reads "source: unknown", the "only current policies" filter returns nothing, and a contractor gets answers from a manager-only document, because the acl field never reached the index.
The fix is one method: split_documents([doc]) and store.add_documents(chunks, ids=...). The lesson the team wrote into their review checklist: any chunk without source and acl in its metadata fails ingestion.
Follow-up questions to expect
- "How would you add a chunk number to each chunk?" — After splitting, loop over the chunks of each source and set
metadata["chunk_no"]; it is safe because metadata was deep-copied. - "Is there a streaming version?" — Splitters work on lists, but you can call
split_documentsper batch of documents fromlazy_load()to keep memory flat. - "What does
transform_documentsdo?" — It is the same operation under the document-transformer interface, so a splitter can be used wherever a transformer is expected.