Course Content
RAG Systems
12 sections · 66 lessons
How does the add_start_index=True parameter help?
What you need to know
How the offset is computed
When add_start_index=True, the splitter searches the parent text for each chunk, starting just before where the previous chunk ended (to allow for overlap), and records the position it finds. The offset is in characters and is relative to the Document it split. For PyPDFLoader output, that is one page, so a real citation needs both page and start_index.
Using it
1from langchain_core.documents import Document2from langchain_text_splitters import RecursiveCharacterTextSplitter34text = open("handbook.txt").read()5parent = Document(page_content=text, metadata={"source": "hr/handbook-v7.md"})6chunks = RecursiveCharacterTextSplitter(7 chunk_size=300, chunk_overlap=40, add_start_index=True8).split_documents([parent])910hit = next(c for c in chunks if "paternity" in c.page_content)11start = hit.metadata["start_index"]12end = start + len(hit.page_content)13assert text[start:end] == hit.page_content # offset points at the exact span1415def expand(doc_text: str, start: int, end: int, pad: int = 400) -> str:16 return doc_text[max(0, start - pad): end + pad] # hit plus context on both sides1718print(f"cite: {hit.metadata['source']} chars {start}-{end}")In a run on the sample handbook, the paternity-leave chunk had start_index 603 and was 292 characters long. expand turned it into about 1,090 characters for the model, and the citation read hr/handbook-v7.md chars 603-895.
Four uses
- Citations — link to the exact span, and highlight it in a PDF viewer using page plus offset.
- Stable IDs —
hash(source + page + start_index)is the same on every run, so upserts replace old chunks. - Window expansion — search on small chunks, send the model a wider window.
- Debugging — sort chunks by
start_indexand compare with lengths to see real overlap and any gaps.
A caveat
If you change a chunk's text after splitting, for example by prepending a heading, the offset still points into the original text, not the modified chunk. Keep the original span length in metadata if you need to slice the source later.
A real-life example
A legal-contract search tool shows each answer next to the contract PDF. Lawyers complained that "page 14" was not enough: page 14 has four clauses, and they had to read all of them to find the one used.
The team turned on add_start_index=True, stored page and start_index per chunk, and used them to highlight the exact passage in the PDF viewer. Review time per answer dropped noticeably in their user study, and trust rose, because lawyers could check each claim in a few seconds. The same offsets let them build deterministic chunk IDs, which removed duplicates that had built up from nightly re-indexing.
Follow-up questions to expect
- "Is
start_indexstable if the document changes?" — No. An edit near the start shifts every later offset. Pair offsets with a document version or content hash. - "Why not just store the chunk text?" — You do, but the offset connects it back to the original, which is what citations and window expansion need.
- "Does it work with token-based splitting?" — Yes; the offset is still in characters of the original text.