Course Content
Advanced RAG
3 sections · 38 lessons
What is sentence window retrieval, and how does it differ from parent-child chunking?
What you need to know
Why "small-to-big" at all
There is a tension in chunk size. Small units retrieve precisely: a one-sentence vector is about one thing. Large units answer well: the model needs the conditions, exceptions and steps around the fact. Small-to-big separates the two jobs — the retrieval unit is small, the generation unit is larger.
Sentence window retrieval
Index each sentence. Store the surrounding sentences in its metadata. After retrieval, swap the sentence for its window.
In LlamaIndex this is SentenceWindowNodeParser.from_defaults(window_size=3) at ingestion, plus MetadataReplacementPostProcessor(target_metadata_key="window") at query time: three sentences either side of the hit.
Parent-child retrieval
Split each document into parents (sections or pages) and split each parent into small children (for example 200 tokens). Index the children; keep parents in a document store. On a hit, return the parent. Several child hits in the same parent collapse into one parent.
In LangChain this is the ParentDocumentRetriever — a vector store for children plus a docstore for parents. Since LangChain 1.0 it lives in the langchain-classic package.
The difference, in code
1sections = {2 "rotate-certs": [3 "Certificates for Kafka brokers expire every 90 days.",4 "Rotation is done by the platform on-call engineer.",5 "Before rotating, pause the consumer autoscaler.",6 "Run the rotate-certs job for one broker at a time.",7 "Do not rotate brokers in different zones at once.",8 "Resume the autoscaler after all brokers report healthy.",9 ],10}1112def sentence_window(sec, hit, k=1):13 s = sections[sec]14 return " ".join(s[max(0, hit - k): hit + k + 1])1516def parent(sec, hit):17 return " ".join(sections[sec])1819hit = 3 # the retriever matched "Run the rotate-certs job..."20print("WINDOW:", sentence_window("rotate-certs", hit))21print("PARENT words:", len(parent("rotate-certs", hit).split()))The window returns three sentences: pause, run, don't mix zones. It cuts off the last step — "resume the autoscaler". The parent returns all six sentences (49 words here), including that step. For a procedure, the parent is right. For a long page of prose where the parent would be 4,000 tokens, the window is right.
| Sentence window | Parent-child | |
|---|---|---|
| Returned unit | Hit ± k sentences | A section or page |
| Size | Fixed, easy to budget | Varies; can be large |
| Follows structure | No | Yes |
| Duplicates | Nearby hits give overlapping windows | Hits in one parent merge |
| Best for | Flat prose, single facts | Manuals, runbooks, policies |
Practical tips
- Dedupe before building the prompt: merge overlapping windows; return each parent once.
- Cap parent size. If a "section" is 5,000 tokens, split it into smaller parents.
- Keep the matched sentence or child highlighted inside the returned text, so the model knows where to look.
A real-life example
A software company's engineering-wiki assistant indexes 30,000 pages. Two types of content behave differently:
- Runbooks are short sections of numbered steps. With sentence windows (k = 2), the assistant sometimes gave four of six steps. An engineer followed them during an incident and left the autoscaler paused. The team switched runbooks to parent-child retrieval, with the "Procedure" section as parent. Now the answer always contains the complete procedure.
- Design docs are long, flowing prose with sections of 3,000–6,000 tokens. Returning whole sections blew the context budget and buried the answer. For these, sentence windows (k = 3) work better.
They store a doc_type field at ingestion and use the retriever that suits each type. Evaluation on 300 questions checks not only "did we find the right page" but "did the returned text contain every step" for procedural questions.
Follow-up questions to expect
- "How do you choose k for the window?" — Try 1–5 on an evaluation set and watch both answer accuracy and tokens. Too small cuts conditions; too large adds noise.
- "What if the parent is huge?" — Use a middle layer (child → sub-section parent), or compress the parent with extractive filtering before sending it.
- "Is this the same as late chunking?" — No. Late chunking changes how a chunk is embedded so it carries document context. Small-to-big changes what is returned to the model.