RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What trade-offs exist between larger chunks and smaller chunks?


What you need to know

Small chunks (~100–300 tokens)Large chunks (~800–2,000 tokens)
Match precisionHigh: the vector is focusedLower: topics blur
Self-contained?Often not ("it", "this policy")Usually yes
Prompt tokens at k = 5500–1,5004,000–10,000
Number of chunksMany (bigger index)Fewer
Typical failureRight passage, missing contextRight document, wrong passage ranked higher

The small-to-big pattern

  1. Index small — split each section into small child chunks (for example 150 tokens) and embed those.
  2. Remember the parent — store the ID of the parent section (for example 800 tokens) in each child's metadata.
  3. Search small — the query matches the most precise child chunks.
  4. Return big — replace each child with its parent, remove duplicate parents, and send those to the model.

In LangChain 1.x, ParentDocumentRetriever lives in the langchain_classic package. You can also build it yourself: store parent_id in metadata and look the parent up in a key-value store after search.

A lighter version is the sentence window: embed single sentences or small chunks, and at answer time expand each hit by a few hundred characters on each side, using its start_index.

Contextual chunks as another answer

Instead of making chunks bigger, you can make each small chunk carry its context. Contextual retrieval (described by Anthropic in 2024) uses an LLM to write one or two sentences per chunk explaining where it sits in the document ("This clause is from the 2025 MSA with Vendor A, section 9, Limitation of Liability"), and prepends that before embedding and keyword indexing. Chunks stay small but stop being ambiguous. The cost is one LLM call per chunk at ingest, which prompt caching of the full document makes much cheaper.

A real-life example

A bank's product-FAQ bot has product terms split into 800-token chunks. Precision is poor: "What is the late payment fee on the Platinum card?" retrieves a chunk that covers fees for three cards, and the model sometimes quotes the Gold card's fee.

The team re-indexes with 150-token child chunks (one fee row or rule each) and 800-token parent sections, and prepends "Platinum credit card — Fees and charges" to each child. The search now lands on the exact fee row for the Platinum card; the model receives that row's parent section, so it also sees the conditions ("waived if paid within 3 days of the due date"). Prompt size stays about the same, because three parent sections replace five mixed chunks.

Follow-up questions to expect

  • "What if several children share a parent?" — Deduplicate: send the parent once. Several matching children in one parent is a strong relevance signal.
  • "Doesn't returning parents bring back the noise?" — Some, but ranking was done on the precise child, and parents are chosen to be coherent sections, not arbitrary windows.
  • "How would you choose the parent size?" — Use the document's structure (section or clause) where it exists, and keep it within your per-chunk prompt budget.