Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

How does hierarchical retrieval (summary → chunk → sentence) work?


Narrowing 2,000 policies to one table rowSummaries: 2,000 docs, keep top 8Chunks: filter doc_id to those 8, top 20Rerank chunks, keep top 5Sentences: pinpoint row, add window
Each level searches only what the level above allowed, so the first level must be wide — its misses can never be recovered.

What you need to know

The problem with one flat index

A flat index puts millions of 400-token chunks in one pool. That fails in two ways on large collections of long documents:

  • Document-level questions have no matching chunk. No single chunk is "about" a 90-page policy's overall approach.
  • Look-alike chunks from the wrong documents crowd out the right one. In a bank, fifty policies each have a "Record keeping" section that reads almost the same.

Searching summaries first narrows the pool to the right documents, so the second search competes only among relevant chunks.

The levels

  1. Summary level — one vector per document (or per section) from an LLM-written summary. Search it for the top 5–10 documents.
  2. Chunk level — search chunks with a filter such as doc_id IN (...) from step 1. Most vector stores support this filtered search natively.
  3. Sentence level (optional) — score sentences inside the best chunks to find the exact line, then expand to a small window around it for context.
Text
summaries  (2,000 vectors)      → top 8 docschunks     (filter doc_id in 8) → top 20 chunks → rerank → top 5sentences  (inside top 5)       → exact line + 2 sentences either side

Related designs

  • RAPTOR (2024) builds the hierarchy bottom-up: it clusters chunks, summarises each cluster, then clusters and summarises the summaries, forming a tree. Queries can match any level, so both detailed and "big picture" questions find a good unit.
  • Contextual chunk headers are a cheap one-level alternative: prepend the document title and a one-line summary to every chunk before embedding. That pushes document identity into each chunk without a second index.

Costs

  • An LLM summary per document or section at ingestion, and regeneration when it changes.
  • Two or three indexes to keep in sync.
  • Sequential hops: stage 2 cannot start until stage 1 finishes, so latencies add up (often 50–150 ms per hop).
  • Error cascade: if stage 1 misses the right document, stage 2 can never find it. Keep stage 1 generous (top 8–10, not top 2), or run a flat search in parallel and fuse.

For a few thousand short FAQ articles, one chunk index plus a reranker is simpler and just as good.

A real-life example

A bank's compliance assistant covers 2,000 internal policies and master circulars, many over 100 pages. Two officers ask:

  • "What is our overall approach to retaining customer records?"
  • "How long must we keep CCTV footage from branch ATMs?"

With one flat chunk index, the first question returns five "Record retention" sections from five unrelated policies — none of them the Records Management Policy's overview. The second returns chunks about CCTV installation standards.

With hierarchical retrieval:

  1. The summary index puts the Records Management Policy and the Branch Security Policy in the top 3 for the respective questions.
  2. Filtered chunk search inside those documents finds the overview section for the first question and the retention table for the second.
  3. For the second, sentence-level search pinpoints the row for ATM footage, and the window adds the row above it, which states the exception for footage under investigation.

On 250 labelled questions, the right document appears in the stage-1 top 8 for nearly all of them, and final recall@5 improves most on the broad, document-level questions.

Follow-up questions to expect

  • "What if stage 1 picks the wrong documents?" — The whole query fails. Keep stage 1 wide, monitor its recall separately, and fuse with a flat search as a safety net.
  • "How is this different from parent-child retrieval?" — Parent-child matches small units and returns bigger ones. Hierarchical retrieval searches at several levels, using the higher level to filter the lower.
  • "Summaries per document or per section?" — Per section for long, varied documents; one document summary averages too much when a policy covers ten topics.