Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

What is late chunking, and why does it preserve more context?


Split first, or embed firstNormal chunking• Split the report into chunks• Embed each chunk alone• 'This arm' has no study name• Query naming ZT-302 misses itLate chunking• Encode the whole section once• Every token attends to page 1• Average tokens per chunk boundary• Chunk vector now carries ZT-302
The chunk text never changes; only its vector learns what the rest of the document said.

What you need to know

Why isolated chunks lose meaning

Documents are full of references back to earlier text: "the company", "this arm", "the above limit", "in that segment". When you split a document and embed each chunk alone, the chunk's vector reflects only what is literally in it. A question that names the study or the drug will not match a chunk that only says "this arm".

How late chunking works

  1. Encode the long text — pass the full document (up to the model's context limit, often 8,000 tokens) through the embedding model's transformer once.
  2. Keep token embeddings — each token's vector has attended to every other token, so "this arm" already carries information about the study named on page 1.
  3. Mark chunk boundaries — decide boundaries as usual (sentences, paragraphs, a token count).
  4. Pool per chunk — average the token embeddings inside each boundary to get one vector per chunk.

"Split late, embed early." The chunks and their text are unchanged; only their vectors improve.

Requirements: an embedding model with a long context window that gives access to token-level outputs, and uses mean pooling. Jina's embedding models (v3 onwards) support it directly as an option. Documents longer than the model's window must be processed in overlapping macro-sections.

Contextual retrieval: the LLM alternative

Anthropic's contextual retrieval (2024) solves the same problem with text. For each chunk, an LLM reads the whole document and writes one or two sentences that situate the chunk ("This chunk is from study ZT-302's safety section and describes the 20 mg arm."). That sentence is prepended to the chunk before both embedding and BM25 indexing. Prompt caching of the document keeps the cost manageable, since the same document is sent once per chunk.

Late chunking

  • No LLM calls; one long forward pass per document
  • Helps dense retrieval only
  • Needs a specific kind of embedding model
  • Chunk text unchanged

Contextual retrieval

  • One LLM call per chunk (cached document)
  • Helps dense and BM25; the generator sees the context too
  • Works with any embedding model
  • Chunk text gets a situating preamble

They are not competitors. A reasonable order: try late chunking first because it is cheap; if BM25 or the generator still struggles with context-free chunks, add contextual headers.

A real-life example

A pharma company indexes clinical study reports, each 200–400 pages. A typical chunk from the safety section reads:

Text
In the 20 mg arm, discontinuations due to adverse events were higher thanplacebo, driven mainly by nausea in the first two weeks.

The study ID, drug name and population appear only on earlier pages. When a safety scientist searches "ZT-302 nausea discontinuation 20 mg", this chunk ranks poorly; chunks from other studies that happen to mention "ZT-302" in their references rank above it.

The team tests two fixes on 200 questions written by the safety group:

  • Late chunking with a long-context embedding model, processing each report in overlapping sections of about 8,000 tokens. The safety chunk's vector now reflects "ZT-302" and the drug name, and it moves into the top 5.
  • Contextual headers from a small LLM with the report cached. This also helps the keyword leg of hybrid search, because "ZT-302" is now literally in the chunk text.

Late chunking alone fixes most of the dense-retrieval misses at almost no extra cost. They add contextual headers only for the safety sections, where keyword search on study IDs matters most.

Follow-up questions to expect

  • "What if the document is longer than the embedding model's window?" — Split it into long macro-sections that fit, with some overlap, and late-chunk inside each. Context from outside the section is lost, so choose sections along natural boundaries.
  • "Does late chunking help BM25?" — No. BM25 sees only the chunk text, which is unchanged. That is where contextual headers help.
  • "Do you need to re-embed everything to adopt it?" — Yes. It changes how vectors are computed, so it is a full re-embedding, planned like an embedding-model migration.