RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

Why must long documents be split before embedding?


Similarity gap between the right and a wrong chunk0.6320.5270.1040.6420.5790.0620.6320.6190.0130.6540.6170.037right chunkwrong chunkgap25 tokens62 tokens111 tokens413 tokensbge-small-en-v1.5, query about the liability cap, one toy run.
Mixing in more clauses barely moves the right chunk's score but lets wrong chunks catch up.

What you need to know

Truncation happens without an error

This runs all-MiniLM-L6-v2, whose limit is 256 tokens. The document is 349 tokens: contract boilerplate first, and the liability clause at the end.

Python
from sentence_transformers import SentenceTransformerm = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")print(m.max_seq_length)                               # 256filler = ("This master services agreement is entered into between the customer "          "and the vendor. The vendor shall provide the services described in "          "each statement of work. ") * 12answer = ("The vendor's total liability under this agreement shall not exceed "          "the fees paid in the twelve months before the claim.")q = "What is the liability cap in the contract?"e = m.encode([q, filler + answer, filler, answer], normalize_embeddings=True)print(e[0] @ e[1], e[0] @ e[2], e[0] @ e[3], e[1] @ e[2])

The results from that run:

Text
query vs full document   0.241query vs filler only     0.241query vs answer only     0.578full document vs filler  1.000

The full document's vector is identical to the filler's vector. The liability clause was cut off, and the only sign was a tokenizer warning in the logs. The document will never be found for liability questions.

Dilution happens even inside the limit

Now take bge-small-en-v1.5 (limit 512) and build chunks of growing size: one with the liability clause plus other contract clauses, one with only other clauses. In one small run, the similarity gap between the right chunk and a wrong one was:

Chunk sizeRight chunkWrong chunkGap
~25 tokens (clause only)0.6320.5270.104
~62 tokens0.6420.5790.062
~111 tokens0.6320.6190.013
~413 tokens0.6540.6170.037

These numbers come from one toy example and will differ on your data, but the direction is typical: as a chunk mixes in more topics, wrong chunks start to look almost as relevant as right ones. With thousands of chunks, a small gap means the right one can drop out of the top k.

How to check your own pipeline

Python
tok = m.tokenizertoo_long = [c for c in chunks if len(tok(c)["input_ids"]) > m.max_seq_length]print(f"{len(too_long)} of {len(chunks)} chunks will be truncated")

Use the embedding model's own tokenizer. tiktoken counts for OpenAI tokenizers, and other models split text differently.

A 2024 alternative: late chunking

Late chunking (proposed by Jina AI) runs a long-input embedding model over the whole document once, then averages the token vectors inside each chunk's span. Each chunk gets its own vector, but that vector was computed with the whole document in view, so "it" and "this agreement" carry their meaning. It still produces one vector per chunk; it does not remove the need to split.

A real-life example

An HR policy assistant indexes a 90-page employee handbook exported as one Markdown file per chapter. The embedding model's limit is 512 tokens, and chapters run 3,000 to 8,000 tokens. Nobody splits them, because "the vector store accepted them".

Questions about anything past the first page of a chapter fail. "What is the meal allowance abroad?" finds nothing, because the expenses chapter's vector only describes travel class, the first topic in it. A token count shows 14 of 16 chapters were truncated. After splitting by section into chunks of at most 400 tokens, the same question retrieves the meal-allowance section first.

Follow-up questions to expect

  • "Does every provider truncate silently?" — Behaviour varies: some libraries truncate with only a warning, and some APIs return an error for over-long input. Never rely on either; check lengths yourself.
  • "Could you average several chunk vectors into one document vector?" — You can, for document-level routing, but it has the same dilution problem. Keep chunk vectors for answering.
  • "Is a 512-token limit a problem?" — Only if your natural sections are longer. Then split them, or pick a model with a longer limit.