Course Content
RAG Systems
12 sections · 66 lessons
Why must long documents be split before embedding?
What you need to know
Truncation happens without an error
This runs all-MiniLM-L6-v2, whose limit is 256 tokens. The document is 349 tokens: contract boilerplate first, and the liability clause at the end.
1from sentence_transformers import SentenceTransformer23m = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")4print(m.max_seq_length) # 2565filler = ("This master services agreement is entered into between the customer "6 "and the vendor. The vendor shall provide the services described in "7 "each statement of work. ") * 128answer = ("The vendor's total liability under this agreement shall not exceed "9 "the fees paid in the twelve months before the claim.")10q = "What is the liability cap in the contract?"11e = m.encode([q, filler + answer, filler, answer], normalize_embeddings=True)12print(e[0] @ e[1], e[0] @ e[2], e[0] @ e[3], e[1] @ e[2])The results from that run:
query vs full document 0.241query vs filler only 0.241query vs answer only 0.578full document vs filler 1.000The full document's vector is identical to the filler's vector. The liability clause was cut off, and the only sign was a tokenizer warning in the logs. The document will never be found for liability questions.
Dilution happens even inside the limit
Now take bge-small-en-v1.5 (limit 512) and build chunks of growing size: one with the liability clause plus other contract clauses, one with only other clauses. In one small run, the similarity gap between the right chunk and a wrong one was:
| Chunk size | Right chunk | Wrong chunk | Gap |
|---|---|---|---|
| ~25 tokens (clause only) | 0.632 | 0.527 | 0.104 |
| ~62 tokens | 0.642 | 0.579 | 0.062 |
| ~111 tokens | 0.632 | 0.619 | 0.013 |
| ~413 tokens | 0.654 | 0.617 | 0.037 |
These numbers come from one toy example and will differ on your data, but the direction is typical: as a chunk mixes in more topics, wrong chunks start to look almost as relevant as right ones. With thousands of chunks, a small gap means the right one can drop out of the top k.
How to check your own pipeline
tok = m.tokenizertoo_long = [c for c in chunks if len(tok(c)["input_ids"]) > m.max_seq_length]print(f"{len(too_long)} of {len(chunks)} chunks will be truncated")Use the embedding model's own tokenizer. tiktoken counts for OpenAI tokenizers, and other models split text differently.
A 2024 alternative: late chunking
Late chunking (proposed by Jina AI) runs a long-input embedding model over the whole document once, then averages the token vectors inside each chunk's span. Each chunk gets its own vector, but that vector was computed with the whole document in view, so "it" and "this agreement" carry their meaning. It still produces one vector per chunk; it does not remove the need to split.
A real-life example
An HR policy assistant indexes a 90-page employee handbook exported as one Markdown file per chapter. The embedding model's limit is 512 tokens, and chapters run 3,000 to 8,000 tokens. Nobody splits them, because "the vector store accepted them".
Questions about anything past the first page of a chapter fail. "What is the meal allowance abroad?" finds nothing, because the expenses chapter's vector only describes travel class, the first topic in it. A token count shows 14 of 16 chapters were truncated. After splitting by section into chunks of at most 400 tokens, the same question retrieves the meal-allowance section first.
Follow-up questions to expect
- "Does every provider truncate silently?" — Behaviour varies: some libraries truncate with only a warning, and some APIs return an error for over-long input. Never rely on either; check lengths yourself.
- "Could you average several chunk vectors into one document vector?" — You can, for document-level routing, but it has the same dilution problem. Keep chunk vectors for answering.
- "Is a 512-token limit a problem?" — Only if your natural sections are longer. Then split them, or pick a model with a longer limit.