RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

Why is it necessary to split documents into smaller chunks?


What you need to know

Four reasons to split

  1. Model input limits. Every embedding model has a maximum input length. all-MiniLM-L6-v2 stops at 256 tokens, bge-base-en-v1.5 at 512, OpenAI's text-embedding-3 models at 8,191, and BGE-M3 at 8,192. Text past the limit is usually dropped without an error.
  2. Precision. A user asking "what is the liability cap?" wants one clause, not a 40-page PDF.
  3. Focus of the vector. An embedding squeezes a whole text into one point. A text about ten topics lands somewhere between all ten and close to none of them.
  4. Prompt budget and cost. Five 400-token chunks are 2,000 tokens. Five whole documents might be 100,000 tokens: slower, more expensive and harder for the model to use.

How big is a document, in tokens?

Tokens are the pieces a model reads; in English one token is about three quarters of a word.

DocumentRough size
One FAQ answer50–150 tokens
One policy section150–500 tokens
One dense page500–700 tokens
40-page contract20,000–28,000 tokens
300-page HR handbook150,000–200,000 tokens

So even a model with an 8,192-token input limit sees only the first dozen pages of the contract.

The target size

Aim for "the smallest self-contained passage that could answer a question". For FAQs that is one question and answer. For policies it is one numbered section. For contracts it is one clause. That is usually 150 to 600 tokens, but the document's structure should decide, not a fixed number.

A real-life example

A legal-contract search tool first stored one embedding per contract, using a long-input model so nothing was cut. Lawyers asked "Which contracts cap liability at 12 months of fees?" and the top results were contracts that simply mentioned "liability" many times, such as insurance agreements.

The team switched to clause-level chunks: they split each contract on its numbered clause headings, giving about 60 chunks of 150 to 500 tokens per contract. Now the query matched the limitation-of-liability clause itself, and each result showed the clause text and page number. The index grew from 12,000 vectors to about 700,000, which is still small for any vector database.

Follow-up questions to expect

  • "Long-context embedding models exist. Why not embed whole documents?" — You can, but one vector still averages everything, so ranking for specific questions gets worse. Long-input models are better used to embed large chunks, or for late chunking, where the whole document is read once and each chunk's vector is pooled from it.
  • "Can a chunk be too small?" — Yes. A single sentence often lacks the subject ("This applies after one year"), so it matches but cannot answer.
  • "Do you split tables?" — Keep a table in one chunk when it fits; if it does not, repeat the header row in every piece.