Course Content
RAG Systems
12 sections · 66 lessons
How does chunk size impact retrieval accuracy?
What you need to know
Chunk size changes three things at once:
- Focus. Smaller chunks give sharper vectors: the query matches the passage, not its neighbours.
- Completeness. Larger chunks are more likely to contain everything needed to answer.
- Prompt cost. Top-5 of 800-token chunks is 4,000 tokens per question; top-5 of 200-token chunks is 1,000.
There is no single best size, because it depends on how your documents are written and how your users ask.
Measure it: a small sweep
This code splits a handbook at several sizes and measures recall@1 (did the top chunk contain the answer?) on labelled questions:
1import numpy as np2from sentence_transformers import SentenceTransformer3from langchain_text_splitters import RecursiveCharacterTextSplitter45model = SentenceTransformer("BAAI/bge-small-en-v1.5")6QP = "Represent this sentence for searching relevant passages: "7text = open("handbook.txt").read()8evalset = [("How many days of paternity leave do I get?", "10 working days of paternity"),9 ("What notice period does a band 4 employee serve?", "band 4 and above must serve 90"),10 ] # ... 10 pairs in the real run1112def recall_at_k(size: int, k: int = 1) -> float:13 chunks = RecursiveCharacterTextSplitter.from_tiktoken_encoder(14 encoding_name="cl100k_base", chunk_size=size, chunk_overlap=size // 815 ).split_text(text)16 C = model.encode(chunks, normalize_embeddings=True)17 Q = model.encode([QP + q for q, _ in evalset], normalize_embeddings=True)18 top = np.argsort(-(Q @ C.T), axis=1)[:, :k]19 return np.mean([any(g in chunks[i] for i in row) for row, (_, g) in zip(top, evalset)])2021for size in [32, 64, 128, 256, 512]:22 print(size, recall_at_k(size))On a toy 650-token handbook with 10 questions, where each policy section is about 40 to 70 tokens, the run gave:
| Chunk size (tokens) | Chunks | Recall@1 |
|---|---|---|
| 32 | 45 | 0.8 |
| 64 | 16 | 1.0 |
| 128 | 7 | 0.8 |
| 256 | 3 | 0.9 |
| 512 | 2 | 0.8 |
The corpus is far too small to prove anything general. It does show the usual shape: the best score came where one chunk matched one policy section. The 256 and 512 rows look fine only because each chunk holds a third or half of the whole handbook, so the "hit" costs 5 to 10 times more prompt tokens. On a real corpus, run the same sweep with 50 to 100 questions and report recall at the k you will actually use.
Starting points
| Content | Start at |
|---|---|
| FAQs, short policies | One entry or section, 100–400 tokens |
| Technical docs, manuals | 300–800 tokens, split by heading |
| Contracts | One clause, 150–600 tokens |
| Transcripts, long prose | 400–1,000 tokens with overlap |
A real-life example
An e-commerce product Q&A team indexes seller descriptions, spec sheets and reviews with 1,000-token chunks. Questions like "What is the battery capacity?" often return a chunk that covers the whole spec sheet, and the model sometimes picks the wrong number (it confuses "5000 mAh battery" with "5000 Pa suction" on a robot vacuum listing that shares a chunk).
They sweep 128, 256, 512 and 1,000 tokens on 80 labelled questions. Recall@5 is similar from 256 to 1,000, but answer accuracy is highest at 256, because each retrieved chunk holds fewer competing numbers. They also stop splitting spec tables by size and keep each spec table whole, since tables are short.
Follow-up questions to expect
- "Would you use different sizes for different document types?" — Yes. Chunking is per source: FAQs by entry, code by function, contracts by clause.
- "Does the embedding model affect the best size?" — Yes. Models trained on short passages usually do better with shorter chunks; long-input models tolerate larger ones.
- "Can you have both small and large?" — Yes: search small chunks, then send the larger parent section to the model. That is the small-to-big pattern in the next lesson.