Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Chunking documents so answers survive
An employee asked PolicyPal, "What is the daily meal allowance in Mumbai for grade L4?" The travel policy answers it in a clear table. PolicyPal said the policy did not cover it. When the team looked at the index, they found why: the naive chunker had cut the table in half. One chunk held the column headers "Grade, Metro, Tier 2, Tier 3". The next chunk held rows of numbers like "L4 1,800 1,400 1,100" with no headers at all. Neither chunk, on its own, meant anything.
A second miss was worse. The leave policy says: "Unused earned leave above 30 days lapses on 31 March, unless an extension is approved by the CHRO in writing." The chunk boundary fell after "lapses on 31 March". PolicyPal answered, confidently and with a citation, that leave above 30 days always lapses.
Chunking looks like a boring pre-processing step. It is actually where many RAG answers are won or lost, because the chunk is the unit the model sees. If a rule or a table is cut in half, the answer is cut in half too.
How chunking destroys answers
Nine of the 23 retrieval misses in the last lesson, and several wrong answers that did retrieve something, came from a small set of chunking failures.
- A split rule. The condition lands in one chunk and the exception in the next, as with the leave lapse rule.
- A headless table. Rows are separated from their column names, so "1,800" is just a number.
- A chunk with no context. "This limit applies to all grades except L1" is useless if the chunk does not say which limit or which policy.
- Mixed topics. A 500-token window that covers the end of "Travel booking" and the start of "Hotel stays" gets an embedding that is a blur of both, and matches neither well.
- Noise. Page footers like "Confidential. Page 7 of 22" repeated in every chunk add tokens and mean nothing.
Every one of these comes from ignoring the document's structure. The fix is to chunk along that structure.
Structure first: split on headings
Policy documents are written in sections, and a section is usually the right unit of meaning. Harbourline's policies follow a template with numbered headings like "4.2 Carry forward of earned leave", which makes sections easy to find with a regular expression. For documents without such structure, layout-aware PDF converters such as PyMuPDF4LLM or Docling can recover headings and tables as Markdown, and the same idea applies.
1# policypal/chunking.py2import re34from policypal.ingest import tok56HEADING = re.compile(r"^(\d+(?:\.\d+)*)\s+([A-Z][^\n]{2,80})$", re.M)7MAX_TOKENS, MIN_TOKENS = 450, 6089def n_tokens(text: str) -> int:10 return len(tok(text, add_special_tokens=False)["input_ids"])1112def sections(text: str) -> list[tuple[str, str]]:13 marks = list(HEADING.finditer(text))14 out = []15 for i, m in enumerate(marks):16 end = marks[i + 1].start() if i + 1 < len(marks) else len(text)17 out.append((f"{m.group(1)} {m.group(2).strip()}", text[m.end():end].strip()))18 return out1920def split_long(body: str) -> list[str]:21 parts, current = [], ""22 for para in body.split("\n\n"):23 if current and n_tokens(current + "\n\n" + para) > MAX_TOKENS:24 parts.append(current)25 current = current.split("\n\n")[-1] # one paragraph of overlap26 current = f"{current}\n\n{para}" if current else para27 return parts + [current] if current else parts2829def chunk_policy(title: str, text: str) -> list[dict]:30 chunks, carry = [], ""31 for heading, body in sections(text):32 body = f"{carry}\n\n{body}".strip() if carry else body33 if n_tokens(body) < MIN_TOKENS: # too small: merge into the next34 carry = f"{heading}\n{body}"35 continue36 carry = ""37 for part in split_long(body):38 chunks.append({"section": heading, "text": f"{title} > {heading}\n{part}"})39 if carry: # a small last section40 chunks.append({"section": carry.split("\n", 1)[0], "text": f"{title} > {carry}"})41 return chunksThe function makes three decisions. It splits at numbered headings, so a chunk never mixes two sections. It merges tiny sections into the next one, so a two-line section does not become a chunk too small to match anything. And it splits long sections at paragraph boundaries, never mid-sentence, with one paragraph of overlap so a rule that spans two paragraphs appears whole in at least one chunk.
The most valuable line is the last one. Every chunk starts with its title path, for example "Leave Policy (India) > 4.2 Carry forward of earned leave". That line gives every chunk its context, so "This limit applies to all grades" now says which policy and section it belongs to. It also helps the embedding: the words "leave" and "carry forward" are now in the vector even if the paragraph only says "unused days".
The chunker runs on the whole document text, not page by page, so a section that crosses a page break stays together. Before chunking, the ingest step removes repeated headers and footers and rejoins words hyphenated across line ends.
Tables deserve their own rule
Text extraction flattens tables into lines of values, which is exactly how the meal allowance answer was lost. PolicyPal handles tables in one of two ways, depending on size.
- Small tables (under the token limit) stay whole in one chunk, headers included.
- Large tables are converted row by row into short sentences at ingest time: "Travel Policy (India) > 6.1 Meal allowance: Grade L4, Metro city (includes Mumbai): ₹1,800 per day."
Row sentences are slightly more tokens, but each one is a complete, searchable fact. The conversion is ordinary code over the extracted table cells, and it is worth writing for the handful of tables that carry most numeric questions: allowances, leave entitlements by grade and device eligibility.
Size and overlap, by measurement
There is no universally right chunk size. Small chunks are precise but lose context; large chunks keep context but blur the embedding and cost more tokens per retrieved result. There is also a hard ceiling: bge-small-en-v1.5 reads at most 512 tokens, and anything after that is silently cut off before embedding. A 900-token chunk is embedded as if only its first 512 tokens existed.
So PolicyPal measured, using the same 60-question recall@5 check.
| Chunking | Chunks | Recall@5 |
|---|---|---|
| Fixed 500 tokens, per page, no overlap (baseline) | 8,900 | 0.62 |
| Fixed 500 tokens, 80 tokens of overlap | 10,400 | 0.66 |
| Heading-aware, max 450 tokens | 6,050 | 0.71 |
| Heading-aware, max 450, with title path | 6,200 | 0.76 |
| Heading-aware, max 250, with title path | 9,800 | 0.72 |
| Heading-aware, max 800, with title path | 4,100 | 0.69 |
Overlap alone helped a little. Respecting structure helped more, and the title path helped most for its cost. Going smaller split sections again, and going larger ran into the 512-token cut-off. The shipped setting is heading-aware, a maximum of 450 tokens and a title path, which gives about 6,200 chunks averaging about 420 tokens. That average is the 420 you used in Section 1's cost budget.
Keep versions and dates in the metadata
Three of the original misses retrieved a superseded policy. Chunking cannot fix that, but ingest can. PolicyPal indexes only the current version of each policy, and stores version and effective_from on every chunk. When a new policy is announced before it takes effect, both versions are indexed with their dates, and the template shows the dates to the model, so it can say "from 1 April, the new rule is ...".
Check your understanding
0 of 3 answered
1.A team uses 1,000-token chunks with bge-small-en-v1.5 because "bigger chunks keep more context". What actually happens?
2.Why does adding the title path ("Leave Policy (India) > 4.2 Carry forward") to each chunk improve recall so much?
3.A travel policy has a 40-row allowance table by grade and city. What is PolicyPal's approach?