Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Chunking documents so answers survive


An employee asked PolicyPal, "What is the daily meal allowance in Mumbai for grade L4?" The travel policy answers it in a clear table. PolicyPal said the policy did not cover it. When the team looked at the index, they found why: the naive chunker had cut the table in half. One chunk held the column headers "Grade, Metro, Tier 2, Tier 3". The next chunk held rows of numbers like "L4 1,800 1,400 1,100" with no headers at all. Neither chunk, on its own, meant anything.

A second miss was worse. The leave policy says: "Unused earned leave above 30 days lapses on 31 March, unless an extension is approved by the CHRO in writing." The chunk boundary fell after "lapses on 31 March". PolicyPal answered, confidently and with a citation, that leave above 30 days always lapses.

Chunking looks like a boring pre-processing step. It is actually where many RAG answers are won or lost, because the chunk is the unit the model sees. If a rule or a table is cut in half, the answer is cut in half too.

Recall@5 on 60 questions, by chunking strategy0.620.660.710.760.720.69012345fixed 500plus overlapby headingplustitle pathmax 250max 800Past 512 tokens the embedding model silently ignores the rest of a chunk.
Following the document's own headings and prefixing each chunk with its title path beat every fixed window.

How chunking destroys answers

Nine of the 23 retrieval misses in the last lesson, and several wrong answers that did retrieve something, came from a small set of chunking failures.

  • A split rule. The condition lands in one chunk and the exception in the next, as with the leave lapse rule.
  • A headless table. Rows are separated from their column names, so "1,800" is just a number.
  • A chunk with no context. "This limit applies to all grades except L1" is useless if the chunk does not say which limit or which policy.
  • Mixed topics. A 500-token window that covers the end of "Travel booking" and the start of "Hotel stays" gets an embedding that is a blur of both, and matches neither well.
  • Noise. Page footers like "Confidential. Page 7 of 22" repeated in every chunk add tokens and mean nothing.

Every one of these comes from ignoring the document's structure. The fix is to chunk along that structure.

Structure first: split on headings

Policy documents are written in sections, and a section is usually the right unit of meaning. Harbourline's policies follow a template with numbered headings like "4.2 Carry forward of earned leave", which makes sections easy to find with a regular expression. For documents without such structure, layout-aware PDF converters such as PyMuPDF4LLM or Docling can recover headings and tables as Markdown, and the same idea applies.

Python
# policypal/chunking.pyimport refrom policypal.ingest import tokHEADING = re.compile(r"^(\d+(?:\.\d+)*)\s+([A-Z][^\n]{2,80})$", re.M)MAX_TOKENS, MIN_TOKENS = 450, 60def n_tokens(text: str) -> int:    return len(tok(text, add_special_tokens=False)["input_ids"])def sections(text: str) -> list[tuple[str, str]]:    marks = list(HEADING.finditer(text))    out = []    for i, m in enumerate(marks):        end = marks[i + 1].start() if i + 1 < len(marks) else len(text)        out.append((f"{m.group(1)} {m.group(2).strip()}", text[m.end():end].strip()))    return outdef split_long(body: str) -> list[str]:    parts, current = [], ""    for para in body.split("\n\n"):        if current and n_tokens(current + "\n\n" + para) > MAX_TOKENS:            parts.append(current)            current = current.split("\n\n")[-1]        # one paragraph of overlap        current = f"{current}\n\n{para}" if current else para    return parts + [current] if current else partsdef chunk_policy(title: str, text: str) -> list[dict]:    chunks, carry = [], ""    for heading, body in sections(text):        body = f"{carry}\n\n{body}".strip() if carry else body        if n_tokens(body) < MIN_TOKENS:                 # too small: merge into the next            carry = f"{heading}\n{body}"            continue        carry = ""        for part in split_long(body):            chunks.append({"section": heading, "text": f"{title} > {heading}\n{part}"})    if carry:                                            # a small last section        chunks.append({"section": carry.split("\n", 1)[0], "text": f"{title} > {carry}"})    return chunks

The function makes three decisions. It splits at numbered headings, so a chunk never mixes two sections. It merges tiny sections into the next one, so a two-line section does not become a chunk too small to match anything. And it splits long sections at paragraph boundaries, never mid-sentence, with one paragraph of overlap so a rule that spans two paragraphs appears whole in at least one chunk.

The most valuable line is the last one. Every chunk starts with its title path, for example "Leave Policy (India) > 4.2 Carry forward of earned leave". That line gives every chunk its context, so "This limit applies to all grades" now says which policy and section it belongs to. It also helps the embedding: the words "leave" and "carry forward" are now in the vector even if the paragraph only says "unused days".

The chunker runs on the whole document text, not page by page, so a section that crosses a page break stays together. Before chunking, the ingest step removes repeated headers and footers and rejoins words hyphenated across line ends.

Tables deserve their own rule

Text extraction flattens tables into lines of values, which is exactly how the meal allowance answer was lost. PolicyPal handles tables in one of two ways, depending on size.

  • Small tables (under the token limit) stay whole in one chunk, headers included.
  • Large tables are converted row by row into short sentences at ingest time: "Travel Policy (India) > 6.1 Meal allowance: Grade L4, Metro city (includes Mumbai): ₹1,800 per day."

Row sentences are slightly more tokens, but each one is a complete, searchable fact. The conversion is ordinary code over the extracted table cells, and it is worth writing for the handful of tables that carry most numeric questions: allowances, leave entitlements by grade and device eligibility.

Size and overlap, by measurement

There is no universally right chunk size. Small chunks are precise but lose context; large chunks keep context but blur the embedding and cost more tokens per retrieved result. There is also a hard ceiling: bge-small-en-v1.5 reads at most 512 tokens, and anything after that is silently cut off before embedding. A 900-token chunk is embedded as if only its first 512 tokens existed.

So PolicyPal measured, using the same 60-question recall@5 check.

ChunkingChunksRecall@5
Fixed 500 tokens, per page, no overlap (baseline)8,9000.62
Fixed 500 tokens, 80 tokens of overlap10,4000.66
Heading-aware, max 450 tokens6,0500.71
Heading-aware, max 450, with title path6,2000.76
Heading-aware, max 250, with title path9,8000.72
Heading-aware, max 800, with title path4,1000.69

Overlap alone helped a little. Respecting structure helped more, and the title path helped most for its cost. Going smaller split sections again, and going larger ran into the 512-token cut-off. The shipped setting is heading-aware, a maximum of 450 tokens and a title path, which gives about 6,200 chunks averaging about 420 tokens. That average is the 420 you used in Section 1's cost budget.

Keep versions and dates in the metadata

Three of the original misses retrieved a superseded policy. Chunking cannot fix that, but ingest can. PolicyPal indexes only the current version of each policy, and stores version and effective_from on every chunk. When a new policy is announced before it takes effect, both versions are indexed with their dates, and the template shows the dates to the model, so it can say "from 1 April, the new rule is ...".

Check your understanding

0 of 3 answered

1.A team uses 1,000-token chunks with bge-small-en-v1.5 because "bigger chunks keep more context". What actually happens?

2.Why does adding the title path ("Leave Policy (India) > 4.2 Carry forward") to each chunk improve recall so much?

3.A travel policy has a 40-row allowance table by grade and city. What is PolicyPal's approach?