Course Content
Advanced RAG
3 sections · 38 lessons
What is agentic chunking, and how does it differ from rule-based splitting?
What you need to know
The spectrum of chunking methods
| Method | How boundaries are chosen | Cost | Weakness |
|---|---|---|---|
| Fixed size | Every N tokens | Free | Cuts mid-sentence |
| Recursive splitter | Paragraph → line → sentence → character, under a size limit | Free | Ignores tables, lists, sections |
| Structure-aware | Document structure: headings, list items, table blocks, code blocks | Cheap | Needs a good parser; messy files lack structure |
| Semantic | Split where embedding similarity between neighbouring sentences drops | Embedding calls | Noisy on short sentences |
| Agentic | An LLM reads the text and decides units, titles, summaries | LLM calls | Cost, nondeterminism |
An advanced interview assumes you know the first rows. The real question is when an LLM's judgement is worth paying for.
How it works in practice
- Parse the document into clean text with layout hints (headings, tables in Markdown).
- Propose units — the LLM reads a window of a few pages and returns a JSON list of chunks: start and end markers, title, summary, type.
- Handle special content — keep a table whole (or turn it into row-level facts), keep a numbered procedure with its preamble, keep a code block with the sentence that explains it.
- Validate — check that chunks cover the text with no gaps or overlaps, and that each is under the size limit; split any oversized chunk with rules.
- Cache by document hash — reuse the result until the document changes, so re-runs don't produce a different index.
Step 4 matters: LLMs sometimes skip a paragraph or repeat one. Always rebuild chunk text from the original document using the LLM's boundaries, never from text the LLM rewrote.
When it is worth it
- Heterogeneous documents where meaning and layout disagree: contracts with schedules, manuals mixing prose and tables, papers with figures.
- High-value corpora where retrieval errors are expensive.
- Documents with little reliable structure (converted PDFs, emails).
For clean HTML help pages or Markdown with good headings, a structure-aware splitter gets nearly all of the benefit for free.
Costs
- An LLM pass over every document, repeated on every change.
- Nondeterminism: two runs can produce different boundaries, which makes index diffs and debugging harder. Low temperature where available, and caching by hash, reduce this.
- It is an index-time cost, paid once per document version and shared by every future query — often the best place in a RAG pipeline to spend LLM budget.
A real-life example
A bank's compliance team indexes 9,000 corporate loan agreements. Each has clauses, sub-clauses, defined terms and schedules with covenant tables. A recursive splitter at 800 tokens:
- cuts clause 14.2 ("Financial Covenants") in half, so the ratio threshold and its "tested quarterly" condition land in different chunks;
- splits the covenant table across two chunks, separating column headers from values;
- leaves the definition of "EBITDA" 30 pages away from the clauses that use it.
The team runs agentic chunking with a mid-sized model over one agreement at a time. It emits one chunk per clause (with sub-clauses kept together), one chunk per schedule table (kept whole, with its caption), and a "defined terms" chunk per agreement that is attached as context whenever a clause using a defined term is retrieved.
On 250 questions from credit analysts ("What is the minimum interest coverage ratio for Borrower X, and how often is it tested?"), the share of answers with both the number and its condition correct rises sharply. The ingestion bill for 9,000 agreements is a one-time cost that the team compares with the analyst hours saved. For the bank's internal HTML policy pages, which have clean headings, they keep a free structure-aware splitter.
Follow-up questions to expect
- "How is this different from semantic chunking?" — Semantic chunking uses embedding similarity between neighbouring sentences to find topic shifts; it cannot recognise a table or a clause. Agentic chunking uses an LLM's reading of structure and meaning.
- "How do you evaluate a chunking strategy?" — Hold everything else fixed and compare retrieval recall and answer accuracy on the same question set. Also check a sample of chunks by eye for broken tables and orphaned steps.
- "Can you mix strategies?" — Yes, and you should: route documents by type at ingestion — structure-aware for clean HTML, agentic for contracts, row-level facts for spreadsheets.