Course Content
Advanced RAG
3 sections · 38 lessons
What is proposition-based chunking, and when is it better than traditional chunking?
What you need to know
The two problems it fixes
- Averaging. An embedding of a 400-token chunk is one point that represents everything in it. If the chunk covers dosage, contraindications and storage, a question about storage matches it only weakly.
- Dangling references. Sentences like "It must not be used in children" or "The above limit applies" mean nothing out of context. As a chunk, or part of one, they are hard to retrieve for a question that names the drug.
Propositions fix both. Each fact becomes its own small unit, and each unit names its subject.
Source: "Zentrofen is indicated for moderate pain. It must not be used in children under 12, and its maximum daily dose is 1,200 mg."Props: "Zentrofen is indicated for moderate pain." "Zentrofen must not be used in children under 12." "The maximum daily dose of Zentrofen is 1,200 mg."How to build it
- Split documents into normal passages (sections or paragraphs).
- Decompose — prompt an LLM: "Rewrite this passage as a list of simple, standalone facts. Replace pronouns with the names they refer to." Ask for JSON output.
- Index each proposition as its own vector, with metadata pointing to its parent passage and document.
- Retrieve on propositions; return the parent passage (deduplicated) to the generator.
Step 4 is the important design choice: small-to-big. Propositions give precise matching, but a list of isolated facts loses the conditions and connections that give them meaning. The parent passage restores that.
When it is better — and when it is not
| Good fit | Poor fit |
|---|---|
| Product labels, spec sheets, policy clauses | Narrative text: case studies, reports |
| FAQs and reference data | Procedures where order matters |
| Questions that ask for one fact | Questions that need reasoning across paragraphs |
Costs: an LLM call per passage at ingestion (a real bill across millions of documents), several times more vectors, re-running on every change, and the risk that the LLM drops or changes a qualifier while rewriting. Spot-check a sample of propositions against their source.
A real-life example
A pharma company's regulatory search indexes 4,000 product labels. Pharmacovigilance staff ask narrow questions such as "maximum daily dose of Zentrofen in hepatic impairment".
With 500-token chunks, the matching sentence sits in a chunk that also covers indications, interactions and storage. Its vector is a blend, and the right chunk ranks 12th. Worse, the key sentence says "In these patients, the dose should not exceed 600 mg", with "these patients" defined two sentences earlier.
After proposition indexing, that fact becomes "In patients with hepatic impairment, the maximum daily dose of Zentrofen is 600 mg." It ranks first. The model receives the whole parent section, so it also sees the surrounding warning about monitoring liver function.
The team measures on 200 questions from the safety team: recall@5 rises clearly for single-fact questions, and stays flat for "summarise the safety profile" questions — which they route to section-level retrieval instead. Ingestion cost for the 4,000 labels is a one-off batch job run overnight on a small model.
Follow-up questions to expect
- "How is this different from indexing hypothetical questions?" — Propositions restate facts from the document; hypothetical questions predict what users will ask. Both make the index unit closer to the query; you can use both as extra vectors pointing at the same parent.
- "Does late chunking or contextual retrieval solve the same problem?" — They fix dangling references by adding context to each chunk's embedding, without rewriting text. They do not split multi-fact chunks, so propositions still win for very dense reference text.
- "How do you stop the LLM from changing facts?" — Low temperature where available, an instruction to copy numbers and units exactly, and an automated check that every number in a proposition appears in its source passage.