Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
You're indexing legal contracts where a single clause spans 3 pages. Fixed-size chunking breaks meaning. Sentence-level chunking loses context. How do you chunk documents so retrieval returns semantically complete units?
What you need to know
Fixed 500-token windows cut a limitation-of-liability clause in the middle of its exceptions. Sentence chunks give you "Such party shall not be liable under this Section" with no idea which party or which section. Both lose the thing that makes a clause meaningful: its boundaries and its context.
The pipeline
- Parse layout, not plain text — a layout-aware parser such as Docling, Unstructured or Azure Document Intelligence recovers numbered headings:
12. Indemnity,12.3 Limitation of Liability. - Split on clause boundaries — using the recovered hierarchy.
- Handle extremes — a three-page clause is split into sub-chunks that share its
clause_id; a two-line clause merges with its siblings under the parent heading. - Prepend context — each chunk starts with the document title and heading path, so "such party" has a subject.
- Search small, read big — retrieve on sub-chunks, then expand to the full clause before sending it to the model.
Context on every chunk
Master Services Agreement — Acme Ltd / Globex Pvt Ltd12. Indemnity > 12.3 Limitation of Liability (part 2 of 4)...except that the cap shall not apply to breaches of Clause 9 (Confidentiality)...The embedding now "knows" this is about liability caps in an MSA between named parties. A further step, often called contextual retrieval, asks a small model to write one sentence describing where the chunk sits in the document and prepends that too. It helps most with text full of references like "such party" and "the foregoing".
Search small, read big
hits = index.search(query, k=20) # sub-chunks: precise matchesclause_ids = list(dict.fromkeys(h.payload["clause_id"] for h in hits))[:4]context = [clause_store.full_text(cid) for cid in clause_ids] # whole clauses for the modelThis is the parent-document retriever pattern, with the clause as the parent. Small chunks embed sharply; whole clauses give the model the exceptions and cross-references it needs to answer correctly.
Trade-offs
| Strategy | Retrieval precision | Context completeness | Risk |
|---|---|---|---|
| Fixed-size windows | Medium | Poor: clauses cut mid-way | Answers miss exceptions |
| Sentence-level | High for single facts | Very poor | No subject, no section |
| Clause-based with parent expansion | High | Complete clause | Depends on parsing quality |
Failure modes
- Scanned contracts where OCR loses numbering, so the hierarchy collapses into one long block. Detect it (few headings found in a long document) and flag it for review.
- Cross-references such as "subject to Clause 9". Store references as metadata and optionally pull the referenced clause in too.
A real-life example
Scenario (illustrative numbers). A corporate legal team indexes 3,000 vendor contracts with 800-token fixed chunks. On 120 questions labelled by in-house lawyers, the governing clause is retrieved in full only 52% of the time; answers often quote a liability cap while missing the exception on the next page.
They switch to clause-based chunking with Docling, heading-path prefixes and parent expansion. Clause-level recall@4 rises to 88%. The remaining misses are mostly 140 scanned contracts where numbering was lost; those are re-processed with a better OCR step and manual heading checks.
Follow-up questions to expect
- "What if a clause is longer than the model's context budget?" — Send the matching sub-chunks plus the clause's heading and a short summary of the rest, and mark it as partial.
- "Why not semantic chunking by embedding similarity?" — It helps for unstructured prose, but contracts already have explicit structure, which is more reliable than inferred topic shifts.
- "How do you evaluate chunking?" — Clause-level recall on expert-labelled questions: did the full governing clause reach the model?