Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

You're indexing legal contracts where a single clause spans 3 pages. Fixed-size chunking breaks meaning. Sentence-level chunking loses context. How do you chunk documents so retrieval returns semantically complete units?


Search small, read bigMasterServices Agreement9 Confidentiality12 Indemnity9.19.212.112.3 cap: 3pages, 4 sub-chunks
A query matches one sub-chunk of 12.3, but the model receives the whole clause, including the exception on the next page.

What you need to know

Fixed 500-token windows cut a limitation-of-liability clause in the middle of its exceptions. Sentence chunks give you "Such party shall not be liable under this Section" with no idea which party or which section. Both lose the thing that makes a clause meaningful: its boundaries and its context.

The pipeline

  1. Parse layout, not plain text — a layout-aware parser such as Docling, Unstructured or Azure Document Intelligence recovers numbered headings: 12. Indemnity, 12.3 Limitation of Liability.
  2. Split on clause boundaries — using the recovered hierarchy.
  3. Handle extremes — a three-page clause is split into sub-chunks that share its clause_id; a two-line clause merges with its siblings under the parent heading.
  4. Prepend context — each chunk starts with the document title and heading path, so "such party" has a subject.
  5. Search small, read big — retrieve on sub-chunks, then expand to the full clause before sending it to the model.

Context on every chunk

Text
Master Services Agreement — Acme Ltd / Globex Pvt Ltd12. Indemnity > 12.3 Limitation of Liability (part 2 of 4)...except that the cap shall not apply to breaches of Clause 9 (Confidentiality)...

The embedding now "knows" this is about liability caps in an MSA between named parties. A further step, often called contextual retrieval, asks a small model to write one sentence describing where the chunk sits in the document and prepends that too. It helps most with text full of references like "such party" and "the foregoing".

Search small, read big

Python
hits = index.search(query, k=20)                        # sub-chunks: precise matchesclause_ids = list(dict.fromkeys(h.payload["clause_id"] for h in hits))[:4]context = [clause_store.full_text(cid) for cid in clause_ids]  # whole clauses for the model

This is the parent-document retriever pattern, with the clause as the parent. Small chunks embed sharply; whole clauses give the model the exceptions and cross-references it needs to answer correctly.

Trade-offs

StrategyRetrieval precisionContext completenessRisk
Fixed-size windowsMediumPoor: clauses cut mid-wayAnswers miss exceptions
Sentence-levelHigh for single factsVery poorNo subject, no section
Clause-based with parent expansionHighComplete clauseDepends on parsing quality

Failure modes

  • Scanned contracts where OCR loses numbering, so the hierarchy collapses into one long block. Detect it (few headings found in a long document) and flag it for review.
  • Cross-references such as "subject to Clause 9". Store references as metadata and optionally pull the referenced clause in too.

A real-life example

Scenario (illustrative numbers). A corporate legal team indexes 3,000 vendor contracts with 800-token fixed chunks. On 120 questions labelled by in-house lawyers, the governing clause is retrieved in full only 52% of the time; answers often quote a liability cap while missing the exception on the next page.

They switch to clause-based chunking with Docling, heading-path prefixes and parent expansion. Clause-level recall@4 rises to 88%. The remaining misses are mostly 140 scanned contracts where numbering was lost; those are re-processed with a better OCR step and manual heading checks.

Follow-up questions to expect

  • "What if a clause is longer than the model's context budget?" — Send the matching sub-chunks plus the clause's heading and a short summary of the rest, and mark it as partial.
  • "Why not semantic chunking by embedding similarity?" — It helps for unstructured prose, but contracts already have explicit structure, which is more reliable than inferred topic shifts.
  • "How do you evaluate chunking?" — Clause-level recall on expert-labelled questions: did the full governing clause reach the model?