Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

How would you summarise 300-page contracts when the long-context model keeps missing details in the middle?


Map-reduce over one vendor contract300-pagecontract,about 150K tokensSplit on clauseheadings, 15% overlapMap: one JSONnote per sectionReduce notes,keep section idsAnswer citesclause 12.3Notes are cached per contract version.
Structured notes keep "30 days from invoice" word for word, where a summary of a summary would blur it.

What you need to know

Why a bigger context window does not fix this

A long context window means the model accepts 300 pages. It does not mean it uses every page equally well. Research published in 2023 as "Lost in the Middle" showed that models recall information best near the start and end of a long prompt and worse in the middle. Newer models are better at this, but the pattern has not disappeared, and one missed indemnity clause is a serious error in a contract.

There is also cost. A 300-page contract is roughly 150,000 tokens. Sending it on every follow-up question multiplies that bill by the number of questions. Prompt caching lowers the price of a repeated prefix but does not make the model read it more carefully.

The plan

  1. Split on structure — use headings and clause numbers ("12. Indemnity", "12.3 Limitation of Liability"), not fixed 1,000-token windows. Add about 15% overlap where a clause crosses a page break.
  2. Map to structured notes — ask for a small JSON object per section, not prose.
  3. Reduce — combine the notes into the final summary or answer, keeping section ids so every claim can cite its source.
  4. Add a second reduce tier — if the notes themselves are too long, reduce them in groups first.
Python
notes = [extract_note(s) for s in sections]   # run in parallel# each note: {"section": "12.3", "obligations": [...], "dates": [...], "amounts": [...]}answer = reduce_notes(notes, question)          # cites note["section"] for each claim

Structured notes matter. A prose summary of a summary slowly paraphrases away the details ("30 days" becomes "a short period"). A field called dates keeps "30 days from invoice" as it is. The section id on each note is what lets the final answer say "see clause 12.3".

Trade-offs

ApproachGood atWeak at
Whole document in one promptShort documents, quick prototypesMiddle-of-document recall, repeated cost
Map-reduce with structured notesLong documents, citations, parallel speedCross-section reasoning needs a careful reduce prompt
RAG over clausesSpecific questions ("what is the notice period?")"Summarise everything" questions

For a question like "what is the termination notice period?", retrieving the right clause is cheaper than summarising all of it. For "summarise all obligations on the supplier", map-reduce is the right tool. Many products use both.

Failure modes to name

  • A clause that spans a section boundary, fixed with heading-aware splitting and overlap.
  • The reduce step itself exceeding context, fixed with a second reduce tier.
  • Scanned PDFs where OCR loses the clause numbering, so the structure disappears. Check it at upload.

The map step is also cacheable per document version. If ten users ask about the same contract, you extract the notes once.

A real-life example

Scenario (illustrative numbers). A legal-tech startup lets procurement teams upload vendor contracts of 200 to 300 pages. Its first version sends the whole PDF to a long-context model. On a 30-question test set, where a lawyer has marked the correct page for each answer, it gets 21 right. Most misses are clauses between pages 90 and 200, such as liquidated damages and audit rights.

The team switches to map-reduce. Each of about 60 sections becomes a JSON note with obligations, dates, amounts and the section id. The reduce step answers from the notes. The score rises to 28 of 30, and each answer now cites a clause the lawyer can open. Because notes are cached per contract version, follow-up questions cost a fraction of a full re-read.

Follow-up questions to expect

  • "Why not just use RAG?" — RAG is best for targeted questions. A full summary needs every section considered, and top-k retrieval returns only a few passages.
  • "How do you keep the reduce step from losing details?" — Use structured fields, keep section ids, and tell the reduce prompt to copy numbers and dates exactly rather than rephrase them.
  • "How do you measure quality?" — Questions with known page locations; score both the answer and whether the cited section is correct.