Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your RAG system retrieves highly relevant chunks — but answers still feel incomplete because critical context exists across multiple documents. How do you perform context stitching and synthesis across retrieved sources?


What you need to know

Why "relevant" is not "complete"

Chunks are cut to a fixed size so they embed well. A chunk that matches "How many days of paternity leave do I get?" may say "Eligible employees receive 15 days", while "eligible" is defined two pages earlier and a 2026 update changed the number. Each chunk is relevant; none is complete.

Stitching the evidence

Python
def build_context(hits, budget_tokens=6000):    expanded = []    for h in hits:        section = store.parent_section(h.chunk_id)            # the enclosing section        neighbours = store.adjacent(h.chunk_id, before=1, after=1)        expanded.append(merge(section, neighbours))    groups = cluster_by_similarity(expanded, threshold=0.9)   # near-duplicates together    best = [max(g, key=lambda c: (c.effective_date, c.authority)) for g in groups]    best.sort(key=lambda c: c.score, reverse=True)    return fit_to_budget(best, budget_tokens)

Each hit is widened to its section and neighbours, near-duplicates are grouped, and each group keeps its newest and most authoritative member. fit_to_budget stops before the context gets too long.

Order matters too. Models use information at the start and end of a long context more reliably than the middle, so put the strongest evidence first and last.

Map-reduce synthesis for multi-document questions

  1. Map — a small model reads each document and writes notes: {claim, source, effective_date, confidence}.
  2. Reduce — the large model reads only the notes, prefers newer and more authoritative sources, and writes the answer.
  3. Flag contradictions — if two sources disagree, say so and name both, instead of quietly averaging.
  4. Name the version — "Per the leave policy effective 1 April 2026..."

Map-reduce multiplies model calls, so use it only for questions a router marks as multi-document, and use a small model for the map step.

TechniqueFixesCost
Parent-document expansionCut-off definitions and contextMore tokens per hit
Dedupe and clusterNear-copies crowding out factsOne similarity pass
Map-reduceFacts spread across many documentsSeveral extra calls
effective_date on chunksOld versions outranking new onesMetadata at ingestion

Measure answer completeness against a rubric of required facts, citation coverage (every claim traceable to a source), and contradiction detection rate.

A real-life example

Scenario, numbers made up. An IT services company's HR assistant answers leave questions. Reviewers mark 35% of answers "incomplete": they give the leave count but miss eligibility rules and the 2026 change.

The team adds parent-section expansion and effective_date on every chunk. The top 8 hits often contained four near-copies of the same paragraph from three versions of the policy; clustering keeps only the newest. For questions touching more than one policy, map-reduce extracts dated claims and the answer names the version used. On a 150-question rubric, completeness rises from 65% to 88%, and cost per answer rises about 30%, mostly from the multi-document path.

Follow-up questions to expect

  • "Why not just raise top-K?" — More hits add more near-duplicates and more middle-of-context noise. Expansion adds the right surrounding text.
  • "How big should the parent be?" — A logical section, capped in tokens. A whole 80-page document is too much; a heading-level section usually works.
  • "How do you detect a contradiction?" — Compare map-step claims about the same attribute; different values with different dates are a version change, same dates are a real conflict to flag.