Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your RAG system retrieves highly relevant chunks — but answers still feel incomplete because critical context exists across multiple documents. How do you perform context stitching and synthesis across retrieved sources?
What you need to know
Why "relevant" is not "complete"
Chunks are cut to a fixed size so they embed well. A chunk that matches "How many days of paternity leave do I get?" may say "Eligible employees receive 15 days", while "eligible" is defined two pages earlier and a 2026 update changed the number. Each chunk is relevant; none is complete.
Stitching the evidence
1def build_context(hits, budget_tokens=6000):2 expanded = []3 for h in hits:4 section = store.parent_section(h.chunk_id) # the enclosing section5 neighbours = store.adjacent(h.chunk_id, before=1, after=1)6 expanded.append(merge(section, neighbours))7 groups = cluster_by_similarity(expanded, threshold=0.9) # near-duplicates together8 best = [max(g, key=lambda c: (c.effective_date, c.authority)) for g in groups]9 best.sort(key=lambda c: c.score, reverse=True)10 return fit_to_budget(best, budget_tokens)Each hit is widened to its section and neighbours, near-duplicates are grouped, and each group keeps its newest and most authoritative member. fit_to_budget stops before the context gets too long.
Order matters too. Models use information at the start and end of a long context more reliably than the middle, so put the strongest evidence first and last.
Map-reduce synthesis for multi-document questions
- Map — a small model reads each document and writes notes:
{claim, source, effective_date, confidence}. - Reduce — the large model reads only the notes, prefers newer and more authoritative sources, and writes the answer.
- Flag contradictions — if two sources disagree, say so and name both, instead of quietly averaging.
- Name the version — "Per the leave policy effective 1 April 2026..."
Map-reduce multiplies model calls, so use it only for questions a router marks as multi-document, and use a small model for the map step.
| Technique | Fixes | Cost |
|---|---|---|
| Parent-document expansion | Cut-off definitions and context | More tokens per hit |
| Dedupe and cluster | Near-copies crowding out facts | One similarity pass |
| Map-reduce | Facts spread across many documents | Several extra calls |
effective_date on chunks | Old versions outranking new ones | Metadata at ingestion |
Measure answer completeness against a rubric of required facts, citation coverage (every claim traceable to a source), and contradiction detection rate.
A real-life example
Scenario, numbers made up. An IT services company's HR assistant answers leave questions. Reviewers mark 35% of answers "incomplete": they give the leave count but miss eligibility rules and the 2026 change.
The team adds parent-section expansion and effective_date on every chunk. The top 8 hits often contained four near-copies of the same paragraph from three versions of the policy; clustering keeps only the newest. For questions touching more than one policy, map-reduce extracts dated claims and the answer names the version used. On a 150-question rubric, completeness rises from 65% to 88%, and cost per answer rises about 30%, mostly from the multi-document path.
Follow-up questions to expect
- "Why not just raise top-K?" — More hits add more near-duplicates and more middle-of-context noise. Expansion adds the right surrounding text.
- "How big should the parent be?" — A logical section, capped in tokens. A whole 80-page document is too much; a heading-level section usually works.
- "How do you detect a contradiction?" — Compare map-step claims about the same attribute; different values with different dates are a version change, same dates are a real conflict to flag.