Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 7: Parallel Research Orchestration
What you need to know
The scenario: a research agent answers one sub-question at a time and takes minutes. You need to run the sub-questions in parallel and combine the results.
The shape: map, then reduce
1import operator2from typing import Annotated, TypedDict3from langgraph.types import Send45class State(TypedDict):6 question: str7 subqueries: list[str]8 findings: Annotated[list[dict], operator.add] # each branch appends9 answer: str1011def fan_out(state: State):12 return [Send("research", {"subq": q}) for q in state["subqueries"][:6]] # cap the width1314async def research(inp: dict):15 try:16 hits = await asyncio.wait_for(search(inp["subq"]), timeout=20)17 return {"findings": [{"subq": inp["subq"], "text": h.text, "source": h.id} for h in hits]}18 except Exception as e:19 return {"findings": [{"subq": inp["subq"], "error": str(e)}]} # partial, not fatal2021builder.add_conditional_edges("plan", fan_out, ["research"])22builder.add_edge("research", "synthesise")The synthesis node runs once every branch in the superstep has finished.
The guardrails
| Guardrail | Why |
|---|---|
| Cap the width (e.g. 6) | Planners love to split into 20 sub-questions, and cost grows with each |
| Per-branch timeout | One slow source should not hold the whole answer |
| Tolerate partials | Synthesis answers from what returned and names the gaps |
| Deduplicate findings | Three sources repeating one fact should not look like three facts |
| Carry source IDs | The final answer can cite each claim |
The economics
Latency becomes roughly the slowest branch plus the synthesis step, instead of the sum of all branches. Cost is roughly N times one branch plus synthesis. That trade is usually right for research reports and wrong for quick chat replies, where a single retrieval is enough.
A real-life example
Scenario, numbers made up. An equity-research tool answers "How exposed is this company to the EV transition?" by researching suppliers, regulation, competitors, financials and news one after another. Each branch takes 15–25 seconds; the full run takes about 100 seconds.
With Send, a width cap of 6 and a 20-second branch timeout, the run takes about 30 seconds. On a test where the news API is switched off, the report still arrives, with a line saying "news coverage unavailable". Deduplication removes about a quarter of findings that were the same fact from different sources. Token cost per report rises about 10% because of the planning step, which the team accepts for the time saved.
Follow-up questions to expect
- "When would you not fan out?" — When sub-questions depend on each other, such as "find the supplier, then check its filings"; those must run in sequence.
- "How do you handle rate limits with parallel branches?" — Use a semaphore or the model client's concurrency limit, so six branches do not all hit the same API at once.
- "How do you test it?" — Make one branch fail on purpose and assert that the answer is still complete and names the gap; track partial-result rate in production.