Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 7: Parallel Research Orchestration


Map with Send, reduce in one joinPlannersplits intosub-questionsSend: atmost 6 branchesEach branch:search, 20 s timeoutfindings listvia operator.addSynthesise,naming any gapsLatency becomes the slowest branch; cost grows with the branch count.
The run fell from about 100 to 30 seconds, and switching off one source now yields a report with a stated gap instead of a failure.

What you need to know

The scenario: a research agent answers one sub-question at a time and takes minutes. You need to run the sub-questions in parallel and combine the results.

The shape: map, then reduce

Python
import operatorfrom typing import Annotated, TypedDictfrom langgraph.types import Sendclass State(TypedDict):    question: str    subqueries: list[str]    findings: Annotated[list[dict], operator.add]     # each branch appends    answer: strdef fan_out(state: State):    return [Send("research", {"subq": q}) for q in state["subqueries"][:6]]   # cap the widthasync def research(inp: dict):    try:        hits = await asyncio.wait_for(search(inp["subq"]), timeout=20)        return {"findings": [{"subq": inp["subq"], "text": h.text, "source": h.id} for h in hits]}    except Exception as e:        return {"findings": [{"subq": inp["subq"], "error": str(e)}]}          # partial, not fatalbuilder.add_conditional_edges("plan", fan_out, ["research"])builder.add_edge("research", "synthesise")

The synthesis node runs once every branch in the superstep has finished.

The guardrails

GuardrailWhy
Cap the width (e.g. 6)Planners love to split into 20 sub-questions, and cost grows with each
Per-branch timeoutOne slow source should not hold the whole answer
Tolerate partialsSynthesis answers from what returned and names the gaps
Deduplicate findingsThree sources repeating one fact should not look like three facts
Carry source IDsThe final answer can cite each claim

The economics

Latency becomes roughly the slowest branch plus the synthesis step, instead of the sum of all branches. Cost is roughly N times one branch plus synthesis. That trade is usually right for research reports and wrong for quick chat replies, where a single retrieval is enough.

A real-life example

Scenario, numbers made up. An equity-research tool answers "How exposed is this company to the EV transition?" by researching suppliers, regulation, competitors, financials and news one after another. Each branch takes 15–25 seconds; the full run takes about 100 seconds.

With Send, a width cap of 6 and a 20-second branch timeout, the run takes about 30 seconds. On a test where the news API is switched off, the report still arrives, with a line saying "news coverage unavailable". Deduplication removes about a quarter of findings that were the same fact from different sources. Token cost per report rises about 10% because of the planning step, which the team accepts for the time saved.

Follow-up questions to expect

  • "When would you not fan out?" — When sub-questions depend on each other, such as "find the supplier, then check its filings"; those must run in sequence.
  • "How do you handle rate limits with parallel branches?" — Use a semaphore or the model client's concurrency limit, so six branches do not all hit the same API at once.
  • "How do you test it?" — Make one branch fail on purpose and assert that the answer is still complete and names the gap; track partial-result rate in production.