Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your multi-agent system has planner, researcher, and executor agents. Latency explodes because agents keep re-explaining context to each other. How do you design efficient inter-agent communication without excessive token overhead?
What you need to know
Why the cost grows so fast
In a prose design, each agent writes a message that repeats the context, and the next agent's prompt includes all earlier messages. If every hop re-sends a growing transcript, hop 1 sends a little, hop 2 sends more, and hop 9 sends almost everything. Total tokens grow roughly with the square of the number of hops. Latency follows, because every one of those tokens must be prefilled before the next agent starts writing.
Prose handoffs
- "Here is a summary of what we know so far..."
- Every agent re-reads everything
- Tokens grow roughly quadratically with hops
- Detail lost in each retelling
Shared typed state
- Agents read and write fields in one workspace
- Each agent sees only its projected view
- Tokens grow roughly linearly with hops
- Facts stored once, referenced by id
The design
1from pydantic import BaseModel23class Finding(BaseModel):4 claim: str5 source_id: str # a reference, not the scraped page67class Workspace(BaseModel):8 goal: str9 plan: list[str]10 current_step: int = 011 findings: list[Finding] = []12 artifacts: dict[str, str] = {} # name -> storage id1314def executor_view(ws: Workspace) -> dict:15 return {"goal": ws.goal,16 "step": ws.plan[ws.current_step],17 "evidence": [f.claim for f in ws.findings[-5:]]}The workspace is the context, so nothing needs re-explaining. executor_view is a projection: the executor gets the current step and recent evidence, not the researcher's raw notes. In LangGraph this is the graph state; in CrewAI it is structured task outputs plus shared memory.
- Shared state instead of messages between agents.
- References, not payloads — an artifact id instead of 20K tokens of scraped text; the consumer fetches what it needs.
- Per-role views — usually the single biggest saving.
- Typed messages for what remains:
{intent, inputs, constraints, refs}. A schema keeps the sender brief. - Stable cached prefix — shared system prompts and tool definitions go first and stay identical, so provider prompt caching cuts cost and prefill time on every call.
- Delete hops — if the planner only passes a query string to the researcher, make it a function call.
Measure tokens per handoff, total tokens per run, hops per run and p95 latency.
A real-life example
Scenario, numbers made up. A market-research product runs planner, researcher, writer and reviewer agents. One report takes 9 hops, 140K tokens and 48 seconds at p95. Traces show the researcher pastes full scraped pages into its messages, and every later agent reads them.
The team moves to a shared workspace, stores pages in object storage with ids, and gives the writer only the claims and their source ids. They merge the planner into the researcher, because it only produced search queries. Tokens per run drop to 38K, p95 latency to 17 seconds, and reviewer complaints about missing sources fall because every claim now carries its source_id.
Follow-up questions to expect
- "What if two agents write the same field?" — Give each field one owner, or use merge rules (LangGraph reducers do this, such as appending to a list).
- "When is a multi-agent design worth it at all?" — When roles need different tools, prompts or models, or can run in parallel. Otherwise one agent with good tools is faster and cheaper.
- "How does prompt caching help here?" — The provider reuses its processed version of an identical prompt prefix, so repeated calls with the same system prompt and tools start faster and cost less.