Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your multi-agent system has planner, researcher, and executor agents. Latency explodes because agents keep re-explaining context to each other. How do you design efficient inter-agent communication without excessive token overhead?


Retelling the story versus sharing the notebookProse handoffs• Each agent re-sends everything so far• Tokens grow roughly with hops squared• Scraped pages pasted into messages• Details lost in every retellingShared typed state• Agents read and write one workspace• Tokens grow roughly linearly• Artifacts passed as ids• Each role sees only its own view
Nothing needs re-explaining when the state is the context, so the cheapest handoff is the one that sends no prose at all.

What you need to know

Why the cost grows so fast

In a prose design, each agent writes a message that repeats the context, and the next agent's prompt includes all earlier messages. If every hop re-sends a growing transcript, hop 1 sends a little, hop 2 sends more, and hop 9 sends almost everything. Total tokens grow roughly with the square of the number of hops. Latency follows, because every one of those tokens must be prefilled before the next agent starts writing.

Prose handoffs

  • "Here is a summary of what we know so far..."
  • Every agent re-reads everything
  • Tokens grow roughly quadratically with hops
  • Detail lost in each retelling

Shared typed state

  • Agents read and write fields in one workspace
  • Each agent sees only its projected view
  • Tokens grow roughly linearly with hops
  • Facts stored once, referenced by id

The design

Python
from pydantic import BaseModelclass Finding(BaseModel):    claim: str    source_id: str                  # a reference, not the scraped pageclass Workspace(BaseModel):    goal: str    plan: list[str]    current_step: int = 0    findings: list[Finding] = []    artifacts: dict[str, str] = {}  # name -> storage iddef executor_view(ws: Workspace) -> dict:    return {"goal": ws.goal,            "step": ws.plan[ws.current_step],            "evidence": [f.claim for f in ws.findings[-5:]]}

The workspace is the context, so nothing needs re-explaining. executor_view is a projection: the executor gets the current step and recent evidence, not the researcher's raw notes. In LangGraph this is the graph state; in CrewAI it is structured task outputs plus shared memory.

  1. Shared state instead of messages between agents.
  2. References, not payloads — an artifact id instead of 20K tokens of scraped text; the consumer fetches what it needs.
  3. Per-role views — usually the single biggest saving.
  4. Typed messages for what remains: {intent, inputs, constraints, refs}. A schema keeps the sender brief.
  5. Stable cached prefix — shared system prompts and tool definitions go first and stay identical, so provider prompt caching cuts cost and prefill time on every call.
  6. Delete hops — if the planner only passes a query string to the researcher, make it a function call.

Measure tokens per handoff, total tokens per run, hops per run and p95 latency.

A real-life example

Scenario, numbers made up. A market-research product runs planner, researcher, writer and reviewer agents. One report takes 9 hops, 140K tokens and 48 seconds at p95. Traces show the researcher pastes full scraped pages into its messages, and every later agent reads them.

The team moves to a shared workspace, stores pages in object storage with ids, and gives the writer only the claims and their source ids. They merge the planner into the researcher, because it only produced search queries. Tokens per run drop to 38K, p95 latency to 17 seconds, and reviewer complaints about missing sources fall because every claim now carries its source_id.

Follow-up questions to expect

  • "What if two agents write the same field?" — Give each field one owner, or use merge rules (LangGraph reducers do this, such as appending to a list).
  • "When is a multi-agent design worth it at all?" — When roles need different tools, prompts or models, or can run in parallel. Otherwise one agent with good tools is faster and cheaper.
  • "How does prompt caching help here?" — The provider reuses its processed version of an identical prompt prefix, so repeated calls with the same system prompt and tools start faster and cost less.