Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Build a multi-agent collaboration system with defined roles.
What you need to know
A multi-agent system splits one task among several LLM "agents", each with a narrower job. The benefits are real but specific:
- Smaller, focused context — the critic does not need the researcher's raw search results, only the draft and the key notes.
- Different tools or models per role — the researcher has web search; the writer does not. A cheap model can do research summaries; a strong one can write.
- Checkable hand-offs — each output is stored and can be inspected or replayed.
The costs are just as real: more calls, more latency, and information lost at every hand-off.
The blackboard pattern is the simplest coordination design: a shared store of named entries. Each agent reads the keys it needs and writes its output under a new key. Because no agent calls another directly, you can test each one alone and replay any step.
1from collections.abc import Callable2from dataclasses import dataclass, field34LLM = Callable[[str, str], str] # (system, user) -> reply56@dataclass7class Agent:8 name: str9 system: str10 llm_fn: LLM1112 def run(self, task: str, context: str = "") -> str:13 return self.llm_fn(self.system, f"{context}\n\nTask: {task}".strip())1415@dataclass16class Blackboard:17 task: str18 notes: dict[str, str] = field(default_factory=dict)1920 def context(self, keys: list[str]) -> str:21 return "\n\n".join(f"{k}:\n{self.notes[k]}" for k in keys if k in self.notes)2223def run_crew(task: str, researcher: Agent, writer: Agent, critic: Agent,24 max_revisions: int = 2) -> Blackboard:25 """Research -> draft -> (review -> revise) up to max_revisions times."""26 board = Blackboard(task)27 board.notes["research"] = researcher.run(task)28 if not board.notes["research"].strip():29 raise ValueError("researcher returned nothing; refusing to draft from no facts")30 board.notes["draft"] = writer.run(task, board.context(["research"]))31 for i in range(max_revisions):32 verdict = critic.run(33 "Reply APPROVED if the draft is accurate and complete; otherwise list the fixes.",34 board.context(["research", "draft"]),35 )36 board.notes[f"review_{i}"] = verdict37 if verdict.strip().upper().startswith("APPROVED"):38 break39 board.notes["draft"] = writer.run(40 "Revise the draft to address every point in the review.",41 board.context(["research", "draft", f"review_{i}"]),42 )43 return boardThe tricky parts:
context(keys)passes each agent only the keys it needs. Passing the whole board to every agent grows the prompt with every round and brings back the context problem multi-agent was meant to solve.- The empty-research check stops a confident draft written from nothing.
startswith("APPROVED")afterstrip().upper()tolerates "approved." and leading spaces; still, a structured JSON verdict is more reliable in production.- Returning the board, not just the draft, keeps the whole run inspectable.
Complexity: 2 calls for research and draft, then up to 2 per revision round (review + revise). With max_revisions = 2 that is at most 6 model calls, against 1 for a single agent. Space is the sum of the notes.
A real-life example
Stub models let us trace the run; the critic rejects once, then approves:
1calls: list[str] = []2def fake(role: str, replies: list[str]) -> LLM:3 it = iter(replies)4 def llm(system: str, user: str) -> str:5 calls.append(role)6 return next(it)7 return llm89researcher = Agent("researcher", "Find facts.", fake("researcher", ["UPI refunds: T+5 days."]))10writer = Agent("writer", "Write clearly.", fake("writer", ["Refunds take 5 days.",11 "UPI refunds take 5 working days."]))12critic = Agent("critic", "Be strict.", fake("critic", ["Say 'working days' and name UPI.", "APPROVED"]))13board = run_crew("Explain refund timelines", researcher, writer, critic)14print(board.notes["draft"], calls)15# UPI refunds take 5 working days. ['researcher', 'writer', 'critic', 'writer', 'critic']| call | agent | reads | writes |
|---|---|---|---|
| 1 | researcher | task | research |
| 2 | writer | research | draft v1 |
| 3 | critic | research, draft | review_0: fixes |
| 4 | writer | research, draft, review_0 | draft v2 |
| 5 | critic | research, draft | review_1: APPROVED → stop |
Five calls for one paragraph. That is the number to have in mind when someone proposes a crew of seven agents.
A content team drafting product descriptions for thousands of SKUs uses this shape: a researcher pulls specs from the catalogue, a writer drafts, and a critic checks against brand rules.
Follow-up questions to expect
- "How do you know the critic is useful?" — Feed it known-bad drafts and measure how often it catches them. A critic that approves everything is theatre; a critic that never approves just costs money.
- "Is multi-agent better than one agent with all the tools?" — Usually not. Start with one agent; split only when the prompt is overloaded or roles need different tools or permissions.
- "The last revision is never reviewed — is that a bug?" — If
max_revisionsis exhausted, the final draft was not approved. Return it with a flag likeapproved: Falseso the caller can route it to a human.