Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Agent A waits for Agent B, Agent B asks Agent C, Agent C calls back into Agent A. Nothing finishes. How do you design multi-agent orchestration that terminates?
What you need to know
A deadlock needs a cycle of waiting. Remove the possibility of a cycle and you remove the deadlock.
Default: the supervisor pattern
Peer-to-peer agents
- Any agent can call any other
- Call graph can form cycles
- Hard to trace or attribute cost
- Deadlocks and ping-pong loops
Supervisor with specialists
- One orchestrator calls specialists
- Specialists return results, never call peers
- The call graph is a tree
- Easy to trace, budget and test
Most systems described as "multi-agent" work well as a supervisor plus specialists. Ask whether agents really need to talk to each other, or only need results passed through a coordinator.
When collaboration is real, constrain it
- Make the graph a DAG (a directed acyclic graph) and validate it at startup. Reject a cycle when the service boots, not at 2 AM.
- Carry depth and budget in state. Each delegation increments
depth; past the cap, an agent must answer with what it has. One shared step and token budget covers the whole run, so one subtree cannot starve the rest. - Detect cycles at runtime. Pass the call path with each request and refuse a call to an agent already on it.
- Don't block. Agents publish requests and continue; results arrive as events. Nobody holds a thread waiting on a peer.
- Deadlines everywhere. Each call has a timeout with a defined degraded result, so a stuck child yields a partial answer.
1def delegate(state: RunState, target: str, task: str):2 if target in state.call_path:3 return {"error": "cycle", "message": f"{target} is already working on this run."}4 if state.depth >= MAX_DEPTH or state.budget.steps_left <= 0:5 return {"error": "budget", "message": "Answer with what you have."}6 child = state.child(target, depth=state.depth + 1, call_path=state.call_path + [target])7 return run_agent(target, task, child, timeout=state.deadline_remaining())The call path and the shared budget travel with the state, so every agent enforces the same limits.
What frameworks give you
LangGraph makes the graph explicit, enforces a recursion_limit on steps, checkpoints state, and supports supervisor patterns. You still choose the topology; a framework will happily run a badly designed cycle until the limit trips.
- Topology — supervisor and specialists by default.
- Validation — acyclic graph checked at build time.
- Runtime limits — depth, shared budget, call-path cycle check, deadlines.
- Chaos test — force a specialist to hang and assert the run ends with a partial result inside the deadline.
A real-life example
Scenario (illustrative numbers). A market-research product has three agents: Planner, Researcher and Writer. The Writer asks the Researcher for missing data, the Researcher asks the Planner to clarify scope, and the Planner waits for the Writer's draft. About 4% of reports never finish, and each stuck run holds workers for 20 minutes.
The team rebuilds it as a supervisor that calls Research, then Write, then a Review step, with the Writer returning "missing data" to the supervisor instead of calling the Researcher. They add a depth cap of 3, a shared budget of 40 steps and a 5-minute deadline. Stuck runs fall to zero; 0.5% of reports finish with a "partial: data unavailable" note, which the team reviews weekly.
Follow-up questions to expect
- "Isn't a timeout enough?" — A timeout ends the wait but wastes the whole run and hides the design flaw; it is a backstop, not the fix.
- "When do you really need peer-to-peer agents?" — Rarely, for example in simulations or negotiation between agents. Even then, bound the turns and keep a coordinator that can end the conversation.
- "How do you trace a multi-agent run?" — One trace id for the run, a span per agent call, and the call path recorded on every span.