Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Agent A waits for Agent B, Agent B asks Agent C, Agent C calls back into Agent A. Nothing finishes. How do you design multi-agent orchestration that terminates?


A supervisor makes the call graph a treeSupervisor:plan and budgetResearchWrite and reviewsearch tooldata tooldraftcheck
Specialists return to the supervisor instead of calling each other, and a tree has no cycle to deadlock on.

What you need to know

A deadlock needs a cycle of waiting. Remove the possibility of a cycle and you remove the deadlock.

Default: the supervisor pattern

Peer-to-peer agents

  • Any agent can call any other
  • Call graph can form cycles
  • Hard to trace or attribute cost
  • Deadlocks and ping-pong loops

Supervisor with specialists

  • One orchestrator calls specialists
  • Specialists return results, never call peers
  • The call graph is a tree
  • Easy to trace, budget and test

Most systems described as "multi-agent" work well as a supervisor plus specialists. Ask whether agents really need to talk to each other, or only need results passed through a coordinator.

When collaboration is real, constrain it

  • Make the graph a DAG (a directed acyclic graph) and validate it at startup. Reject a cycle when the service boots, not at 2 AM.
  • Carry depth and budget in state. Each delegation increments depth; past the cap, an agent must answer with what it has. One shared step and token budget covers the whole run, so one subtree cannot starve the rest.
  • Detect cycles at runtime. Pass the call path with each request and refuse a call to an agent already on it.
  • Don't block. Agents publish requests and continue; results arrive as events. Nobody holds a thread waiting on a peer.
  • Deadlines everywhere. Each call has a timeout with a defined degraded result, so a stuck child yields a partial answer.
Python
def delegate(state: RunState, target: str, task: str):    if target in state.call_path:        return {"error": "cycle", "message": f"{target} is already working on this run."}    if state.depth >= MAX_DEPTH or state.budget.steps_left <= 0:        return {"error": "budget", "message": "Answer with what you have."}    child = state.child(target, depth=state.depth + 1, call_path=state.call_path + [target])    return run_agent(target, task, child, timeout=state.deadline_remaining())

The call path and the shared budget travel with the state, so every agent enforces the same limits.

What frameworks give you

LangGraph makes the graph explicit, enforces a recursion_limit on steps, checkpoints state, and supports supervisor patterns. You still choose the topology; a framework will happily run a badly designed cycle until the limit trips.

  1. Topology — supervisor and specialists by default.
  2. Validation — acyclic graph checked at build time.
  3. Runtime limits — depth, shared budget, call-path cycle check, deadlines.
  4. Chaos test — force a specialist to hang and assert the run ends with a partial result inside the deadline.

A real-life example

Scenario (illustrative numbers). A market-research product has three agents: Planner, Researcher and Writer. The Writer asks the Researcher for missing data, the Researcher asks the Planner to clarify scope, and the Planner waits for the Writer's draft. About 4% of reports never finish, and each stuck run holds workers for 20 minutes.

The team rebuilds it as a supervisor that calls Research, then Write, then a Review step, with the Writer returning "missing data" to the supervisor instead of calling the Researcher. They add a depth cap of 3, a shared budget of 40 steps and a 5-minute deadline. Stuck runs fall to zero; 0.5% of reports finish with a "partial: data unavailable" note, which the team reviews weekly.

Follow-up questions to expect

  • "Isn't a timeout enough?" — A timeout ends the wait but wastes the whole run and hides the design flaw; it is a backstop, not the fix.
  • "When do you really need peer-to-peer agents?" — Rarely, for example in simulations or negotiation between agents. Even then, bound the turns and keep a coordinator that can end the conversation.
  • "How do you trace a multi-agent run?" — One trace id for the run, a span per agent call, and the call path recorded on every span.