Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
What is agent orchestration, and how do multiple components work together?
What you need to know
The parts
- Router — classifies the request and chooses an agent, tool set or model tier.
- State store — keeps messages, plan, step status and artefacts so a run can resume after a crash or a pause.
- Executor — runs tool calls: in parallel when independent, with timeouts, retries and idempotency keys.
- Policy layer — checks permissions, requires approval for side effects, enforces step and cost budgets.
- Observability — one trace per run, with a span for every model and tool call, including tokens, latency and cost.
An executor for one turn of tool calls
1import asyncio23async def execute_turn(blocks, user, run):4 async def one(block):5 tool = REGISTRY[block.name]6 decision = policy.check(user, tool, block.input, run.budget) # allow / confirm / deny7 if decision.deny:8 return err(block.id, decision.reason)9 if decision.confirm:10 run.pause_for_approval(block) # persisted; resumes later11 return None12 try:13 out = await asyncio.wait_for(tool.run(user, **block.input), timeout=tool.timeout_s)14 return ok(block.id, out)15 except asyncio.TimeoutError:16 return err(block.id, f"{block.name} timed out after {tool.timeout_s}s")17 return await asyncio.gather(*(one(b) for b in blocks))Every tool call passes through policy.check — there is no path around it. Results, including errors, go back to the model as tool_result blocks in one message.
Frameworks
You can write this yourself, or use a framework. LangGraph models the orchestration as a graph of nodes with shared typed state and a checkpointer that saves state after each step — which gives crash recovery and human-in-the-loop pauses. Provider SDKs have tool runners that handle the loop with hooks for approval and logging. Hosted options, such as Anthropic's Managed Agents, also run the loop for you.
A real-life example
A travel company's orchestrator handles "Book me Delhi–Singapore next Friday and a hotel near the office".
- Router: a small model tags it
booking, so the booking agent with flight, hotel and payment tools is used. - Executor: the model calls
search_flightsandsearch_hotelstogether; both run in parallel with 8-second timeouts. The hotel API times out; the executor returns an error result, and the model retries once successfully. - Policy: the model calls
hold_booking(allowed, reversible for 20 minutes), thencharge_cardfor ₹48,300. Policy says payments above ₹0 need user confirmation, so the run pauses and state is saved. - State store: the user confirms 6 minutes later from their phone; a different server resumes the run from the checkpoint.
- Observability: the trace shows 5 model calls, 6 tool calls, 11.2 seconds of active time and ₹4.10 in model cost.
Earlier, the system prompt said "always confirm before charging". In one test the model skipped it after a user said "just do it". The policy-layer rule does not have that weakness.
Follow-up questions to expect
- "Where do retries belong — in the model or the executor?" — Transient errors (timeouts, 503s) are retried by the executor with backoff. Logical errors (bad arguments) go back to the model to fix.
- "How do you make resume safe?" — Idempotency keys on side-effect tools, so a step that runs twice after a crash has no double effect.
- "Do you need a framework?" — Not for simple loops. Once you need durable pauses, resume and branching, a framework's checkpointing saves real work.