LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to implement a multi-agent system in LangChain.


Supervisor with subagents as toolssupervisorresearchdraftsearch_contractsget_clauseget_template
The supervisor sees a 400-token report, not the 9,000 tokens of clauses the research subagent read — that isolation is the point.

What you need to know

Why split into several agents

  • Focus. An agent with 4 tools and a short prompt chooses better than one with 25 tools and a 3,000-word prompt.
  • Context isolation. A subagent can read 10 documents and return a 150-word summary; the supervisor never sees the 10 documents.
  • Different models. A cheap model for extraction, a strong one for final reasoning.
  • Separate permissions. Only the billing agent has the refund tool.

Common patterns

PatternHow control movesGood for
Subagents as tools (supervisor)Supervisor calls specialists like tools, gets results backMost cases; easy to reason about
HandoffsThe active agent transfers the conversation to anotherCustomer service with departments
Custom graphYour own StateGraph with agent nodes and fixed edgesKnown workflows, such as research then write then review

The code: supervisor with subagents as tools

Python
from langchain.agents import create_agentfrom langchain.tools import toolresearch_agent = create_agent(model, tools=[search_contracts, get_clause],    system_prompt="Find and quote the relevant clauses. Return quotes with contract IDs.")drafting_agent = create_agent(model, tools=[get_template],    system_prompt="Draft client-ready text from the findings you are given.")@tooldef research(task: str) -> str:    """Find clauses and facts in the firm's contracts. Input: a clear research task."""    out = research_agent.invoke({"messages": [{"role": "user", "content": task}]})    return out["messages"][-1].content@tooldef draft(brief: str) -> str:    """Draft a memo or email from findings. Input: the findings and what to write."""    out = drafting_agent.invoke({"messages": [{"role": "user", "content": brief}]})    return out["messages"][-1].contentdef build_supervisor(checkpointer):    return create_agent(model, tools=[research, draft], checkpointer=checkpointer,        system_prompt="Plan the work. Use research for facts, then draft for writing. "                      "Never draft before research has returned quotes.")

The supervisor only sees each subagent's final answer, not its internal steps. The subagents are stateless between calls; conversation memory lives in the supervisor's checkpointer. Subagents run inside the supervisor's tool calls, so a call-limit middleware on each of them keeps the total bounded.

For handoffs, a tool returns a LangGraph Command(goto="billing_agent", graph=Command.PARENT) that moves control to another agent node in a parent graph. The langgraph-supervisor package also offers a prebuilt supervisor graph.

The honest trade-off

Every agent is several model calls, so a supervisor with two subagents easily makes 8 to 12 calls per request. Information can be lost between agents: the supervisor only knows what a subagent chose to report. Start with one well-scoped agent; split only when an eval shows it is confused by too many tools or too much context.

A real-life example

A law firm builds a "client update" assistant. A partner asks: "Draft an email to Sharma Textiles explaining their termination rights under the 2024 supply agreement."

The supervisor calls research with "Find termination clauses in the Sharma Textiles 2024 supply agreement". The research subagent runs 4 tool calls, reads 3 clauses, and returns 5 quotes with section numbers — about 400 tokens, instead of the 9,000 tokens of clause text it read. The supervisor then calls draft with those quotes. The partner gets an email that cites sections 12.1 and 12.3.

With a single agent, the drafting prompt and 9,000 tokens of clauses shared one context, and in testing it mixed up clauses from two different contracts in 6 of 40 cases. The split version did so in 1 of 40, at about 1.6 times the cost per request.

Follow-up questions to expect

  • "How do agents share information?" — Through the supervisor's messages (tool inputs and outputs), or through shared graph state in a custom StateGraph.
  • "How do you stop agents passing work back and forth forever?" — Limit model and tool calls per agent, set a recursion_limit on the whole graph, and give the supervisor a clear stopping rule.
  • "When is multi-agent a bad idea?" — When one agent with 3 to 6 tools handles the task; splitting then only adds cost, latency and places for information to be lost.