LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you optimize LangChain agents for complex NLP tasks?


The contract assistant before and after, on 80 test questions113860,00071%51614,00084%StepsSecondsPrompt tokensSuccessOne agent, 22 toolsRouted,filtered, subagentsSummaries from a review subagent replaced full clause text in the main context.
Fewer steps and smaller context improved speed and accuracy together, because every step re-reads everything before it.

What you need to know

Where the time and money go

Suppose one model call takes 1.5 seconds and costs Rs 0.40 with a 4,000-token prompt. An agent that takes 6 steps costs about 9 seconds and Rs 2.40 — and the prompt grows each step, because every tool result is added to the history. The levers:

LeverWhat it saves
Fewer stepsLatency and cost, linearly
Smaller tool resultsTokens in every later step
Fewer tools in the promptSchema tokens per call, and wrong choices
Parallel tool callsLatency when lookups are independent
Cheaper model for easy partsCost per call

The main techniques

  • Route before the agent. A classifier or rules send FAQ-style questions to a one-call RAG chain. Only questions that need several lookups reach the agent.
  • Cut and sharpen the tool surface. Merge overlapping tools. With many tools, filter them per request using LLMToolSelectorMiddleware (a small model picks the relevant few) or a wrap_model_call middleware with your own rules.
  • Return small observations. Project the three fields the model needs, not the 6 KB API payload.
  • Use parallel tool calls. Modern models return several tool calls in one turn; the tool node runs them concurrently. Say in the system prompt "call independent tools together".
  • Keep context short. SummarizationMiddleware for long threads; subagents for big side tasks, so their intermediate steps never enter the main context.
  • Use the right model per step. A small model for tool selection or summarisation; a strong one for the final reasoning.
  • Cap and cache. ModelCallLimitMiddleware(run_limit=...), and caching for repeated read-only tool calls. Keep the system prompt and tool list stable so provider prompt caching can reuse the prefix.
Python
from langchain.agents import create_agentfrom langchain.agents.middleware import (LLMToolSelectorMiddleware,    ModelCallLimitMiddleware, SummarizationMiddleware)agent = create_agent(    model=settings.chat_model, tools=all_22_tools,        # model names from config    system_prompt=SYSTEM,                                   # stable prefix for caching    middleware=[        LLMToolSelectorMiddleware(model=settings.small_model, max_tools=4,                                  always_include=["search_help_centre"]),        SummarizationMiddleware(model=settings.small_model, trigger=("tokens", 6000)),        ModelCallLimitMiddleware(run_limit=8, exit_behavior="end"),    ],)

Measure, do not guess

Track per question type: success rate, average steps, tokens and cost per successful task, and p95 latency. Change one thing at a time and re-run the same eval set. An optimisation that saves 30% cost but drops success by 5 points is usually a bad trade.

A real-life example

A law firm's contract assistant handles questions like "Summarise the indemnity terms across all our vendor contracts renewing in Q1." The first version was one agent with 22 tools. Average run: 11 steps, 38 seconds, about 60,000 prompt tokens, because every clause it read stayed in context.

Changes, measured on 80 test questions:

  1. Simple lookups ("Who signed contract 4471?") routed to a chain: 45% of traffic now takes 2 seconds.
  2. LLMToolSelectorMiddleware with max_tools=4: wrong-tool calls fell by half.
  3. A review_contract subagent reads one contract and returns a 150-word summary; the main agent only sees summaries. Prompt tokens per run fell from 60,000 to 14,000.
  4. Parallel calls to review_contract for the 9 matching contracts cut latency to 16 seconds.

Average steps per run fell from 11 to 5, success rate went from 71% to 84%, and cost per successful answer fell by about 65%.

Follow-up questions to expect

  • "How do you reduce latency without changing the model?" — Fewer steps, parallel tool calls, smaller tool outputs, streaming the answer, and routing easy questions away.
  • "When does a bigger model save money?" — When it solves the task in 3 steps where a small model takes 9 or fails; measure cost per successful task, not per call.
  • "How many tools is too many?" — There is no fixed number; when the eval shows wrong-tool choices rising, filter tools per request.