LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Implement a LangChain agent to handle API-based tools.


What the tracking tool hands back to the modelRaw API payload• About 6 KB of nested JSON• 40 scan events, oldest first• Model read an old Delhi scan• About 3,500 prompt tokensCompact projection• Three fields: status, city, eta• Only the latest scan• Current city every time• About 900 prompt tokens
Trimming a tool's output is not only cheaper — it removes the stale fields the model was misreading.

What you need to know

One endpoint, one tool

A model choosing between get_order, track_shipment and create_return is far more reliable than one filling in http_request(method, url, body). Narrow tools also let you validate arguments, and they limit what a confused or manipulated model can do.

The code

Python
import httpx, osfrom typing import Literalfrom pydantic import BaseModel, Fieldfrom langchain.agents import create_agentfrom langchain.agents.middleware import HumanInTheLoopMiddlewarefrom langchain_core.tools import tool, ToolExceptionfrom langgraph.checkpoint.memory import InMemorySaverAPI = "https://api.example-store.in/v2"client = httpx.Client(base_url=API, timeout=5.0,                      headers={"Authorization": f"Bearer {os.environ['STORE_API_KEY']}"})class TrackArgs(BaseModel):    awb: str = Field(description="Courier airway bill number, 10-12 digits",                     pattern=r"^\d{10,12}$")@tool(args_schema=TrackArgs)def track_shipment(awb: str) -> dict:    """Get the latest courier status for a shipment by airway bill number."""    try:        r = client.get(f"/shipments/{awb}")        r.raise_for_status()    except httpx.HTTPStatusError as e:        raise ToolException(f"Tracking returned {e.response.status_code}; ask the user to check the number.")    except httpx.TransportError:        raise ToolException("Tracking service unreachable; say so and offer to check later.")    d = r.json()    return {"status": d["status"], "city": d["last_scan"]["city"], "eta": d["eta_date"]}@tooldef create_return(order_id: str, reason: Literal["damaged", "wrong_item", "not_needed"]) -> str:    """Create a return request for a delivered order."""    r = client.post("/returns", json={"order_id": order_id, "reason": reason})    r.raise_for_status()    return f"Return {r.json()['id']} created."for t in (track_shipment, create_return):    t.handle_tool_error = Trueagent = create_agent(model, tools=[track_shipment, create_return],    middleware=[HumanInTheLoopMiddleware(interrupt_on={"create_return": True})],    checkpointer=InMemorySaver())   # interrupts need a checkpointer; use Postgres in prod

Why each piece is there

  • Shared httpx.Client with timeout=5.0 — connection reuse and no hanging calls. Without a timeout, one slow endpoint can freeze the run.
  • Key in the client headers — the model never sees it and cannot be tricked into sending it somewhere else.
  • Pydantic schema with pattern and Literal — bad arguments are rejected before any HTTP call, and the model gets the validation error to fix.
  • Compact projection — the real tracking response is about 6 KB of nested JSON with 40 scan events. The tool returns three fields. Less context, fewer misreadings.
  • ToolException per failure type — the message tells the model what to do next.
  • HumanInTheLoopMiddleware — the run pauses before create_return executes; a person or the user confirms, and the run resumes from the checkpoint.

Generic API toolkits

langchain-community has RequestsToolkit and OpenAPI-based toolkits that let a model call arbitrary endpoints. They require an explicit allow_dangerous_requests=True for a reason: a model that can call any URL is a server-side request forgery (SSRF) risk and can reach internal services. Use them only in sandboxes, or behind a strict allowlist.

A real-life example

An online electronics store's support bot gets "Where is my order? AWB 41893022715." The agent calls track_shipment, gets {"status": "out_for_delivery", "city": "Pune", "eta": "2026-09-25"}, and answers in one sentence.

The first version returned the full tracking payload. On shipments with many scans, the model sometimes read an old "in transit — Delhi" event as the current status and told customers the wrong city. After switching to the three-field projection, those complaints stopped, and the average prompt size for tracking questions dropped from about 3,500 tokens to 900.

Follow-up questions to expect

  • "How do you handle API rate limits?" — Retry 429s with backoff (respecting Retry-After) in the client or with ToolRetryMiddleware, and cache read-only calls.
  • "How do you pass the user's identity to the API?" — From runtime context set by your authenticated backend, read inside the tool through ToolRuntime — never as a model argument.
  • "Sync or async tools?" — Async (httpx.AsyncClient) when the agent runs in an async server, so parallel tool calls do not block each other.