Course Content
LangChain Mastery
7 sections · 109 lessons
How do you monitor LangChain agent performance in production?
What you need to know
Why normal API monitoring is not enough
A normal API either works or returns an error code. An agent can return HTTP 200 with a wrong answer, or take 9 steps where 3 were enough. You need to see inside each run, and you need quality signals, not just uptime.
Layer 1: traces
export LANGSMITH_TRACING=trueexport LANGSMITH_API_KEY=...export LANGSMITH_PROJECT=support-bot-prod1agent.invoke(inputs, config={2 "run_name": "support_chat",3 "tags": ["prod", "v42"],4 "metadata": {"request_id": req_id, "user_tier": "gold"},5 "configurable": {"thread_id": thread_id},6})Tags and metadata let you filter traces ("all runs of prompt v42 for gold users") and join them with application logs by request_id. LangSmith can also export to OpenTelemetry, so traces can go to an existing backend like Datadog or Grafana. Mask personal data before it leaves your system, and sample traces if volume is high.
Layer 2: metrics that map to failures
| Metric | What a change usually means |
|---|---|
| p95 latency, per run and per tool | One slow tool or a new loop |
| Average steps per run; % hitting the call limit | Worse tool selection after a prompt or model change |
| Tool error and timeout rate, per tool | A dependency is failing, or bad arguments from a vague description |
| Tokens and cost per successful task | A prompt change that quietly doubled spend |
| Task success rate | The number the business actually cares about |
Layer 3: quality signals
- User feedback — thumbs up or down, "talk to a human" clicks, or the user asking the same question again.
- Online evaluators — an LLM judge scores a sample (say 5%) of production runs for correctness or groundedness. LangSmith can run these automatically on new traces.
- Business outcomes — ticket reopened within 24 hours, return created after a "your order is fine" reply.
Layer 4: closing the loop
- Alert — on p95 latency, error rate, limit-hit rate and cost per task moving beyond a threshold.
- Inspect — open the failing traces and find the first wrong step.
- Capture — add those cases to the regression dataset with the expected answer or tool sequence.
- Gate — run the dataset in CI on every change to prompts, tools or models.
Also set hard budget limits per run (ModelCallLimitMiddleware) and per day. A misbehaving agent loop is a cost incident, not only a quality issue.
A real-life example
A help-centre support bot handles 20,000 conversations a day. One Monday, the dashboard shows average steps per run rising from 3.1 to 4.6 and cost per resolved conversation up 45%, while latency and error rates look normal.
Filtering traces by the v42 tag shows the cause: a weekend prompt edit removed the line "use get_plan only for the customer's own plan", and the agent now calls get_plan and search_help_centre for every pricing question. The team reverts the line, adds 30 pricing questions with the expected tool sequence to the regression set, and makes that set a required CI check. Without step-count monitoring, the only signal would have been the end-of-month bill.
Follow-up questions to expect
- "Which single metric would you watch?" — Cost per successful task, because it moves when quality, steps or prompt size change.
- "How do you measure success without labels?" — Combine implicit feedback (escalations, repeat questions) with an LLM judge on a sample, and spot-check the judge against human ratings.
- "How do you keep traces from leaking personal data?" — Mask or hash fields before export, restrict who can view traces, and set retention limits.