Course Content
LangChain Mastery
7 sections · 109 lessons
How do you log errors in LangChain applications?
What you need to know
An error-logging handler
Python
1import logging2from langchain_core.callbacks import BaseCallbackHandler34log = logging.getLogger("llm")56class ErrorLogger(BaseCallbackHandler):7 def __init__(self, request_id: str, tenant: str):8 self.base = {"request_id": request_id, "tenant": tenant}910 def _log(self, kind, error, run_id, parent_run_id):11 log.error(kind, exc_info=error, extra={12 **self.base, "run_id": str(run_id), "parent_run_id": str(parent_run_id),13 "error_type": type(error).__name__})1415 def on_llm_error(self, error, *, run_id, parent_run_id=None, **kw):16 self._log("llm_error", error, run_id, parent_run_id)1718 def on_tool_error(self, error, *, run_id, parent_run_id=None, **kw):19 self._log("tool_error", error, run_id, parent_run_id)2021 def on_retriever_error(self, error, *, run_id, parent_run_id=None, **kw):22 self._log("retriever_error", error, run_id, parent_run_id)2324 def on_chain_error(self, error, *, run_id, parent_run_id=None, **kw):25 if parent_run_id is None: # only the top-level run26 self._log("chain_error", error, run_id, parent_run_id)2728chain.invoke(inputs, config={"callbacks": [ErrorLogger(req_id, tenant)], "run_id": run_id})Details worth mentioning
on_chain_errorfires at every level. A failure in step 3 of a sequence raises an error event for step 3 and for the sequence around it. Logging only whenparent_run_id is Noneavoids duplicate lines; the specific model or tool error is logged by its own hook.run_idin config — pass your own UUID asrun_id. The LangSmith trace then has the same id as your log line, so you can jump from one to the other.- Handlers passed in
configare inherited by every child step, including agent tool calls. - Structured logs — JSON with a fixed set of fields is searchable; free-text messages are not.
What never to log
- API keys or tokens (from config objects or headers).
- Unredacted prompts and outputs that may contain personal data such as phone numbers, PAN or Aadhaar numbers. Redact at the logging boundary with one shared function.
A real-life example
A telecom company's support bot handles 50,000 chats a day. Errors were logged with print(e) in three places, so an alert said "Error: timeout" with no chat id, step or model.
After adding ErrorLogger with request id, tenant and error_type, the on-call engineer can filter: 92% of yesterday's errors were tool_error from the "check bill" tool, all with ReadTimeout, all in one region. The billing API in that region was slow. They added a 5-second tool timeout and a friendly fallback message, and the next incident was diagnosed in minutes from the dashboard instead of hours of guessing.
Follow-up questions to expect
- "Callback handler or LangSmith?" — Both: LangSmith for deep per-trace debugging, your own handler to feed the company's log and alerting stack.
- "Do callbacks slow down the chain?" — Handlers run inline by default; keep them fast, or make the handler write to a queue.
- "How do you log errors from agent middleware?" — The same handler receives the model and tool errors inside the agent, because the config is passed down.