LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you log errors in LangChain applications?


What you need to know

An error-logging handler

Python
import loggingfrom langchain_core.callbacks import BaseCallbackHandlerlog = logging.getLogger("llm")class ErrorLogger(BaseCallbackHandler):    def __init__(self, request_id: str, tenant: str):        self.base = {"request_id": request_id, "tenant": tenant}    def _log(self, kind, error, run_id, parent_run_id):        log.error(kind, exc_info=error, extra={            **self.base, "run_id": str(run_id), "parent_run_id": str(parent_run_id),            "error_type": type(error).__name__})    def on_llm_error(self, error, *, run_id, parent_run_id=None, **kw):        self._log("llm_error", error, run_id, parent_run_id)    def on_tool_error(self, error, *, run_id, parent_run_id=None, **kw):        self._log("tool_error", error, run_id, parent_run_id)    def on_retriever_error(self, error, *, run_id, parent_run_id=None, **kw):        self._log("retriever_error", error, run_id, parent_run_id)    def on_chain_error(self, error, *, run_id, parent_run_id=None, **kw):        if parent_run_id is None:                    # only the top-level run            self._log("chain_error", error, run_id, parent_run_id)chain.invoke(inputs, config={"callbacks": [ErrorLogger(req_id, tenant)], "run_id": run_id})

Details worth mentioning

  • on_chain_error fires at every level. A failure in step 3 of a sequence raises an error event for step 3 and for the sequence around it. Logging only when parent_run_id is None avoids duplicate lines; the specific model or tool error is logged by its own hook.
  • run_id in config — pass your own UUID as run_id. The LangSmith trace then has the same id as your log line, so you can jump from one to the other.
  • Handlers passed in config are inherited by every child step, including agent tool calls.
  • Structured logs — JSON with a fixed set of fields is searchable; free-text messages are not.

What never to log

  • API keys or tokens (from config objects or headers).
  • Unredacted prompts and outputs that may contain personal data such as phone numbers, PAN or Aadhaar numbers. Redact at the logging boundary with one shared function.

A real-life example

A telecom company's support bot handles 50,000 chats a day. Errors were logged with print(e) in three places, so an alert said "Error: timeout" with no chat id, step or model.

After adding ErrorLogger with request id, tenant and error_type, the on-call engineer can filter: 92% of yesterday's errors were tool_error from the "check bill" tool, all with ReadTimeout, all in one region. The billing API in that region was slow. They added a 5-second tool timeout and a friendly fallback message, and the next incident was diagnosed in minutes from the dashboard instead of hours of guessing.

Follow-up questions to expect

  • "Callback handler or LangSmith?" — Both: LangSmith for deep per-trace debugging, your own handler to feed the company's log and alerting stack.
  • "Do callbacks slow down the chain?" — Handlers run inline by default; keep them fast, or make the handler write to a queue.
  • "How do you log errors from agent middleware?" — The same handler receives the model and tool errors inside the agent, because the config is passed down.