LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you implement logging for distributed LangChain applications?


One request id across three servicesGatewaycreatesrequest_idAgent: run_id,metadata, tagsTrace headers tobooking serviceBooking toolcalls theairline APIOne searchabletrace, end to endLangSmith headers or OpenTelemetry context carry the trace between processes.
Without a shared id, one customer complaint means searching three log systems by timestamp.

What you need to know

A "distributed" LLM app is common: an API gateway, a retrieval service, an agent service, and tool services such as payments or HR. One user message can touch four processes. Without a shared id, you have four unrelated log streams.

Tag every run

Python
import uuidrequest_id = uuid.uuid4()result = chain.invoke(inputs, config={    "run_id": request_id,                  # becomes the root run's id in LangSmith    "run_name": "support_answer",    "metadata": {"tenant": tenant_id, "app_version": "2.4.1", "user_tier": "gold"},    "tags": ["prod", "rag"],    "callbacks": [json_log_handler],})

Metadata and tags are inherited by every child run (model, retriever, tools) and are searchable in LangSmith. Your own logs include the same request_id, so you can jump between them.

Continue the trace across services

With LangSmith, the calling service sends the current trace position in headers and the receiving service continues it:

Python
# service A (caller)from langsmith.run_helpers import get_current_run_treeheaders = {}if run_tree := get_current_run_tree():    headers.update(run_tree.to_headers())        # langsmith-trace + baggagehttpx.post(f"{HR_SERVICE}/file-leave", json=payload, headers=headers)# service B (FastAPI handler)import langsmith as lswith ls.tracing_context(parent=request.headers):    result = leave_agent.invoke(...)

With OpenTelemetry, you use the standard inject/extract propagation instead, and LangSmith can receive OTel spans, so LLM spans sit next to your database and HTTP spans in the same trace.

What every log line should have

FieldWhy
request_id / trace idJoin logs across services
tenant, app_versionFind whose requests and which release
model, stepWhich call failed or was slow
input_tokens, output_tokens, latency_msCost and performance
error_typeGroup failures

Privacy and volume

  • Sample full prompts and outputs (for example 5–10%), but keep metrics for 100% of requests.
  • Redact emails, phone numbers, PAN or card numbers in one shared function at the logging boundary, not at each call site.
  • Retention — keep payload logs for a short, agreed period.

A real-life example

A travel-booking platform has three services: chat gateway, an agent service, and a booking service that calls airline APIs. A customer reports being charged but not getting a ticket. Before the change, engineers searched three log systems by timestamp and guessed.

After the change, the gateway creates request_id, the agent passes it as run_id and metadata, and trace headers go to the booking service. Searching one id shows a single trace: the model called book_flight, the booking service got a 504 from the airline after the payment step, and the agent told the user "booked" because the tool returned an ambiguous message. Diagnosis took ten minutes. The fix was a clear tool error for "payment taken, ticket pending" and a refund workflow.

Follow-up questions to expect

  • "LangSmith or OpenTelemetry?" — LangSmith gives LLM-specific views (prompts, tokens, evaluations); OpenTelemetry fits your existing observability stack. Many teams send to both.
  • "How do you avoid logging secrets?" — Never log config objects or headers wholesale; use an allow-list of fields and a redaction filter.
  • "What about async tasks and queues?" — Put the request id and trace headers in the message payload and restore the tracing context in the worker.