Course Content
LangChain Mastery
7 sections · 109 lessons
How do you implement logging for distributed LangChain applications?
What you need to know
A "distributed" LLM app is common: an API gateway, a retrieval service, an agent service, and tool services such as payments or HR. One user message can touch four processes. Without a shared id, you have four unrelated log streams.
Tag every run
1import uuid23request_id = uuid.uuid4()4result = chain.invoke(inputs, config={5 "run_id": request_id, # becomes the root run's id in LangSmith6 "run_name": "support_answer",7 "metadata": {"tenant": tenant_id, "app_version": "2.4.1", "user_tier": "gold"},8 "tags": ["prod", "rag"],9 "callbacks": [json_log_handler],10})Metadata and tags are inherited by every child run (model, retriever, tools) and are searchable in LangSmith. Your own logs include the same request_id, so you can jump between them.
Continue the trace across services
With LangSmith, the calling service sends the current trace position in headers and the receiving service continues it:
1# service A (caller)2from langsmith.run_helpers import get_current_run_tree3headers = {}4if run_tree := get_current_run_tree():5 headers.update(run_tree.to_headers()) # langsmith-trace + baggage6httpx.post(f"{HR_SERVICE}/file-leave", json=payload, headers=headers)78# service B (FastAPI handler)9import langsmith as ls10with ls.tracing_context(parent=request.headers):11 result = leave_agent.invoke(...)With OpenTelemetry, you use the standard inject/extract propagation instead, and LangSmith can receive OTel spans, so LLM spans sit next to your database and HTTP spans in the same trace.
What every log line should have
| Field | Why |
|---|---|
request_id / trace id | Join logs across services |
tenant, app_version | Find whose requests and which release |
model, step | Which call failed or was slow |
input_tokens, output_tokens, latency_ms | Cost and performance |
error_type | Group failures |
Privacy and volume
- Sample full prompts and outputs (for example 5–10%), but keep metrics for 100% of requests.
- Redact emails, phone numbers, PAN or card numbers in one shared function at the logging boundary, not at each call site.
- Retention — keep payload logs for a short, agreed period.
A real-life example
A travel-booking platform has three services: chat gateway, an agent service, and a booking service that calls airline APIs. A customer reports being charged but not getting a ticket. Before the change, engineers searched three log systems by timestamp and guessed.
After the change, the gateway creates request_id, the agent passes it as run_id and metadata, and trace headers go to the booking service. Searching one id shows a single trace: the model called book_flight, the booking service got a 504 from the airline after the payment step, and the agent told the user "booked" because the tool returned an ambiguous message. Diagnosis took ten minutes. The fix was a clear tool error for "payment taken, ticket pending" and a refund workflow.
Follow-up questions to expect
- "LangSmith or OpenTelemetry?" — LangSmith gives LLM-specific views (prompts, tokens, evaluations); OpenTelemetry fits your existing observability stack. Many teams send to both.
- "How do you avoid logging secrets?" — Never log config objects or headers wholesale; use an allow-list of fields and a redaction filter.
- "What about async tasks and queues?" — Put the request id and trace headers in the message payload and restore the tracing context in the worker.