LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you profile LangChain chain performance?


What you need to know

What to measure

  • Latency per step — prompt, retriever, model, parser, tool.
  • Time to first token (TTFT) — how long before the user sees anything; this is what feels slow.
  • Tokens in and out per model call — input tokens drive cost and some latency; output tokens drive most of the generation time.
  • p50 and p95 — averages hide the slow tail that users complain about.

A timing callback

Python
import timefrom langchain_core.callbacks import BaseCallbackHandlerclass StepTimer(BaseCallbackHandler):    """Records wall-clock milliseconds for chains, chat models and retrievers."""    def __init__(self):        self.started, self.spans = {}, []    def _start(self, run_id, name):        self.started[run_id] = (name, time.perf_counter())    def _end(self, output, *, run_id, **kw):        name, t0 = self.started.pop(run_id)        self.spans.append((name, round((time.perf_counter() - t0) * 1000, 1)))    def on_chain_start(self, serialized, inputs, *, run_id, **kw):        self._start(run_id, kw.get("name") or "chain")    def on_chat_model_start(self, serialized, messages, *, run_id, **kw):        self._start(run_id, "chat_model")    def on_retriever_start(self, serialized, query, *, run_id, **kw):        self._start(run_id, "retriever")    on_chain_end = on_llm_end = on_retriever_end = _endtimer = StepTimer()chain.invoke(inputs, config={"callbacks": [timer]})print(timer.spans)   # [('ChatPromptTemplate', 0.1), ('retriever', 180.4), ('chat_model', 2310.7), ...]

Chat models fire on_chat_model_start (not on_llm_start), and every run ends with an *_end event carrying the same run_id, which is how start and end are matched.

The usual fixes, by impact

FindingFix
Two model calls that don't depend on each other run one after anotherRunnableParallel — latency becomes the slower one, not the sum
Same question asked oftenCache (exact or semantic)
Long promptsFewer retrieved chunks, shorter system prompt, provider prompt caching for the fixed prefix
Long answersAsk for shorter output; set max_tokens
Simple step on a large modelUse a smaller, faster model for classification or routing
User waits for the full answerStream tokens (stream / astream) to cut perceived latency
Many inputs processed one by onebatch / abatch with max_concurrency

A real-life example

A support bot's p95 latency was 9 seconds. The LangSmith trace of a slow request showed: query rewrite 1.8 s (large model), retrieval 0.2 s, answer 4.5 s, then a "safety check" model call of 2.3 s — all sequential.

Changes: the query rewrite moved to a small model (0.5 s); the safety check ran in parallel with the start of streaming and could cancel the reply; answers were streamed. p95 total time fell to 5.4 seconds and time to first token from 8.8 s to 1.1 s. The Python code itself took under 30 ms, so cProfile would have found nothing useful.

Follow-up questions to expect

  • "How do you profile in production?" — Sample traces into LangSmith or export to OpenTelemetry, and track p50/p95 per step on a dashboard.
  • "Does async make chains faster?" — Not one request by itself; it lets one worker serve many concurrent requests and run independent calls at the same time.
  • "How do you measure TTFT?" — Time from request to the first streamed chunk, recorded in the streaming handler or callbacks (on_llm_new_token).