Course Content
LangChain Mastery
7 sections · 109 lessons
How do you profile LangChain chain performance?
What you need to know
What to measure
- Latency per step — prompt, retriever, model, parser, tool.
- Time to first token (TTFT) — how long before the user sees anything; this is what feels slow.
- Tokens in and out per model call — input tokens drive cost and some latency; output tokens drive most of the generation time.
- p50 and p95 — averages hide the slow tail that users complain about.
A timing callback
Python
1import time2from langchain_core.callbacks import BaseCallbackHandler34class StepTimer(BaseCallbackHandler):5 """Records wall-clock milliseconds for chains, chat models and retrievers."""6 def __init__(self):7 self.started, self.spans = {}, []8 def _start(self, run_id, name):9 self.started[run_id] = (name, time.perf_counter())10 def _end(self, output, *, run_id, **kw):11 name, t0 = self.started.pop(run_id)12 self.spans.append((name, round((time.perf_counter() - t0) * 1000, 1)))13 def on_chain_start(self, serialized, inputs, *, run_id, **kw):14 self._start(run_id, kw.get("name") or "chain")15 def on_chat_model_start(self, serialized, messages, *, run_id, **kw):16 self._start(run_id, "chat_model")17 def on_retriever_start(self, serialized, query, *, run_id, **kw):18 self._start(run_id, "retriever")19 on_chain_end = on_llm_end = on_retriever_end = _end2021timer = StepTimer()22chain.invoke(inputs, config={"callbacks": [timer]})23print(timer.spans) # [('ChatPromptTemplate', 0.1), ('retriever', 180.4), ('chat_model', 2310.7), ...]Chat models fire on_chat_model_start (not on_llm_start), and every run ends with an *_end event carrying the same run_id, which is how start and end are matched.
The usual fixes, by impact
| Finding | Fix |
|---|---|
| Two model calls that don't depend on each other run one after another | RunnableParallel — latency becomes the slower one, not the sum |
| Same question asked often | Cache (exact or semantic) |
| Long prompts | Fewer retrieved chunks, shorter system prompt, provider prompt caching for the fixed prefix |
| Long answers | Ask for shorter output; set max_tokens |
| Simple step on a large model | Use a smaller, faster model for classification or routing |
| User waits for the full answer | Stream tokens (stream / astream) to cut perceived latency |
| Many inputs processed one by one | batch / abatch with max_concurrency |
A real-life example
A support bot's p95 latency was 9 seconds. The LangSmith trace of a slow request showed: query rewrite 1.8 s (large model), retrieval 0.2 s, answer 4.5 s, then a "safety check" model call of 2.3 s — all sequential.
Changes: the query rewrite moved to a small model (0.5 s); the safety check ran in parallel with the start of streaming and could cancel the reply; answers were streamed. p95 total time fell to 5.4 seconds and time to first token from 8.8 s to 1.1 s. The Python code itself took under 30 ms, so cProfile would have found nothing useful.
Follow-up questions to expect
- "How do you profile in production?" — Sample traces into LangSmith or export to OpenTelemetry, and track p50/p95 per step on a dashboard.
- "Does async make chains faster?" — Not one request by itself; it lets one worker serve many concurrent requests and run independent calls at the same time.
- "How do you measure TTFT?" — Time from request to the first streamed chunk, recorded in the streaming handler or callbacks (
on_llm_new_token).