Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 9: Performance Bottleneck Investigation
What you need to know
The scenario: a LangGraph workflow takes 20 seconds end to end, and the team wants it under 8.
Measure first
A trace records each node as a span with start time, duration and token counts. Laid out as a waterfall, it shows which nodes run in sequence, which overlap, and which one dominates. LangSmith does this for LangGraph out of the box; OpenTelemetry spans per node work too. Look at p95 traces, not the average: slow runs are often slow for a different reason.
The usual causes
| Cause | How it shows in the trace | Fix |
|---|---|---|
| Independent LLM calls in sequence | Spans end-to-end that don't depend on each other | Parallel edges or Send |
| Oversized prompt | One LLM span with huge input tokens and slow first token | Rerank to 3–5 chunks |
| Slow tool or database | Long non-LLM span | Index, cache or timeout it |
| Repeated work | Same embedding or retrieval twice | Compute once, pass in state |
| Heavy checkpoint writes | Gaps between nodes | Keep state small; store blobs by reference |
Fix in payoff order
- Parallelise — run independent nodes in the same superstep.
- Cut context — fewer, better chunks shorten prefill and cost.
- Cache deterministic nodes — embeddings and retrieval keyed by a hash of the input.
- Right-size models — classification and extraction nodes rarely need the largest model.
- Stream the last node — users see text start sooner, even if total time is the same.
LangGraph supports node-level caching directly:
1from langgraph.types import CachePolicy2from langgraph.cache.memory import InMemoryCache34builder.add_node("retrieve", retrieve, cache_policy=CachePolicy(ttl=600))5graph = builder.compile(checkpointer=saver, cache=InMemoryCache())67async for chunk, meta in graph.astream(inputs, config, stream_mode="messages"):8 if meta["langgraph_node"] == "answer":9 send_to_client(chunk.content) # stream only the final node's tokensMake per-node p50 and p95 latency a standing dashboard with a target per node, fix the top offender, re-measure, and repeat.
A real-life example
Scenario, numbers made up. A loan-eligibility assistant takes about 20 seconds per answer. The p95 trace shows: classify (1.5 s), then policy lookup (4 s), then credit-bureau summary (5 s), then an answer node with 22 retrieved chunks and 18,000 input tokens (8 s), then a formatting LLM call (1.5 s).
The team runs policy lookup and the bureau summary in parallel (saving 4 s), reranks down to 5 chunks (the answer node drops to about 4 s), folds formatting into the answer prompt, and moves classify to a small model. End-to-end time falls to about 8 seconds, and with streaming, users see the first words in under 3 seconds. The answer quality score on their golden set does not change.
Follow-up questions to expect
- "How do you know it's the model and not the network?" — Compare time-to-first-token with total span time and input token counts; long prefill with huge inputs points at the prompt.
- "Is streaming a real fix?" — It fixes perceived latency, which is what users feel, but not total time or cost; do it alongside real fixes.
- "When would you switch to a smaller model for the final answer?" — Only if the golden-set score holds; measure quality before and after every speed change.