Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 9: Performance Bottleneck Investigation


Seconds per node in the p95 trace, all in sequence1.54581.501234policy lookupbureau summary22 chunks,18k tokensThe two middle nodes are independent and can run together.
Parallelising two nodes and reranking to five chunks took the run from about 20 to 8 seconds without touching the answer model.

What you need to know

The scenario: a LangGraph workflow takes 20 seconds end to end, and the team wants it under 8.

Measure first

A trace records each node as a span with start time, duration and token counts. Laid out as a waterfall, it shows which nodes run in sequence, which overlap, and which one dominates. LangSmith does this for LangGraph out of the box; OpenTelemetry spans per node work too. Look at p95 traces, not the average: slow runs are often slow for a different reason.

The usual causes

CauseHow it shows in the traceFix
Independent LLM calls in sequenceSpans end-to-end that don't depend on each otherParallel edges or Send
Oversized promptOne LLM span with huge input tokens and slow first tokenRerank to 3–5 chunks
Slow tool or databaseLong non-LLM spanIndex, cache or timeout it
Repeated workSame embedding or retrieval twiceCompute once, pass in state
Heavy checkpoint writesGaps between nodesKeep state small; store blobs by reference

Fix in payoff order

  1. Parallelise — run independent nodes in the same superstep.
  2. Cut context — fewer, better chunks shorten prefill and cost.
  3. Cache deterministic nodes — embeddings and retrieval keyed by a hash of the input.
  4. Right-size models — classification and extraction nodes rarely need the largest model.
  5. Stream the last node — users see text start sooner, even if total time is the same.

LangGraph supports node-level caching directly:

Python
from langgraph.types import CachePolicyfrom langgraph.cache.memory import InMemoryCachebuilder.add_node("retrieve", retrieve, cache_policy=CachePolicy(ttl=600))graph = builder.compile(checkpointer=saver, cache=InMemoryCache())async for chunk, meta in graph.astream(inputs, config, stream_mode="messages"):    if meta["langgraph_node"] == "answer":        send_to_client(chunk.content)                 # stream only the final node's tokens

Make per-node p50 and p95 latency a standing dashboard with a target per node, fix the top offender, re-measure, and repeat.

A real-life example

Scenario, numbers made up. A loan-eligibility assistant takes about 20 seconds per answer. The p95 trace shows: classify (1.5 s), then policy lookup (4 s), then credit-bureau summary (5 s), then an answer node with 22 retrieved chunks and 18,000 input tokens (8 s), then a formatting LLM call (1.5 s).

The team runs policy lookup and the bureau summary in parallel (saving 4 s), reranks down to 5 chunks (the answer node drops to about 4 s), folds formatting into the answer prompt, and moves classify to a small model. End-to-end time falls to about 8 seconds, and with streaming, users see the first words in under 3 seconds. The answer quality score on their golden set does not change.

Follow-up questions to expect

  • "How do you know it's the model and not the network?" — Compare time-to-first-token with total span time and input token counts; long prefill with huge inputs points at the prompt.
  • "Is streaming a real fix?" — It fixes perceived latency, which is what users feel, but not total time or cost; do it alongside real fixes.
  • "When would you switch to a smaller model for the final answer?" — Only if the golden-set score holds; measure quality before and after every speed change.