Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 5: High Latency in Production


Scenario: a LangChain RAG chain that felt fine in development shows a p95 latency of several seconds in production. How do you bring it down?

One slow request, span by span, in ms9008060250370001234rewrite onlarge modelvector search: 1%sequential lookup15 chunks in prompt
The vector search teams rush to optimise is 1% of the time; the model call and prompt size are nearly all of it.

What you need to know

Development hides latency: one user, a warm cache, short test questions, a small index. Production adds long questions, cold paths, rate limits and concurrency.

Read the waterfall

A LangSmith trace shows each runnable as a span with its duration. A typical slow RAG request looks like this:

SpanTimeShare
Query rewrite (large model)900 ms18%
Embedding call80 ms2%
Vector search60 ms1%
Profile lookup (sequential)250 ms5%
Generation, 15 chunks in prompt3,700 ms74%

Every row suggests a different fix. Without the waterfall, teams often optimise the vector search, which here is 1% of the time.

Fixes in order of payoff

  1. Stream — send tokens as they are generated. Total time stays the same, but users see text after the time-to-first-token.
  2. Parallelise — independent steps run concurrently instead of one after another.
  3. Cut prefill — rerank and send 4 or 5 chunks, not 15. A shorter prompt means a faster first token, and often a better answer.
  4. Cache — a semantic cache for repeated questions, plus provider prompt caching on the static system prefix.
  5. Right-size models — a small, fast model for rewriting and routing; the large model only for the final answer.
Python
from langchain_core.runnables import RunnableParallel, RunnablePassthroughchain = (    RunnableParallel(docs=retriever, profile=profile_lookup, question=RunnablePassthrough())    | prompt    | llm)async for chunk in chain.astream(user_question):    await send(chunk.content)

RunnableParallel runs the retriever and the profile lookup at the same time, so the slower one sets the time, not their sum. astream delivers tokens as they arrive.

Caching, with care

set_llm_cache with a semantic cache (for example RedisSemanticCache from the langchain-redis package) returns stored answers for near-identical questions. On support traffic with many repeated questions, hit rates can be meaningful. Include the knowledge-base version in the cache scope so a document update does not serve old answers, and keep the similarity threshold tight.

A budget per stage

Write it down: retrieval 200 ms, rerank 150 ms, first token 800 ms. Alert when a stage breaks its budget, not only when the total does. That makes the next regression easy to find.

A real-life example

Scenario (illustrative numbers). An HR-tech company's policy assistant has p95 of 6.2 s in production. The waterfall shows a large model rewriting every query (1.1 s), a sequential employee-profile lookup (300 ms) and 15 chunks in the prompt, making time-to-first-token 2.4 s.

They switch rewriting to a small model (250 ms), run the profile lookup in parallel with retrieval, rerank to 5 chunks, and stream. p95 total time falls to 3.4 s, and p95 time-to-first-token to 0.9 s, which is what users feel. The answer-quality eval improves slightly too, because fewer distractor chunks reach the model.

Follow-up questions to expect

  • "Does streaming work through a whole chain?" — Yes; astream streams the final model's tokens, and astream_events lets you stream intermediate events such as retrieved sources.
  • "Where does prompt caching help most?" — Long, static prefixes such as system prompts and tool definitions; it reduces both cost and time-to-first-token on repeated calls.
  • "What if the provider itself is slow?" — Measure time-to-first-token per provider and region; consider a fallback model or region for the tail.