Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 5: High Latency in Production
Scenario: a LangChain RAG chain that felt fine in development shows a p95 latency of several seconds in production. How do you bring it down?
What you need to know
Development hides latency: one user, a warm cache, short test questions, a small index. Production adds long questions, cold paths, rate limits and concurrency.
Read the waterfall
A LangSmith trace shows each runnable as a span with its duration. A typical slow RAG request looks like this:
| Span | Time | Share |
|---|---|---|
| Query rewrite (large model) | 900 ms | 18% |
| Embedding call | 80 ms | 2% |
| Vector search | 60 ms | 1% |
| Profile lookup (sequential) | 250 ms | 5% |
| Generation, 15 chunks in prompt | 3,700 ms | 74% |
Every row suggests a different fix. Without the waterfall, teams often optimise the vector search, which here is 1% of the time.
Fixes in order of payoff
- Stream — send tokens as they are generated. Total time stays the same, but users see text after the time-to-first-token.
- Parallelise — independent steps run concurrently instead of one after another.
- Cut prefill — rerank and send 4 or 5 chunks, not 15. A shorter prompt means a faster first token, and often a better answer.
- Cache — a semantic cache for repeated questions, plus provider prompt caching on the static system prefix.
- Right-size models — a small, fast model for rewriting and routing; the large model only for the final answer.
1from langchain_core.runnables import RunnableParallel, RunnablePassthrough23chain = (4 RunnableParallel(docs=retriever, profile=profile_lookup, question=RunnablePassthrough())5 | prompt6 | llm7)8async for chunk in chain.astream(user_question):9 await send(chunk.content)RunnableParallel runs the retriever and the profile lookup at the same time, so the slower one sets the time, not their sum. astream delivers tokens as they arrive.
Caching, with care
set_llm_cache with a semantic cache (for example RedisSemanticCache from the langchain-redis package) returns stored answers for near-identical questions. On support traffic with many repeated questions, hit rates can be meaningful. Include the knowledge-base version in the cache scope so a document update does not serve old answers, and keep the similarity threshold tight.
A budget per stage
Write it down: retrieval 200 ms, rerank 150 ms, first token 800 ms. Alert when a stage breaks its budget, not only when the total does. That makes the next regression easy to find.
A real-life example
Scenario (illustrative numbers). An HR-tech company's policy assistant has p95 of 6.2 s in production. The waterfall shows a large model rewriting every query (1.1 s), a sequential employee-profile lookup (300 ms) and 15 chunks in the prompt, making time-to-first-token 2.4 s.
They switch rewriting to a small model (250 ms), run the profile lookup in parallel with retrieval, rerank to 5 chunks, and stream. p95 total time falls to 3.4 s, and p95 time-to-first-token to 0.9 s, which is what users feel. The answer-quality eval improves slightly too, because fewer distractor chunks reach the model.
Follow-up questions to expect
- "Does streaming work through a whole chain?" — Yes;
astreamstreams the final model's tokens, andastream_eventslets you stream intermediate events such as retrieved sources. - "Where does prompt caching help most?" — Long, static prefixes such as system prompts and tool definitions; it reduces both cost and time-to-first-token on repeated calls.
- "What if the provider itself is slow?" — Measure time-to-first-token per provider and region; consider a fallback model or region for the tail.