LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you optimize LangChain chains for low-latency NLP tasks?


Seconds per stage for a typical message1.51.06.08.500.61.92.5rewriteretrieve + lookupanswertotalBeforeAfterAnswer length fell from 350 to 110 output tokens.
The biggest saving came from generating fewer tokens, not from faster retrieval, because output tokens are produced one at a time.

What you need to know

Rough rule: total model time is about the time to read the input plus a fixed time for each output token. A 300-token reply at 20 milliseconds per token is 6 seconds of generation, whatever else you do.

The levers, biggest first

LeverHowWhat it saves
Streamchain.stream() / astream()Perceived wait: first words in under a second
ParalleliseIndependent steps in RunnableParallelSum of steps becomes the slowest step
Fewer output tokensLower max_tokens, ask for short answers, structured outputGeneration time, linearly
Smaller inputFewer retrieved chunks, trimmed history, fewer examplesInput processing time and cost
Smaller model for easy casesRoute by difficultyOften 2 to 5 times faster per call
Cacheset_llm_cache for exact repeats; semantic cache for near repeatsThe whole call on a hit
Prompt cachingKeep the long, fixed part of the prompt first and identicalInput processing on repeat calls
Asyncainvoke, astream under FastAPIServer capacity under load

Streaming and parallel steps in code

Python
from langchain_core.runnables import RunnableParallelprep = RunnableParallel(    docs=itemgetter("question") | retriever,        # about 300 ms    profile=itemgetter("user_id") | fetch_profile,  # about 250 ms    question=itemgetter("question"),)chain = prep | build_prompt | llm | StrOutputParser()async for token in chain.astream({"question": q, "user_id": uid}):    await websocket.send_text(token)

Retrieval and the profile lookup now overlap, and the answer streams to the client as it is generated.

Remove calls you don't need

A second model call to "rephrase the question" or "check the answer" can double latency. Keep it only if your evaluation shows it earns its place.

A real-life example

A food-delivery app's support bot had a p95 latency of 9 seconds. The trace showed four sequential steps: a question-rewrite call (1.5 s), retrieval (0.4 s), an order lookup (0.6 s), and the answer (6 s, averaging 350 output tokens). The team dropped the rewrite call for messages longer than 8 words, ran retrieval and the order lookup in parallel, told the model to answer in at most 3 sentences (average output fell to 110 tokens), and streamed the reply. p95 total time fell to 3.8 seconds, and time to first token to 0.9 seconds. Answer quality on their 300-question test set did not change.

Follow-up questions to expect

  • "Does streaming make the chain faster?" — Not in total time; it makes the wait feel shorter, which is usually what matters for chat.
  • "What is prompt caching?" — Providers can reuse the processed form of a long, identical prompt prefix across calls, which cuts input latency and cost; keep the fixed system prompt and documents at the start.
  • "How do you choose a smaller model safely?" — Route by a cheap classifier and compare quality on your test set per route.