Course Content
LangChain Mastery
7 sections · 109 lessons
How do you optimize LangChain chains for low-latency NLP tasks?
What you need to know
Rough rule: total model time is about the time to read the input plus a fixed time for each output token. A 300-token reply at 20 milliseconds per token is 6 seconds of generation, whatever else you do.
The levers, biggest first
| Lever | How | What it saves |
|---|---|---|
| Stream | chain.stream() / astream() | Perceived wait: first words in under a second |
| Parallelise | Independent steps in RunnableParallel | Sum of steps becomes the slowest step |
| Fewer output tokens | Lower max_tokens, ask for short answers, structured output | Generation time, linearly |
| Smaller input | Fewer retrieved chunks, trimmed history, fewer examples | Input processing time and cost |
| Smaller model for easy cases | Route by difficulty | Often 2 to 5 times faster per call |
| Cache | set_llm_cache for exact repeats; semantic cache for near repeats | The whole call on a hit |
| Prompt caching | Keep the long, fixed part of the prompt first and identical | Input processing on repeat calls |
| Async | ainvoke, astream under FastAPI | Server capacity under load |
Streaming and parallel steps in code
1from langchain_core.runnables import RunnableParallel23prep = RunnableParallel(4 docs=itemgetter("question") | retriever, # about 300 ms5 profile=itemgetter("user_id") | fetch_profile, # about 250 ms6 question=itemgetter("question"),7)8chain = prep | build_prompt | llm | StrOutputParser()910async for token in chain.astream({"question": q, "user_id": uid}):11 await websocket.send_text(token)Retrieval and the profile lookup now overlap, and the answer streams to the client as it is generated.
Remove calls you don't need
A second model call to "rephrase the question" or "check the answer" can double latency. Keep it only if your evaluation shows it earns its place.
A real-life example
A food-delivery app's support bot had a p95 latency of 9 seconds. The trace showed four sequential steps: a question-rewrite call (1.5 s), retrieval (0.4 s), an order lookup (0.6 s), and the answer (6 s, averaging 350 output tokens). The team dropped the rewrite call for messages longer than 8 words, ran retrieval and the order lookup in parallel, told the model to answer in at most 3 sentences (average output fell to 110 tokens), and streamed the reply. p95 total time fell to 3.8 seconds, and time to first token to 0.9 seconds. Answer quality on their 300-question test set did not change.
Follow-up questions to expect
- "Does streaming make the chain faster?" — Not in total time; it makes the wait feel shorter, which is usually what matters for chat.
- "What is prompt caching?" — Providers can reuse the processed form of a long, identical prompt prefix across calls, which cuts input latency and cost; keep the fixed system prompt and documents at the start.
- "How do you choose a smaller model safely?" — Route by a cheap classifier and compare quality on your test set per route.