LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to implement a parallel chain execution in LangChain.


Parallel changes the wait, not the bill4.5 s3 calls1.8 s3 callswall time per reviewpaid model callsOne after anotherRunnableParallel
Wall time drops to the slowest branch while cost stays fixed, and the request rate triples.

What you need to know

Python
from langchain_core.runnables import RunnableParallelfrom langchain_core.output_parsers import StrOutputParser, JsonOutputParserdef build_review_analysis(llm):    return RunnableParallel(        summary=summary_prompt | llm | StrOutputParser(),        sentiment=sentiment_prompt | llm | StrOutputParser(),        entities=entity_prompt | llm | JsonOutputParser(),    )analysis = build_review_analysis(llm)analysis.invoke({"text": review})# {"summary": "...", "sentiment": "negative", "entities": {...}}

All three prompts use {text}. invoke runs the branches in a thread pool; ainvoke runs them as asyncio tasks, which is the better choice inside an async web server.

The dict shortcut

Python
chain = {"context": retriever, "question": RunnablePassthrough()} | prompt | llm

A dict at the start of or inside a pipeline becomes a RunnableParallel. Here, retrieval and passing the question through run together, and the dict feeds the prompt.

Partial failure

By default, if one branch raises, the whole parallel step raises and you lose the other results. When a missing branch is acceptable:

Python
safe_entities = (entity_prompt | llm | JsonOutputParser()).with_fallbacks(    [RunnableLambda(lambda x: {})])

Now a failed entity extraction returns an empty dict, and the summary and sentiment still come back.

Limits

  • Rate limits: three branches means three times the request rate. With batch on top, concurrency multiplies again.
  • Independence: branches cannot see each other's output. If one needs another's result, that part is a sequence.
  • Threads: sync branches share a thread pool; very large fan-outs are better done async.

A real-life example

An e-commerce marketplace analyses 50,000 product reviews a day for its seller dashboard: a one-line summary, a sentiment label, and product aspects mentioned (delivery, packaging, quality). Run one after another, each review took about 4.5 seconds. With the three branches in a RunnableParallel, it took about 1.8 seconds, the time of the slowest branch. Combined with abatch at a concurrency of 20, the daily job dropped from 9 hours to under 2. The team added an empty-dict fallback to the aspects branch, which failed on about 0.5% of reviews written in mixed scripts, so those reviews still got a summary and sentiment.

Follow-up questions to expect

  • "Does RunnableParallel reduce cost?" — No. It reduces waiting time only; each branch is still a paid call.
  • "How do you stream a parallel step?" — stream yields chunks tagged with their branch key, so you can fill each part of a UI as it arrives.
  • "How do you limit concurrency inside a big parallel batch?" — Set max_concurrency in the config; it applies across the whole run.