LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you implement a chain with multiple prompts in LangChain?


What you need to know

Splitting a job into several prompts makes each prompt simpler and each step testable. The cost is more calls.

Dependent prompts: a sequence

Python
from langchain_core.runnables import RunnablePassthroughextract = extract_prompt | llm.with_structured_output(LeadInfo)score   = score_prompt   | llm | StrOutputParser()     # uses {info} and {enquiry}reply   = reply_prompt   | llm | StrOutputParser()     # uses {score} and {enquiry}chain = (RunnablePassthrough.assign(info=extract)         | RunnablePassthrough.assign(score=score)         | RunnablePassthrough.assign(reply=reply))

Each step adds a key. Latency is the sum of the three calls.

Independent prompts: parallel, then merge

Python
from langchain_core.runnables import RunnableParallelanalysis = RunnableParallel(    summary=summary_prompt | llm | StrOutputParser(),    risks=risk_prompt | llm | StrOutputParser(),    original=lambda x: x["contract"],)chain = analysis | final_prompt | llm | StrOutputParser()   # uses {summary} {risks} {original}

summary and risks both read the same input and run at the same time. Latency is the slower of the two, plus the final call.

Sequence

  • Step B needs step A's result
  • Latency adds up
  • Errors in A flow into B

Parallel

  • Steps only need the original input
  • Latency is the slowest branch
  • Cost is the same as sequence

How many prompts?

  • One prompt is best when the task is simple and a single instruction does it well.
  • Several prompts help when one prompt tries to do too much and quality drops, when steps need different models (a cheap one for classification, a strong one for writing), or when you need to check an intermediate result.
  • Mixing models is easy: each step can have its own llm.

A real-life example

A contract-review tool for a Gurugram legal team first used one long prompt: "summarise, list risks, and suggest edits". Answers often skipped the risks section on long contracts. The team split it into three prompts: summary and risk-listing run in parallel, then a third prompt writes suggested edits using both. On 50 test contracts, missed risk clauses dropped from 14 to 4. Latency went from 11 seconds to 13 because of the extra final call, which the lawyers accepted. They used a cheaper model for the summary and the stronger one for risks and edits, so total cost rose only about 20%.

Follow-up questions to expect

  • "How do you pass the original input to the final prompt?" — Include it as a branch in the parallel dict (as original above) or use assign, which keeps all keys.
  • "Can each step use a different model?" — Yes; each sub-chain has its own model object.
  • "How do you know splitting helped?" — Compare the single-prompt and multi-prompt versions on the same test set for quality, latency and cost.