Course Content
LangChain Mastery
7 sections · 109 lessons
How do you implement a chain with multiple prompts in LangChain?
What you need to know
Splitting a job into several prompts makes each prompt simpler and each step testable. The cost is more calls.
Dependent prompts: a sequence
1from langchain_core.runnables import RunnablePassthrough23extract = extract_prompt | llm.with_structured_output(LeadInfo)4score = score_prompt | llm | StrOutputParser() # uses {info} and {enquiry}5reply = reply_prompt | llm | StrOutputParser() # uses {score} and {enquiry}67chain = (RunnablePassthrough.assign(info=extract)8 | RunnablePassthrough.assign(score=score)9 | RunnablePassthrough.assign(reply=reply))Each step adds a key. Latency is the sum of the three calls.
Independent prompts: parallel, then merge
1from langchain_core.runnables import RunnableParallel23analysis = RunnableParallel(4 summary=summary_prompt | llm | StrOutputParser(),5 risks=risk_prompt | llm | StrOutputParser(),6 original=lambda x: x["contract"],7)8chain = analysis | final_prompt | llm | StrOutputParser() # uses {summary} {risks} {original}summary and risks both read the same input and run at the same time. Latency is the slower of the two, plus the final call.
Sequence
- Step B needs step A's result
- Latency adds up
- Errors in A flow into B
Parallel
- Steps only need the original input
- Latency is the slowest branch
- Cost is the same as sequence
How many prompts?
- One prompt is best when the task is simple and a single instruction does it well.
- Several prompts help when one prompt tries to do too much and quality drops, when steps need different models (a cheap one for classification, a strong one for writing), or when you need to check an intermediate result.
- Mixing models is easy: each step can have its own
llm.
A real-life example
A contract-review tool for a Gurugram legal team first used one long prompt: "summarise, list risks, and suggest edits". Answers often skipped the risks section on long contracts. The team split it into three prompts: summary and risk-listing run in parallel, then a third prompt writes suggested edits using both. On 50 test contracts, missed risk clauses dropped from 14 to 4. Latency went from 11 seconds to 13 because of the extra final call, which the lawyers accepted. They used a cheaper model for the summary and the stronger one for risks and edits, so total cost rose only about 20%.
Follow-up questions to expect
- "How do you pass the original input to the final prompt?" — Include it as a branch in the parallel dict (as
originalabove) or useassign, which keeps all keys. - "Can each step use a different model?" — Yes; each sub-chain has its own model object.
- "How do you know splitting helped?" — Compare the single-prompt and multi-prompt versions on the same test set for quality, latency and cost.