Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to create a simple LangChain LLMChain for text summarization.
What you need to know
LLMChain was the original unit of LangChain: one prompt plus one model. It returned a dict like {"text": "..."} and did not stream properly. LCEL replaced it with a pipe of Runnables that does the same job with less code.
1from langchain.chat_models import init_chat_model2from langchain_core.prompts import ChatPromptTemplate3from langchain_core.output_parsers import StrOutputParser45def build_summarizer(model: str = "openai:gpt-5.4-mini", bullets: int = 3):6 prompt = ChatPromptTemplate.from_messages([7 ("system", "Summarise the text in {n} bullet points. "8 "Use only facts that appear in the text."),9 ("human", "{text}"),10 ]).partial(n=str(bullets))11 llm = init_chat_model(model, temperature=0, timeout=30)12 return prompt | llm | StrOutputParser()1314summarize = build_summarizer()15print(summarize.invoke({"text": article}))The factory function builds the chain once; you call invoke on it as many times as you like. partial fixes the bullet count at build time, so callers only pass text. The parser turns the AIMessage into a plain string.
The legacy version, for comparison
1from langchain_classic.chains import LLMChain # was: from langchain.chains2from langchain_core.prompts import PromptTemplate34chain = LLMChain(llm=llm, prompt=PromptTemplate.from_template("Summarise: {text}"))5chain.invoke({"text": article}) # -> {"text": "...", ...}It still runs with langchain-classic installed, but it raises a deprecation warning and is scheduled for removal in 2.0.
When the text is longer than the context window
1from langchain_text_splitters import RecursiveCharacterTextSplitter23def summarize_long(text: str, summarize, chunk_size: int = 8000) -> str:4 chunks = RecursiveCharacterTextSplitter(5 chunk_size=chunk_size, chunk_overlap=200).split_text(text)6 if len(chunks) == 1:7 return summarize.invoke({"text": chunks[0]})8 partials = summarize.batch([{"text": c} for c in chunks],9 config={"max_concurrency": 5})10 return summarize.invoke({"text": "\n\n".join(partials)})This is map-reduce: summarise each chunk in parallel (map), then summarise the partial summaries (reduce). Many current models have very large context windows, so first check whether the text simply fits; one call is cheaper and keeps cross-chunk context.
A real-life example
A legal-tech startup in Delhi summarises court orders for lawyers. Most orders are 3 to 10 pages and fit in one call. About 5% are 150-page judgments. The team used build_summarizer for normal orders and summarize_long for large ones, switching on a token count done before the call. They also added a guard: if the extracted text is shorter than 200 characters (a scanned PDF with no OCR text), the function returns "no text found" instead of calling the model. Before the guard, those empty files produced confident, invented summaries and cost money.
Follow-up questions to expect
- "How would you stream the summary to the UI?" — Call
summarize.stream({...})and send each chunk; LCEL chains stream through the parser automatically. - "How do you stop it inventing facts?" — Temperature 0, an instruction to use only the text, and for high-stakes use a second check that every bullet is supported by the source.
- "Map-reduce or refine?" — Map-reduce is parallel and fast; refine (update one running summary chunk by chunk) keeps order better but is sequential and slower.