LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to create a simple LangChain LLMChain for text summarization.


Map-reduce for a judgment longer than the context window150-pagejudgmentSplit into8,000-characterchunksSummariseevery chunkin parallelJoin thepartialsummariesSummarise thesummaries once moreOnly about 5% of orders need this path; the rest fit in one call.
Check whether the text simply fits before splitting — one call is cheaper and keeps the links between distant pages.

What you need to know

LLMChain was the original unit of LangChain: one prompt plus one model. It returned a dict like {"text": "..."} and did not stream properly. LCEL replaced it with a pipe of Runnables that does the same job with less code.

Python
from langchain.chat_models import init_chat_modelfrom langchain_core.prompts import ChatPromptTemplatefrom langchain_core.output_parsers import StrOutputParserdef build_summarizer(model: str = "openai:gpt-5.4-mini", bullets: int = 3):    prompt = ChatPromptTemplate.from_messages([        ("system", "Summarise the text in {n} bullet points. "                   "Use only facts that appear in the text."),        ("human", "{text}"),    ]).partial(n=str(bullets))    llm = init_chat_model(model, temperature=0, timeout=30)    return prompt | llm | StrOutputParser()summarize = build_summarizer()print(summarize.invoke({"text": article}))

The factory function builds the chain once; you call invoke on it as many times as you like. partial fixes the bullet count at build time, so callers only pass text. The parser turns the AIMessage into a plain string.

The legacy version, for comparison

Python
from langchain_classic.chains import LLMChain       # was: from langchain.chainsfrom langchain_core.prompts import PromptTemplatechain = LLMChain(llm=llm, prompt=PromptTemplate.from_template("Summarise: {text}"))chain.invoke({"text": article})          # -> {"text": "...", ...}

It still runs with langchain-classic installed, but it raises a deprecation warning and is scheduled for removal in 2.0.

When the text is longer than the context window

Python
from langchain_text_splitters import RecursiveCharacterTextSplitterdef summarize_long(text: str, summarize, chunk_size: int = 8000) -> str:    chunks = RecursiveCharacterTextSplitter(        chunk_size=chunk_size, chunk_overlap=200).split_text(text)    if len(chunks) == 1:        return summarize.invoke({"text": chunks[0]})    partials = summarize.batch([{"text": c} for c in chunks],                               config={"max_concurrency": 5})    return summarize.invoke({"text": "\n\n".join(partials)})

This is map-reduce: summarise each chunk in parallel (map), then summarise the partial summaries (reduce). Many current models have very large context windows, so first check whether the text simply fits; one call is cheaper and keeps cross-chunk context.

A real-life example

A legal-tech startup in Delhi summarises court orders for lawyers. Most orders are 3 to 10 pages and fit in one call. About 5% are 150-page judgments. The team used build_summarizer for normal orders and summarize_long for large ones, switching on a token count done before the call. They also added a guard: if the extracted text is shorter than 200 characters (a scanned PDF with no OCR text), the function returns "no text found" instead of calling the model. Before the guard, those empty files produced confident, invented summaries and cost money.

Follow-up questions to expect

  • "How would you stream the summary to the UI?" — Call summarize.stream({...}) and send each chunk; LCEL chains stream through the parser automatically.
  • "How do you stop it inventing facts?" — Temperature 0, an instruction to use only the text, and for high-stakes use a second check that every bullet is supported by the source.
  • "Map-reduce or refine?" — Map-reduce is parallel and fast; refine (update one running summary chunk by chunk) keeps order better but is sequential and slower.