LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to create a LangChain RAG pipeline.


What you need to know

The function

Python
from langchain_core.prompts import ChatPromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_core.runnables import RunnableParallel, RunnablePassthroughfrom langchain_core.vectorstores import InMemoryVectorStorefrom langchain_text_splitters import RecursiveCharacterTextSplitterPROMPT = ChatPromptTemplate.from_messages([    ("system", "Answer only from the context below. Cite sources as [source]. "               "If the answer is not in the context, say you don't know.\n\n{context}"),    ("human", "{question}"),])def format_docs(docs):    return "\n\n".join(f"[{d.metadata.get('source', '?')}] {d.page_content}" for d in docs)def build_rag(docs, llm, embeddings, k: int = 4):    chunks = RecursiveCharacterTextSplitter(        chunk_size=1000, chunk_overlap=150).split_documents(docs)    store = InMemoryVectorStore.from_documents(chunks, embeddings)    retriever = store.as_retriever(search_kwargs={"k": k})    answer = (lambda x: {"context": format_docs(x["context"]), "question": x["question"]}) \        | PROMPT | llm | StrOutputParser()    return RunnableParallel(context=retriever, question=RunnablePassthrough()).assign(answer=answer)rag = build_rag(docs, llm, embeddings)out = rag.invoke("What is the SLA for P1 tickets?")# out = {"context": [Document, ...], "question": "...", "answer": "4 hours [sla.md]"}

Walking through it

  • Splitter — 1,000-character chunks with 150 overlap is a reasonable default; tune it with evaluation.
  • Store — InMemoryVectorStore keeps the example self-contained. In production, pass in a retriever for Chroma, pgvector or Qdrant instead of building the index inside the request path.
  • RunnableParallel — runs the retriever and passes the raw question through at the same time, producing {"context": [...], "question": "..."}.
  • .assign(answer=...) — adds an answer key while keeping context. That is how you return sources.
  • The prompt — the "answer only from the context" and "say you don't know" lines are the grounding rules. Without them the model fills gaps from its training data.
  • Streaming — rag.stream(q) first yields the context, then the answer token by token, so a UI can show sources immediately.

Better shape for production

Split the function in two: build_index(docs, embeddings) runs offline as a job; build_rag_chain(retriever, llm) runs in the service. The service then starts in seconds and tests can pass a fake retriever.

Older helpers you may be asked about

create_retrieval_chain(retriever, create_stuff_documents_chain(llm, prompt)) does the same thing and returns {"input", "context", "answer"}. In LangChain 1.x it is imported from langchain_classic.chains.

The agent version

If the model should decide when to search, wrap the retriever as a tool:

Python
from langchain.tools import toolfrom langchain.agents import create_agent@tool(response_format="content_and_artifact")def search_policies(query: str):    """Search the HR handbook for leave, notice-period and benefits rules."""    found = store.similarity_search(query, k=4)    return format_docs(found), found      # text for the model, docs kept as artifactagent = create_agent(llm, tools=[search_policies],                     system_prompt="Use search_policies before answering policy questions.")

content_and_artifact sends the formatted text to the model but keeps the raw Documents on the ToolMessage.artifact, so you can still show citations.

A real-life example

An IT helpdesk bot for a 2,000-employee company answers "What is the SLA for a laptop replacement?". Built with build_rag, the first version answered in about 2.5 seconds with a citation to it-sla.md.

In code review, the reviewer asked for the split: indexing 1,400 pages at startup took 90 seconds and cost embedding calls on every deploy. After the change, a nightly job builds the pgvector index and the service only builds the chain. Unit tests pass a FakeListChatModel and a lambda retriever, and run in under a second.

Follow-up questions to expect

  • "How would you add chat history?" — Add a MessagesPlaceholder("history") to the prompt, and rewrite the follow-up question into a standalone one before retrieval.
  • "How do you stop hallucination here?" — The grounded prompt, a score check before generating, citations, and a groundedness check in evaluation.
  • "Why return the context?" — For citations in the UI, and to log chunk ids so a bad answer can be traced to its retrieval.