Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to create a LangChain RAG pipeline.
What you need to know
The function
1from langchain_core.prompts import ChatPromptTemplate2from langchain_core.output_parsers import StrOutputParser3from langchain_core.runnables import RunnableParallel, RunnablePassthrough4from langchain_core.vectorstores import InMemoryVectorStore5from langchain_text_splitters import RecursiveCharacterTextSplitter67PROMPT = ChatPromptTemplate.from_messages([8 ("system", "Answer only from the context below. Cite sources as [source]. "9 "If the answer is not in the context, say you don't know.\n\n{context}"),10 ("human", "{question}"),11])1213def format_docs(docs):14 return "\n\n".join(f"[{d.metadata.get('source', '?')}] {d.page_content}" for d in docs)1516def build_rag(docs, llm, embeddings, k: int = 4):17 chunks = RecursiveCharacterTextSplitter(18 chunk_size=1000, chunk_overlap=150).split_documents(docs)19 store = InMemoryVectorStore.from_documents(chunks, embeddings)20 retriever = store.as_retriever(search_kwargs={"k": k})2122 answer = (lambda x: {"context": format_docs(x["context"]), "question": x["question"]}) \23 | PROMPT | llm | StrOutputParser()24 return RunnableParallel(context=retriever, question=RunnablePassthrough()).assign(answer=answer)2526rag = build_rag(docs, llm, embeddings)27out = rag.invoke("What is the SLA for P1 tickets?")28# out = {"context": [Document, ...], "question": "...", "answer": "4 hours [sla.md]"}Walking through it
- Splitter — 1,000-character chunks with 150 overlap is a reasonable default; tune it with evaluation.
- Store —
InMemoryVectorStorekeeps the example self-contained. In production, pass in a retriever for Chroma, pgvector or Qdrant instead of building the index inside the request path. RunnableParallel— runs the retriever and passes the raw question through at the same time, producing{"context": [...], "question": "..."}..assign(answer=...)— adds ananswerkey while keepingcontext. That is how you return sources.- The prompt — the "answer only from the context" and "say you don't know" lines are the grounding rules. Without them the model fills gaps from its training data.
- Streaming —
rag.stream(q)first yields the context, then the answer token by token, so a UI can show sources immediately.
Better shape for production
Split the function in two: build_index(docs, embeddings) runs offline as a job; build_rag_chain(retriever, llm) runs in the service. The service then starts in seconds and tests can pass a fake retriever.
Older helpers you may be asked about
create_retrieval_chain(retriever, create_stuff_documents_chain(llm, prompt)) does the same thing and returns {"input", "context", "answer"}. In LangChain 1.x it is imported from langchain_classic.chains.
The agent version
If the model should decide when to search, wrap the retriever as a tool:
1from langchain.tools import tool2from langchain.agents import create_agent34@tool(response_format="content_and_artifact")5def search_policies(query: str):6 """Search the HR handbook for leave, notice-period and benefits rules."""7 found = store.similarity_search(query, k=4)8 return format_docs(found), found # text for the model, docs kept as artifact910agent = create_agent(llm, tools=[search_policies],11 system_prompt="Use search_policies before answering policy questions.")content_and_artifact sends the formatted text to the model but keeps the raw Documents on the ToolMessage.artifact, so you can still show citations.
A real-life example
An IT helpdesk bot for a 2,000-employee company answers "What is the SLA for a laptop replacement?". Built with build_rag, the first version answered in about 2.5 seconds with a citation to it-sla.md.
In code review, the reviewer asked for the split: indexing 1,400 pages at startup took 90 seconds and cost embedding calls on every deploy. After the change, a nightly job builds the pgvector index and the service only builds the chain. Unit tests pass a FakeListChatModel and a lambda retriever, and run in under a second.
Follow-up questions to expect
- "How would you add chat history?" — Add a
MessagesPlaceholder("history")to the prompt, and rewrite the follow-up question into a standalone one before retrieval. - "How do you stop hallucination here?" — The grounded prompt, a score check before generating, citations, and a groundedness check in evaluation.
- "Why return the context?" — For citations in the UI, and to log chunk ids so a bad answer can be traced to its retrieval.