LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

What is retrieval-augmented generation (RAG) in LangChain?


One offline job, then one lookup per questionLoad filesinto DocumentsSplit into800-characterchunksEmbed andstore in avector storeRetrievetop-k chunksfor the questionGenerate ananswer thatcites themThe first three steps run offline; only the last two run per question.
If step four misses the right chunk, step five still writes a fluent answer — which is why most RAG bugs are retrieval bugs.

What you need to know

A language model only knows what was in its training data, up to a cut-off date. It has never seen your HR policy, your product catalogue or yesterday's price list. You can either retrain the model on that data (fine-tuning) or show the data to the model at the moment you ask. RAG is the second option.

The two halves of the pipeline

  1. Load — read source files (PDFs, web pages, database rows) into Document objects: text plus metadata such as source and page.
  2. Split — cut long documents into chunks of a few hundred to a thousand characters, so each chunk is about one idea.
  3. Embed and store — turn each chunk into a vector with an embedding model and save it in a vector store. Steps 1–3 run offline, as an indexing job.
  4. Retrieve — at question time, embed the question and fetch the top-k closest chunks.
  5. Generate — put those chunks into the prompt and ask the model to answer from them, with citations.

Why teams choose RAG over fine-tuning

  • Freshness — update the index in minutes; retraining takes days.
  • Citations — you know which chunks produced the answer, so users can check.
  • Access control — you filter what each user may retrieve. A fine-tuned model cannot "forget" a document for one user.
  • Cost — no training runs.

The cost is a new failure mode. If retrieval misses the right chunk, the model still writes a fluent answer — from nothing. Most RAG bugs are retrieval bugs.

The current LangChain shape

In LangChain 1.x a simple RAG pipeline is plain LCEL (the | composition from langchain-core):

Python
from langchain_core.prompts import ChatPromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_core.runnables import RunnableParallel, RunnablePassthroughprompt = ChatPromptTemplate.from_messages([    ("system", "Answer only from the context. If it is not there, say you don't know.\n\n{context}"),    ("human", "{question}"),])def format_docs(docs):    return "\n\n".join(f"[{d.metadata['source']}] {d.page_content}" for d in docs)rag = (    RunnableParallel(context=retriever | format_docs, question=RunnablePassthrough())    | prompt | llm | StrOutputParser())rag.invoke("How many casual leaves do I get?")

retriever comes from a vector store (vector_store.as_retriever()). The parallel step runs retrieval and passes the question through; the prompt then receives both.

There are two other forms you will meet:

  • create_retrieval_chain + create_stuff_documents_chain — pre-built helpers that return {"input", "context", "answer"}. Since LangChain 1.0 they live in the langchain-classic package (from langchain_classic.chains import create_retrieval_chain). They still work, but new code usually writes the LCEL above.
  • Agentic RAG — wrap the retriever in a @tool and give it to create_agent. The model decides whether to search, and can search twice with a better query. It is more flexible but slower and less predictable. RetrievalQA is the oldest form and is legacy.

A real-life example

An HR assistant for a 4,000-person company answers questions like "Can I carry forward unused earned leave?". The HR handbook is 180 pages and changes every quarter.

  • Indexing: the handbook is split into about 900 chunks of 800 characters. Each chunk keeps metadata {"source": "handbook-2026-q3.pdf", "page": 42}.
  • Question time: the retriever returns the 4 closest chunks. The prompt tells the model to answer only from them and to cite the page.
  • Result: "Yes, up to 30 days can be carried forward (handbook p. 42)." When HR updates the policy, the team re-indexes the changed pages that evening — no model training.

The same assistant also files leave requests. That part is an agent with an apply_leave tool; the policy lookup becomes one of its tools. This is a common production shape: a fixed RAG chain for questions, an agent when actions are needed.

Follow-up questions to expect

  • "When would you fine-tune instead?" — When you need the model to learn a style, format or skill, not facts. Facts that change belong in retrieval.
  • "What if the answer is not in the documents?" — The prompt must allow "I don't know", and I check retrieval scores; if nothing relevant came back, I say so instead of generating.
  • "Chain or agent for RAG?" — A chain when every question needs one search (fast, predictable). An agent when some questions need no search or several searches.
  • "What does long context change?" — You can stuff more, but cost and latency grow with every token, and models still miss facts buried in the middle. Retrieval keeps the prompt small and relevant.