RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is the specific role of a Retriever in LangChain?


What you need to know

The interface is deliberately small. A retriever must accept a string and return Documents. It does not have to use vectors, and it does not have to return scores. The older method get_relevant_documents() is deprecated; use invoke().

Building your own

To write a retriever, subclass BaseRetriever and implement _get_relevant_documents. This one wraps any other retriever and reorders its results with a cross-encoder, a reranking step covered later in this section:

Python
from typing import Anyfrom langchain_core.callbacks import CallbackManagerForRetrieverRunfrom langchain_core.documents import Documentfrom langchain_core.retrievers import BaseRetrieverclass RerankRetriever(BaseRetriever):    base: BaseRetriever          # any retriever: vector, BM25, hybrid...    model: Any                   # e.g. sentence_transformers.CrossEncoder    top_n: int = 3    def _get_relevant_documents(self, query: str, *,                                run_manager: CallbackManagerForRetrieverRun) -> list[Document]:        docs = self.base.invoke(query)        scores = self.model.predict([(query, d.page_content) for d in docs])        ranked = sorted(zip(docs, scores), key=lambda p: p[1], reverse=True)        for d, s in ranked:            d.metadata["rerank_score"] = float(s)        return [d for d, _ in ranked[: self.top_n]]

Because RerankRetriever is itself a retriever, the chain that uses it does not know or care that there is a hybrid search and a reranker inside. Swapping strategies is a one-line change, which makes A/B tests of retrieval easy.

Where it sits in a chain

Text
question -> retriever -> format docs as numbered sources -> prompt -> LLM -> answer

The retriever's output feeds a formatting step; the LLM call comes after. Keeping them separate lets you test and trace retrieval on its own.

In an agent

In agentic RAG, a retriever is usually wrapped as a tool (for example search_hr_policies(query)), so the model can decide when to search and with what query. The retriever is still the same object underneath.

A real-life example

An HR policy assistant started with a plain vector-store retriever. Over three months the team added BM25 for policy codes like "HR-POL-017", then a reranker, then a filter by the employee's country.

Each change was a new retriever object: EnsembleRetriever for hybrid, then the RerankRetriever wrapper around it. The chain code (retriever | format_docs | prompt | llm) never changed. They ran each version against the same 120 labelled questions, and because every step was a Runnable, their tracing tool showed exactly which documents each version returned for each question.

Follow-up questions to expect

  • "How do you get scores from a retriever?" — The interface returns only Documents. Put scores into metadata in a custom retriever, as above, or call the vector store's similarity_search_with_score directly.
  • "Where do hybrid and ensemble retrievers live now?" — In LangChain 1.x, legacy helpers such as EnsembleRetriever, ParentDocumentRetriever and MultiQueryRetriever moved to the langchain_classic package; BM25Retriever is in langchain_community.retrievers.
  • "Is a retriever the same as a vector store?" — No. A vector store stores and searches vectors; a retriever is the read interface, and it may not use a vector store at all.