RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is the primary purpose of DocumentLoaders?


What you need to know

What a loader is responsible for

  • Fetching the raw content: a file path, a URL, an API with paging, a database query.
  • Turning it into text: PDF text extraction, HTML parsing, speech-to-text for audio.
  • Attaching metadata: source, page number, title, dates, and anything you need for filters and permissions.

It is not responsible for splitting or embedding. Keeping these separate is what lets you test and replace each step.

load() versus lazy_load()

load() returns a full list, so a 50,000-page crawl sits in memory at once. lazy_load() is a generator that yields one Document at a time, so you can stream documents into the splitter and embed in batches. When you write your own loader, you implement lazy_load() and get load() for free:

Python
from typing import Iteratorfrom langchain_core.document_loaders import BaseLoaderfrom langchain_core.documents import Documentclass FaqApiLoader(BaseLoader):    def __init__(self, rows: list[dict]):        self.rows = rows                 # a real loader would page through an API    def lazy_load(self) -> Iterator[Document]:        for row in self.rows:            yield Document(                page_content=f"Q: {row['question']}\nA: {row['answer']}",                metadata={"source": f"faq/{row['id']}", "product": row["product"],                          "updated": row["updated"], "acl": ["public"]},            )docs = FaqApiLoader([{"id": 101, "product": "savings", "updated": "2026-08-02",                      "question": "What is the minimum balance?",                      "answer": "Rs 10,000 average monthly balance."}]).load()

Notice that the loader writes the question and answer together into page_content. A FAQ answer alone ("Rs 10,000 average monthly balance") would be hard to retrieve, because it does not contain the words of the question.

Choosing a PDF loader

PDFs are where loaders differ most. A simple text extractor such as PyPDFLoader is fast and fine for plain text. For tables, multi-column layouts and scanned pages, use a layout-aware parser (Unstructured, Docling, or a vision-model parser) that keeps tables as tables and reads columns in the right order.

In 2026, most loaders still live in langchain-community, which now prints a warning that it is being sunset in favour of standalone integration packages. The Document contract is unchanged; check each integration's docs for its current package.

A real-life example

An HR policy assistant ingests from three places: 140 policy PDFs, an internal FAQ API, and a SharePoint site of country-specific guides.

The first version used a plain PDF loader and a generic web loader. Two problems appeared. The leave-encashment table in one PDF came through as a run of numbers with no column names, and the SharePoint pages lost their "India only" and "UK only" labels, so a UK employee got the Indian gratuity rule.

The fix was in the loaders, not the model. A layout-aware PDF parser kept the table structure. The SharePoint loader was extended to copy the page's country field and its permission groups into metadata. At query time, retrieval filters on the employee's country. The same questions now return the right country's policy.

Follow-up questions to expect

  • "Can a loader run at query time?" — Yes. For live data, a tool can call an API and wrap the result in a Document on the fly instead of indexing it.
  • "How do you handle a source with no built-in loader?" — Subclass BaseLoader, implement lazy_load(), and yield Documents with the metadata your pipeline expects.
  • "What metadata must every loader provide?" — At minimum source; in practice also a stable document ID, last-modified date, and access groups.