RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What kinds of data sources can be used in a RAG system?


What you need to know

Each source needs three things: a loader that fetches it, a parser that turns it into clean text, and metadata that travels with every chunk.

Sources and their traps

SourceTypical trap
PDFTables become loose numbers; headers and footers repeat on every page; two-column layouts get interleaved
Scanned PDF or imageNo text layer at all; needs OCR or a vision model
Word, PowerPointSpeaker notes and comments may hold the real content
HTML and portalsNavigation, cookie banners and footers repeat on every page
Confluence, SharePoint, DrivePermissions live in the source system and must be copied
Slack, TeamsMeaning lives in the thread, not the single message
CodeMust be split by function or class, not by character count
Database rowsNeed to be rendered into sentences ("Plan Gold: 50 GB data, 84 days, Rs 799")
Audio and videoSpeech-to-text first, with timestamps as metadata

Index it, or fetch it live?

  • Index content that changes slowly and is searched by meaning: policies, manuals, contracts, tickets.
  • Fetch live anything that changes by the minute or is per user: balances, order status, stock, today's rates. Give the model a tool to call, and never embed this data.

Metadata to capture at ingest

source, title, section, page, last_modified, version, language, doc_type, and the ACL (access-control list, meaning which users or groups may read it). Metadata powers citations ("Loan Policy, page 4"), filters (only current versions), and security (only documents this user may see).

Pages as images

For documents where layout carries meaning, such as slides, forms and scanned brochures, a newer option is to embed the page image directly with a vision-language retriever (ColPali-style models, 2024 onward) instead of extracting text first. It keeps charts and tables that text extraction would destroy, at the cost of larger indexes.

A real-life example

A bank builds its product-FAQ bot from five sources.

  1. Public website FAQ pages (HTML), cleaned to the main content block.
  2. Product terms and conditions (PDF), parsed with a layout-aware parser so the fee tables stay as tables.
  3. Scanned older circulars, run through OCR, with a flag marking them as OCR text.
  4. Call-centre scripts from SharePoint, with their SharePoint groups copied into metadata, so only staff users can retrieve them.
  5. Deposit and loan rates, not indexed. They come from the rate API at query time.

Six weeks after launch, a customer gets an answer from a retired credit-card fee schedule. Because every chunk carries version and effective_to, the fix is a one-line filter (effective_to is null) rather than a rebuild.

Follow-up questions to expect

  • "How do you keep the index in sync with the source?" — Incremental ingestion: poll or subscribe to change events, compare last_modified or a content hash, and re-embed only changed documents; delete chunks of removed documents.
  • "How do you handle tables?" — Parse them into Markdown or HTML tables and keep a table in one chunk, often with a short generated description of what it contains.
  • "Would you index emails?" — Only with the owner's permissions attached, strong PII handling, and a clear retention policy.