RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How does the indexing process work, and why is it essential?


What you need to know

The four steps in code

This runs as shown with a small open embedding model. Swap in any hosted embedding model the same way.

Python
import hashlibfrom langchain_community.document_loaders import PyPDFLoaderfrom langchain_text_splitters import RecursiveCharacterTextSplitterfrom langchain_huggingface import HuggingFaceEmbeddingsfrom langchain_chroma import Chromapages = PyPDFLoader("handbook.pdf").load()          # 1. load: one Document per pagesplitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(    encoding_name="cl100k_base", chunk_size=400, chunk_overlap=50,    add_start_index=True)chunks = splitter.split_documents(pages)             # 2. split, keeping metadatadef chunk_id(c):                                     # same chunk -> same ID every run    key = f"{c.metadata['source']}|{c.metadata['page']}|{c.metadata['start_index']}"    return hashlib.sha1(key.encode()).hexdigest()emb = HuggingFaceEmbeddings(model_name="BAAI/bge-small-en-v1.5",                            encode_kwargs={"normalize_embeddings": True})store = Chroma(collection_name="hr_policies", embedding_function=emb,               persist_directory="./chroma_hr",               collection_metadata={"hnsw:space": "cosine"})store.add_documents(chunks, ids=[chunk_id(c) for c in chunks])   # 3 + 4. embed, store

What each line is doing:

  • PyPDFLoader returns one Document per page with source and page in its metadata. Those fields later become citations.
  • from_tiktoken_encoder measures chunk_size in tokens, not characters, which is what embedding models limit.
  • The chunk ID is a hash of source, page and start offset. Running the script twice leaves the count unchanged instead of doubling it.
  • hnsw:space sets cosine distance. Chroma otherwise defaults to L2 distance, so set it on purpose.

In 2026, langchain-community prints a warning that it is being sunset in favour of standalone integration packages. The loader API is the same, but check each integration's docs for its current package.

Why it is worth doing well

Put numbers on it. 10,000 documents of 20 pages each, at about 650 tokens per page, is 130 million tokens. At about $0.02 per million tokens, a typical price for a small hosted embedding model, embedding the whole corpus costs under $3, once. At query time you embed only the question (about 20 tokens) and search.

The index also turns search into a fast lookup. Comparing a query with 5 million vectors one by one takes far too long for a chat reply; an approximate index such as HNSW answers in milliseconds by checking only a small part of the collection.

Keeping it fresh

Store a content hash per document. On each run, skip documents whose hash has not changed, re-embed the ones that changed, and delete chunks for documents that were removed. Without the delete step, old versions stay searchable forever.

A real-life example

A bank's product-FAQ bot indexes 3,200 product documents every night. In the first version, the job used random UUIDs as chunk IDs. After 30 nightly runs the index held 30 copies of every chunk, and the top 5 results for "home loan prepayment charges" were five copies of the same paragraph.

The fix took an hour: deterministic IDs from source|page|start_index, a content hash to skip unchanged files, and a delete step for removed files. The nightly job went from re-embedding everything to re-embedding about 40 changed documents a night.

Follow-up questions to expect

  • "What metadata do you store per chunk?" — Source, title, section heading, page, start offset, version or effective date, language, document type and access-control groups.
  • "How do you handle a change of embedding model?" — Re-embed everything into a new collection, evaluate it, then switch traffic. Vectors from two models cannot be mixed.
  • "How long does indexing take?" — Parsing and embedding dominate. It scales with corpus size and is easy to parallelise, because each document is independent.