Course Content
RAG Systems
12 sections · 66 lessons
How does the indexing process work, and why is it essential?
What you need to know
The four steps in code
This runs as shown with a small open embedding model. Swap in any hosted embedding model the same way.
1import hashlib2from langchain_community.document_loaders import PyPDFLoader3from langchain_text_splitters import RecursiveCharacterTextSplitter4from langchain_huggingface import HuggingFaceEmbeddings5from langchain_chroma import Chroma67pages = PyPDFLoader("handbook.pdf").load() # 1. load: one Document per page89splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(10 encoding_name="cl100k_base", chunk_size=400, chunk_overlap=50,11 add_start_index=True)12chunks = splitter.split_documents(pages) # 2. split, keeping metadata1314def chunk_id(c): # same chunk -> same ID every run15 key = f"{c.metadata['source']}|{c.metadata['page']}|{c.metadata['start_index']}"16 return hashlib.sha1(key.encode()).hexdigest()1718emb = HuggingFaceEmbeddings(model_name="BAAI/bge-small-en-v1.5",19 encode_kwargs={"normalize_embeddings": True})20store = Chroma(collection_name="hr_policies", embedding_function=emb,21 persist_directory="./chroma_hr",22 collection_metadata={"hnsw:space": "cosine"})23store.add_documents(chunks, ids=[chunk_id(c) for c in chunks]) # 3 + 4. embed, storeWhat each line is doing:
PyPDFLoaderreturns oneDocumentper page withsourceandpagein its metadata. Those fields later become citations.from_tiktoken_encodermeasureschunk_sizein tokens, not characters, which is what embedding models limit.- The chunk ID is a hash of source, page and start offset. Running the script twice leaves the count unchanged instead of doubling it.
hnsw:spacesets cosine distance. Chroma otherwise defaults to L2 distance, so set it on purpose.
In 2026, langchain-community prints a warning that it is being sunset in favour of standalone integration packages. The loader API is the same, but check each integration's docs for its current package.
Why it is worth doing well
Put numbers on it. 10,000 documents of 20 pages each, at about 650 tokens per page, is 130 million tokens. At about $0.02 per million tokens, a typical price for a small hosted embedding model, embedding the whole corpus costs under $3, once. At query time you embed only the question (about 20 tokens) and search.
The index also turns search into a fast lookup. Comparing a query with 5 million vectors one by one takes far too long for a chat reply; an approximate index such as HNSW answers in milliseconds by checking only a small part of the collection.
Keeping it fresh
Store a content hash per document. On each run, skip documents whose hash has not changed, re-embed the ones that changed, and delete chunks for documents that were removed. Without the delete step, old versions stay searchable forever.
A real-life example
A bank's product-FAQ bot indexes 3,200 product documents every night. In the first version, the job used random UUIDs as chunk IDs. After 30 nightly runs the index held 30 copies of every chunk, and the top 5 results for "home loan prepayment charges" were five copies of the same paragraph.
The fix took an hour: deterministic IDs from source|page|start_index, a content hash to skip unchanged files, and a delete step for removed files. The nightly job went from re-embedding everything to re-embedding about 40 changed documents a night.
Follow-up questions to expect
- "What metadata do you store per chunk?" — Source, title, section heading, page, start offset, version or effective date, language, document type and access-control groups.
- "How do you handle a change of embedding model?" — Re-embed everything into a new collection, evaluate it, then switch traffic. Vectors from two models cannot be mixed.
- "How long does indexing take?" — Parsing and embedding dominate. It scales with corpus size and is easy to parallelise, because each document is independent.