Course Content
RAG Systems
12 sections · 66 lessons
What is the step-by-step workflow to embed and store documents?
What you need to know
Indexing is the offline half of RAG. Everything the retriever can ever find is decided here. A mistake at this stage cannot be fixed by a better prompt.
- Load — read each file into
Documentobjects, keepingsourceandpagein metadata so you can cite and filter later. - Split — cut documents into chunks that hold one idea each, with some overlap so a sentence is not cut from its context.
- Enrich metadata — add fields such as
doc_type,version,tenant_id,allowed_groups. You cannot filter on a field you did not store. - Embed — turn each chunk into a vector, in batches, with exactly the model used at query time.
- Persist — write to disk or a server, with IDs you can recompute.
- Verify — check the count, then run known questions and read what comes back.
The code
This runs end to end (tested with LangChain 1.x and Chroma 1.x). Swap the fake embedding for your real model.
1import hashlib2from pypdf import PdfReader3from langchain_core.documents import Document4from langchain_core.embeddings import DeterministicFakeEmbedding5from langchain_text_splitters import RecursiveCharacterTextSplitter6from langchain_chroma import Chroma78embeddings = DeterministicFakeEmbedding(size=384) # use your real model here910docs = [Document(page_content=p.extract_text(),11 metadata={"source": "handbook.pdf", "page": i + 1})12 for i, p in enumerate(PdfReader("handbook.pdf").pages)]1314splitter = RecursiveCharacterTextSplitter(15 chunk_size=1000, chunk_overlap=150, add_start_index=True)16chunks = splitter.split_documents(docs)1718for c in chunks:19 c.metadata.update({"doc_type": "leave_policy", "version": "2026-04"})20ids = [hashlib.sha256(f"{c.metadata['source']}|{c.metadata['page']}|"21 f"{c.metadata['start_index']}".encode()).hexdigest()[:16] for c in chunks]2223vs = Chroma(collection_name="hr_handbook", embedding_function=embeddings,24 persist_directory="./chroma")25vs.add_documents(chunks, ids=ids) # same ids overwrite, so re-runs are safe2627assert vs._collection.count() == len(chunks), "index count mismatch"What each part does:
- Loading.
pypdfreads one page at a time, and we keep the page number for citations. LangChain'sPyPDFLoaderdoes the same, but it lives inlangchain-community, which LangChain is now winding down in favour of standalone packages, so reading the PDF directly is the safer choice. - Splitting.
chunk_size=1000characters is roughly 200 to 250 English words.chunk_overlap=150repeats the end of one chunk at the start of the next.add_start_index=Truerecords where each chunk begins, which gives us a stable part of the ID. - IDs. The ID is a hash of source, page and start position. Run the script twice and the second run overwrites the same 21 records instead of adding 21 more. I ran it twice on a 6-page test PDF; the count stayed at 21.
- Verify. The
assertstops a broken index from going live.
Choices that change the answer
- Tables and scanned PDFs need a layout-aware parser or OCR. Plain text extraction turns a table into a jumble of numbers.
- Many teams now add contextual retrieval: before embedding, prepend a one-line summary of where the chunk sits ("From the 2026 Leave Policy, section on carry-over:"). A chunk that says "up to 10 days" becomes findable by a question about carry-over.
- For large corpora, embed in batches of a few hundred texts per API call and retry failed batches, not the whole job.
A real-life example
A hospital builds a search tool over 1,200 clinical guidelines, many of them scanned PDFs. The first index has 18,000 chunks, but 3,100 of them are empty strings, because the plain loader cannot read scanned pages. Nobody notices for a week, because the verify step only checked the count.
They add two checks to step 6: no chunk under 50 characters, and 20 known questions whose answer page must appear in the top 5. The empty-chunk check fails at once and points them to the scanned files. After adding OCR for those files, the 20-question check passes 18 of 20. The two misses are dosage tables, which lead them to a table-aware parser.
Follow-up questions to expect
- "Why store the start index?" — It gives a stable ID for upserts, lets you highlight the exact passage in the source, and helps you find neighbouring chunks.
- "What happens if you embed with one model and query with another?" — Different models place text in different vector spaces, so distances are meaningless. You get an error if the dimensions differ and silent garbage if they match.
- "How do you handle a document that changes?" — Delete all chunks with that
source, then insert the new ones. Covered in detail in the section on large-scale updates.