RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How do you inspect the contents of a vector store?


What you need to know

A vector store is a database. Before you trust it, you check it the way you would check a new table: how many rows, what they look like, and whether a query returns what you expect.

Step 1: read the raw records

Python
from collections import Counterimport chromadbcol = chromadb.PersistentClient(path="./chroma").get_collection("hr_handbook")print("rows:", col.count())                       # compare with len(chunks)print(col.peek(2)["metadatas"])                   # do the fields look right?print(len(col.get(where={"page": 3})["ids"]))     # one source or page onlyrows = col.get(include=["documents"])dupes = [t for t, n in Counter(rows["documents"]).items() if n > 1]short = [i for i, d in zip(rows["ids"], rows["documents"]) if len(d.strip()) < 50]print("duplicate texts:", len(dupes), "| near-empty chunks:", len(short))

count() tells you if the ingest ran once, twice or not at all. peek() shows the first records with their metadata, so you can see a missing source or a wrong version field. The Counter finds exact duplicate text. Chunks under 50 characters are usually page numbers, headers or "Confidential" footers that match many questions weakly.

Step 2: probe with questions you know the answer to

Python
for doc, dist in vs.similarity_search_with_score("maternity leave weeks", k=5):    print(round(dist, 3), doc.metadata["page"], doc.page_content[:80])

Pick 5 to 10 questions where you know which page holds the answer. For each, check that the right chunk appears, and where. Remember Chroma returns a distance by default, so the smallest number is the best match.

Step 3: check the score spread

Embed two texts you know are related ("parental leave" and "maternity leave") and two you know are not ("parental leave" and "laptop refresh policy"). The related pair must be clearly closer. If every question gets top scores within a narrow band, say all between 0.71 and 0.74 cosine similarity, the model is not telling documents apart on this corpus. That points to the embedding model, or to chunks full of shared boilerplate.

What each finding means

What you seeLikely cause
Count is a multiple of expectedIngest re-run with random IDs
Count is zero or smallWrong path, wrong collection, loader failed silently
Chunks of 10 to 40 charactersHeaders, footers, page numbers not cleaned
source or page missingLoader dropped metadata, or custom code rebuilt documents
Known answer ranked 15thChunking, embedding model, or need for hybrid search

A real-life example

A bank's product-FAQ bot answers "What is the minimum balance for the Salary Plus account?" with the number for a different account. The team first rewrites the prompt twice. Nothing changes.

Then they inspect the store. get(where={"product": "salary_plus"}) returns zero rows. The loader for the new product sheets never wrote the product field, so the filter matched nothing and the retriever fell back to a general search over all accounts. peek() also shows 300 chunks whose entire text is "Terms and conditions apply. Page 2 of 4".

They fix the loader, drop chunks under 50 characters, and re-ingest. The correct minimum balance now appears at rank 1 for all 8 test questions on that product. The prompt was never the problem.

Follow-up questions to expect

  • "How do you do this in production?" — Log the retrieved IDs, scores and metadata on every request, through LangSmith, Langfuse or plain logs. Then you can inspect a real failing request instead of guessing.
  • "How do you find near-duplicates, not only exact ones?" — Compare embeddings: pairs with cosine similarity above about 0.98 are usually the same text with small changes. Or hash the text after normalising whitespace and case.
  • "Can you see the embeddings?" — Yes, with include=["embeddings"]. You rarely read them directly, but a 2-D projection (UMAP or PCA) can show whether topics form separate clusters.