RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What are common challenges when working with the Chroma vector store?


The nightly HR ingest after 14 runsRandom ids with add()• Night 1: 210 rows• Night 14: 2,940 rows• Top 4 = one paragraph, four times• Carry-over rule pushed to rank 5Hash ids with upsert()• Night 1: 210 rows• Night 14: 210 rows• Top 4 = four different passages• Carry-over rule reaches the prompt
Duplicates do not just waste space — they push the second-best evidence out of the few slots the model sees.

What you need to know

1. Persistence: which client did you create?

  • chromadb.Client() keeps everything in memory. The index is gone when the process stops.
  • chromadb.PersistentClient(path="./chroma") writes to disk. Every process must use the same path.
  • chromadb.HttpClient(host=...) talks to a Chroma server. Use it when several app instances share one index.

A common bug is that the ingest script writes to ./chroma and the API server reads from /app/chroma. The collection exists but is empty, and every question returns nothing.

2. Duplicates from random IDs

Every record needs an ID. If the ingest job makes a new uuid4() for each chunk, running the job three times stores every chunk three times. I tested this: two chunks added three times gave count() == 6. With a hash of source + page + start_index as the ID and upsert(), three runs still gave 2.

Duplicates are not just waste. With k=4, three slots can hold the same paragraph, so the answer never sees the second-best evidence.

3. One collection, one embedding model

A collection remembers its dimension, which is the length of each vector. Querying a 384-dimension collection with a 1,024-dimension model fails with a clear error. The worse case is two different models with the same dimension: there is no error, and the neighbours are nonsense. Store the model name in the collection metadata and check it at start-up.

4. Distance, not similarity

A new collection uses L2 (squared Euclidean) distance unless you choose another metric. You set it when you create the collection, for example configuration={"hnsw": {"space": "cosine"}}, and you cannot change it later. LangChain's similarity_search_with_score returns this distance, so lower means closer. People often set a threshold like "keep scores above 0.8" and throw away their best results.

5. Metadata shape

Values must be strings, numbers, booleans or (in Chroma 1.x) lists of these. Nested dictionaries are rejected. Flatten {"acl": {"group": "hr"}} into acl_group: "hr" before writing.

6. Scale and operations

Chroma is excellent for prototypes and single-node services. When you need many millions of vectors, replication, backups you trust, or strict multi-tenant isolation, teams usually move to a managed Chroma Cloud, pgvector, Qdrant, Weaviate or OpenSearch. That is an operations decision, not a quality one.

A real-life example

An HR policy assistant for a 4,000-person company indexes a 60-page leave handbook. A cron job re-ingests it every night. After two weeks, employees complain that answers about carry-over leave are "vague".

The engineer opens the store and runs collection.count(): 2,940 rows. The handbook splits into 210 chunks, and 210 × 14 nights = 2,940. Every chunk exists 14 times. For the question "how many leave days carry over?", all four retrieved slots are copies of the same paragraph about annual leave. The paragraph about the 10-day carry-over limit is ranked fifth and never reaches the model.

The fix takes an hour: delete the collection, switch to hash-based IDs with upsert, and add assert collection.count() == len(chunks) at the end of the job. The next night the count stays at 210.

Follow-up questions to expect

  • "When would you move off Chroma?" — When I need replication, high availability, strong tenant isolation, or tens of millions of vectors with heavy filtering. The code change is small if retrieval sits behind a LangChain VectorStore or my own interface.
  • "How do you change the embedding model?" — Build a new collection with the new model, run the evaluation set against it, then switch the app over. Never mix two models in one collection.
  • "Why do my scores look backwards?" — Chroma returns distance by default. Use similarity_search_with_relevance_scores for a 0 to 1 score, or read the distance with "lower is better".