Course Content
RAG Systems
12 sections · 66 lessons
What are common challenges when working with the Chroma vector store?
What you need to know
1. Persistence: which client did you create?
chromadb.Client()keeps everything in memory. The index is gone when the process stops.chromadb.PersistentClient(path="./chroma")writes to disk. Every process must use the same path.chromadb.HttpClient(host=...)talks to a Chroma server. Use it when several app instances share one index.
A common bug is that the ingest script writes to ./chroma and the API server reads from /app/chroma. The collection exists but is empty, and every question returns nothing.
2. Duplicates from random IDs
Every record needs an ID. If the ingest job makes a new uuid4() for each chunk, running the job three times stores every chunk three times. I tested this: two chunks added three times gave count() == 6. With a hash of source + page + start_index as the ID and upsert(), three runs still gave 2.
Duplicates are not just waste. With k=4, three slots can hold the same paragraph, so the answer never sees the second-best evidence.
3. One collection, one embedding model
A collection remembers its dimension, which is the length of each vector. Querying a 384-dimension collection with a 1,024-dimension model fails with a clear error. The worse case is two different models with the same dimension: there is no error, and the neighbours are nonsense. Store the model name in the collection metadata and check it at start-up.
4. Distance, not similarity
A new collection uses L2 (squared Euclidean) distance unless you choose another metric. You set it when you create the collection, for example configuration={"hnsw": {"space": "cosine"}}, and you cannot change it later. LangChain's similarity_search_with_score returns this distance, so lower means closer. People often set a threshold like "keep scores above 0.8" and throw away their best results.
5. Metadata shape
Values must be strings, numbers, booleans or (in Chroma 1.x) lists of these. Nested dictionaries are rejected. Flatten {"acl": {"group": "hr"}} into acl_group: "hr" before writing.
6. Scale and operations
Chroma is excellent for prototypes and single-node services. When you need many millions of vectors, replication, backups you trust, or strict multi-tenant isolation, teams usually move to a managed Chroma Cloud, pgvector, Qdrant, Weaviate or OpenSearch. That is an operations decision, not a quality one.
A real-life example
An HR policy assistant for a 4,000-person company indexes a 60-page leave handbook. A cron job re-ingests it every night. After two weeks, employees complain that answers about carry-over leave are "vague".
The engineer opens the store and runs collection.count(): 2,940 rows. The handbook splits into 210 chunks, and 210 × 14 nights = 2,940. Every chunk exists 14 times. For the question "how many leave days carry over?", all four retrieved slots are copies of the same paragraph about annual leave. The paragraph about the 10-day carry-over limit is ranked fifth and never reaches the model.
The fix takes an hour: delete the collection, switch to hash-based IDs with upsert, and add assert collection.count() == len(chunks) at the end of the job. The next night the count stays at 210.
Follow-up questions to expect
- "When would you move off Chroma?" — When I need replication, high availability, strong tenant isolation, or tens of millions of vectors with heavy filtering. The code change is small if retrieval sits behind a LangChain
VectorStoreor my own interface. - "How do you change the embedding model?" — Build a new collection with the new model, run the evaluation set against it, then switch the app over. Never mix two models in one collection.
- "Why do my scores look backwards?" — Chroma returns distance by default. Use
similarity_search_with_relevance_scoresfor a 0 to 1 score, or read the distance with "lower is better".