RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How does the VectorStoreRetriever function?


What you need to know

Python
retriever = store.as_retriever(    search_type="mmr",    search_kwargs={"k": 5, "fetch_k": 25, "lambda_mult": 0.5,                   "filter": {"country": "IN"}},)docs = retriever.invoke("What notice period do I have to serve?")

The three search types

search_typeWhat it doesKey search_kwargs
"similarity" (default)Top k nearest chunksk, filter
"mmr"Fetches fetch_k, then picks k that are relevant and different from each otherk, fetch_k, lambda_mult, filter
"similarity_score_threshold"Top k, then drops any below the thresholdk, score_threshold, filter

Scores: distance or similarity?

On a Chroma collection set to cosine, with an HR handbook, the query "What notice period do I have to serve?" gave:

Text
similarity_search_with_score            (cosine distance, lower = better)  0.158  hr/handbook-v5.md  Notice period  0.189  hr/handbook-v7.md  Notice period  0.306  hr/handbook-v7.md  Notice buy-outsimilarity_search_with_relevance_scores (0 to 1, higher = better)  0.842  Notice period  0.811  Notice period  0.694  Notice buy-out

Same results, opposite directions. with_relevance_scores converts the store's native score to a 0-to-1 relevance, which is what similarity_score_threshold uses. (Notice also that the old v5 handbook ranked first. That is the metadata-filter problem covered two lessons from now.)

The threshold retriever can return nothing

Python
r = store.as_retriever(search_type="similarity_score_threshold",                       search_kwargs={"k": 3, "score_threshold": 0.75})r.invoke("What notice period do I have to serve?")   # 2 documentsr.invoke("What is the office wifi password?")        # [] plus a warning

An empty list is useful: the chain can reply "I could not find this in the HR policies" instead of giving the model loosely related text. But the right threshold depends on the embedding model, so calibrate it on labelled questions: pick the value that keeps most correct hits and drops most unrelated ones.

It drops scores

invoke returns Documents with no score in metadata. If you need scores for logging or thresholds in your own code, call similarity_search_with_score or similarity_search_with_relevance_scores on the store, or write a small custom retriever that adds them.

A real-life example

A bank's product-FAQ bot answered every question, even "Who won the cricket match yesterday?", by stretching whatever it retrieved. The team switched to similarity_score_threshold.

They picked the threshold from data: on 200 labelled in-scope questions and 100 out-of-scope ones, they plotted how many of each passed at thresholds from 0.5 to 0.85, and chose the point where almost all in-scope questions still passed and most out-of-scope ones did not. Out-of-scope questions now get a polite "I can help with our products and services" reply, and the compliance team stopped seeing invented answers about topics the bank does not cover.

Follow-up questions to expect

  • "What does fetch_k mean?" — The number of candidates MMR considers before choosing k of them. It must be larger than k.
  • "Can I change k per request?" — Yes: pass a new configuration, or call the store directly with a different k.
  • "Is the threshold the same across stores?" — No. It depends on the embedding model and the store's score conversion; recalibrate whenever either changes.