LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you use LangChain to query a vector store?


Top-6 similarity scores for one Tatkal question0.910.90.90.890.840.8012345sameparagraphcopy 4refund tablePlain similarity fills k=4 with one paragraph copied across four FAQ pages; MMR swaps copies for the refund table.
Relevance alone rewards duplicates — MMR trades a little score for chunks that add something new.

What you need to know

Querying the store directly

Python
vector_store.similarity_search("carry forward earned leave", k=4)vector_store.similarity_search_with_score("carry forward earned leave", k=4)vector_store.similarity_search("carry forward", k=4, filter={"country": "IN"})
  • similarity_search returns a list of Document objects.
  • similarity_search_with_score returns (Document, score) pairs. The score is whatever the backend uses: some return a distance (lower is better), some a similarity (higher is better).
  • similarity_search_with_relevance_scores rescales to 0–1 (higher is better), but only on stores that implement it — Chroma and FAISS do, InMemoryVectorStore does not.
  • The filter format depends on the store (dict for Chroma, SQL-like for pgvector, a function for InMemoryVectorStore).

Querying through a retriever

Python
retriever = vector_store.as_retriever(    search_type="mmr",    search_kwargs={"k": 4, "fetch_k": 20, "lambda_mult": 0.5},)docs = retriever.invoke("carry forward earned leave")

The three search types:

search_typeWhat it doesWhen to use
"similarity" (default)Top k by closenessMost cases
"mmr"Fetch fetch_k, then pick k that are relevant and different from each otherResults are near-duplicates
"similarity_score_threshold"Only return docs above score_thresholdYou would rather return nothing than weak matches

MMR means maximal marginal relevance. lambda_mult near 1 favours relevance; near 0 favours diversity.

Legacy method names

retriever.get_relevant_documents(q) is the old API. It was deprecated in favour of invoke, which also gives you callbacks, tracing and config for free.

A real-life example

A railway-enquiry bot indexes train FAQ pages. The question "Can I cancel a Tatkal ticket?" returns four chunks — all copies of the same paragraph, because the paragraph appears on four different FAQ pages. The model sees the same fact four times and misses the refund-rules chunk.

The engineer checks with similarity_search_with_score and sees four near-identical scores. Switching to search_type="mmr" with fetch_k=20, k=4 returns one copy of the cancellation rule plus the refund table and the chart-preparation timing. Answer quality improves without any prompt change.

Later they add a threshold: if the best score is below the tuned cut-off, the bot says "I could not find this in the railway rules" instead of guessing.

Follow-up questions to expect

  • "How do you pick k?" — Start at 3–5, measure recall@k on labelled questions, and remember each extra chunk costs prompt tokens.
  • "Can you reuse a score threshold after changing stores?" — No. Score scales differ per backend and per embedding model, so re-tune it.
  • "Store method or retriever in production?" — Retriever. It composes into chains and agents and is traced; use store methods for debugging.