Course Content
LangChain Mastery
7 sections · 109 lessons
How do you use LangChain to query a vector store?
What you need to know
Querying the store directly
vector_store.similarity_search("carry forward earned leave", k=4)vector_store.similarity_search_with_score("carry forward earned leave", k=4)vector_store.similarity_search("carry forward", k=4, filter={"country": "IN"})similarity_searchreturns a list ofDocumentobjects.similarity_search_with_scorereturns(Document, score)pairs. The score is whatever the backend uses: some return a distance (lower is better), some a similarity (higher is better).similarity_search_with_relevance_scoresrescales to 0–1 (higher is better), but only on stores that implement it — Chroma and FAISS do,InMemoryVectorStoredoes not.- The
filterformat depends on the store (dict for Chroma, SQL-like for pgvector, a function forInMemoryVectorStore).
Querying through a retriever
1retriever = vector_store.as_retriever(2 search_type="mmr",3 search_kwargs={"k": 4, "fetch_k": 20, "lambda_mult": 0.5},4)5docs = retriever.invoke("carry forward earned leave")The three search types:
search_type | What it does | When to use |
|---|---|---|
"similarity" (default) | Top k by closeness | Most cases |
"mmr" | Fetch fetch_k, then pick k that are relevant and different from each other | Results are near-duplicates |
"similarity_score_threshold" | Only return docs above score_threshold | You would rather return nothing than weak matches |
MMR means maximal marginal relevance. lambda_mult near 1 favours relevance; near 0 favours diversity.
Legacy method names
retriever.get_relevant_documents(q) is the old API. It was deprecated in favour of invoke, which also gives you callbacks, tracing and config for free.
A real-life example
A railway-enquiry bot indexes train FAQ pages. The question "Can I cancel a Tatkal ticket?" returns four chunks — all copies of the same paragraph, because the paragraph appears on four different FAQ pages. The model sees the same fact four times and misses the refund-rules chunk.
The engineer checks with similarity_search_with_score and sees four near-identical scores. Switching to search_type="mmr" with fetch_k=20, k=4 returns one copy of the cancellation rule plus the refund table and the chart-preparation timing. Answer quality improves without any prompt change.
Later they add a threshold: if the best score is below the tuned cut-off, the bot says "I could not find this in the railway rules" instead of guessing.
Follow-up questions to expect
- "How do you pick k?" — Start at 3–5, measure recall@k on labelled questions, and remember each extra chunk costs prompt tokens.
- "Can you reuse a score threshold after changing stores?" — No. Score scales differ per backend and per embedding model, so re-tune it.
- "Store method or retriever in production?" — Retriever. It composes into chains and agents and is traced; use store methods for debugging.