LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you use LangChain to implement semantic search?


What you need to know

Keyword vs semantic

Keyword search (BM25)

  • Matches exact words and their frequency
  • Great for IDs, codes, names
  • Misses synonyms: "salary" does not match "payroll"

Semantic search (embeddings)

  • Matches meaning
  • Handles paraphrase and other wording
  • Blurs exact tokens like ERR_5012 or a PNR number

The code

Python
from langchain_chroma import Chromafrom langchain_openai import OpenAIEmbeddingsstore = Chroma.from_documents(chunks, OpenAIEmbeddings(model="text-embedding-3-small"))retriever = store.as_retriever(search_type="mmr",                               search_kwargs={"k": 5, "fetch_k": 25})for doc in retriever.invoke("cancel my subscription"):    print(doc.metadata["source"], "|", doc.page_content[:100])

There is no LLM here: semantic search is useful on its own, for a search box, "related articles", or duplicate-ticket detection. RAG adds an LLM on top to write an answer from these results.

MMR in one paragraph

Plain similarity often returns near-duplicates — the same paragraph copied across pages. MMR (maximal marginal relevance) first fetches fetch_k candidates, then picks results one by one, each time preferring a candidate that is relevant to the query and different from what it has already picked. You lose a little relevance and gain coverage.

Making results trustworthy

  • Show scores while tuning (similarity_search_with_score) to learn what a "good" match looks like for your data.
  • Filter by metadata (language, product, date) so you search only what applies.
  • Add a threshold so nonsense queries return "no results" instead of the least-bad match.

A real-life example

A food-delivery app's help centre had keyword search. Customers typing "food came cold" got zero results, because the article is titled "Order arrived at the wrong temperature". After switching to semantic search, that query finds it first.

Then a new problem appears: agents searching for a specific order id like ORD-88213 get random articles, because embeddings do not understand ids. The team routes queries matching an id pattern to a keyword/database lookup and everything else to semantic search, and later merges both into hybrid search. Zero-result searches drop from 18% to 3% in the first week (made-up numbers for this scenario).

Follow-up questions to expect

  • "How is semantic search different from RAG?" — Semantic search returns documents; RAG passes those documents to an LLM to write an answer.
  • "How do you evaluate it?" — A labelled set of queries with the correct documents, measured with recall@k and MRR.
  • "Why is my semantic search slow?" — Usually embedding the query over the network, or exact search over a large index; use an approximate index such as HNSW and cache query embeddings.