Course Content
LangChain Mastery
7 sections · 109 lessons
What is the role of embeddings in LangChain retrieval?
What you need to know
How similarity search uses them
- At indexing time, each chunk becomes a vector and is stored.
- At query time, the question becomes a vector with the same model.
- The store returns the chunks whose vectors are closest to the question vector.
Keyword search needs shared words. Embeddings capture meaning, so "salary credited late" can match "payroll delay". That is their strength — and also their weakness: exact tokens like ERR_5012, PNR numbers or product SKUs are blurred, because the model cares about meaning, not spelling. That is why hybrid search exists.
The LangChain interface
1from langchain_openai import OpenAIEmbeddings23emb = OpenAIEmbeddings(model="text-embedding-3-small")4doc_vectors = emb.embed_documents(["Refunds take 5-7 days.", "Leave policy ..."])5query_vector = emb.embed_query("how do I get my money back")6print(len(query_vector)) # 1536 for this modelWhy two methods? Some models are trained with different prefixes or instructions for queries and passages (for example the E5 and BGE families). The LangChain wrapper adds the right one for you. You can also create any provider's model with init_embeddings("openai:text-embedding-3-small") from langchain.embeddings.
Trade-offs to know
- Dimensions — bigger vectors can retrieve slightly better but cost more to store and search. Some models (OpenAI's
text-embedding-3family) let you ask for fewer dimensions. - Hosted vs local — a hosted API is easy; a local model (for example via
langchain-huggingface) keeps data in-house and has no per-call cost. - Language — for Hindi, Tamil or mixed "Hinglish" queries, choose a multilingual model and test it on real queries.
- Caching — re-embedding the same chunk costs money.
CacheBackedEmbeddings(now inlangchain_classic.embeddings) stores vectors by a hash of the text.
A real-life example
A bank's internal support bot was built with one embedding model. Six months later a developer switched the config to a newer, "better" model, but only for queries — the index was not rebuilt. Nothing crashed. Similarity scores simply became noise, and the bot started answering questions about credit-card limits with chunks about home-loan interest.
The fix was to store the embedding model name with the index, check it at startup, and fail if the query model is different. Re-embedding 200,000 chunks took about 40 minutes as a batch job, and it became a planned migration instead of a silent bug.
Follow-up questions to expect
- "Can you mix vectors from two models in one index?" — No. The two spaces are unrelated, so distances between them mean nothing.
- "Cosine similarity or dot product?" — For normalised vectors they rank the same. Use what the model and store recommend.
- "How do embeddings handle exact IDs?" — Badly. Add keyword (BM25) search or a metadata filter for identifiers.