Course Content
LangChain Mastery
7 sections · 109 lessons
How do you create a vector store in LangChain?
What you need to know
The four pieces
- Documents —
Document(page_content=..., metadata={...}). Metadata is what you later filter and cite on. - Text splitter — cuts documents into chunks.
RecursiveCharacterTextSplittertries paragraph breaks first, then sentences, then words, so chunks end at natural boundaries. - Embedding model — turns text into a vector. Every LangChain embedding class has the same two methods, so they are swappable.
- Vector store — every integration shares one interface:
from_documents,add_documents,similarity_search,delete,as_retriever.
1from langchain_text_splitters import RecursiveCharacterTextSplitter2from langchain_openai import OpenAIEmbeddings3from langchain_chroma import Chroma45splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=150)6chunks = splitter.split_documents(docs)78embeddings = OpenAIEmbeddings(model="text-embedding-3-small")9vector_store = Chroma.from_documents(10 chunks, embeddings,11 collection_name="hr_policies",12 persist_directory="./chroma",13)14retriever = vector_store.as_retriever(search_kwargs={"k": 4})from_documents embeds all chunks in batches and writes them. Later, vector_store.add_documents(new_chunks) appends more. Each vector store is its own package (langchain-chroma, langchain-postgres, langchain-qdrant, langchain-pinecone), installed next to langchain-core.
Choosing a store
| Store | Runs where | Good for |
|---|---|---|
InMemoryVectorStore | Python process | Tests, demos, a few thousand chunks |
| FAISS | Python process, saved to disk | Fast local search, offline batch jobs |
| Chroma | Local file or server | Prototypes that must survive a restart |
| pgvector | Your Postgres | You already run Postgres; you want SQL joins and transactions |
| Qdrant, Pinecone, Weaviate, Milvus | Managed or self-hosted service | Millions of vectors, replicas, native hybrid search |
Chunk size and overlap
- Too big (say 4,000 characters): one vector averages several topics, so it matches many questions weakly.
- Too small (say 100 characters): a chunk loses the sentence that gives it meaning.
- Overlap (10–15%) keeps a sentence that crosses a boundary in both chunks.
A starting point is 500–1,000 characters with 100–150 overlap, then measure recall and adjust.
A real-life example
An e-commerce support team indexes 12,000 help-centre articles. First try: chunk_size=4000. Questions such as "Can I return a phone after opening the box?" retrieve the whole generic "Returns" article, and the model answers with the default 7-day rule, missing the electronics exception three paragraphs down.
They switch to chunk_size=800, chunk_overlap=120 with a markdown header splitter, so each section ("Electronics", "Fashion", "Groceries") becomes its own chunk with the header stored in metadata. The electronics exception is now a chunk of its own and ranks first. The index grows from 15,000 to 61,000 vectors — still small for pgvector, which they already run for orders.
Follow-up questions to expect
- "Why not store whole documents?" — One vector per long document averages many topics, so search precision drops and you pay to send the whole document to the model.
- "
from_documentsoradd_documents?" —from_documentscreates and fills a new store;add_documentsappends to an existing one. For repeated syncs use the indexing API so you don't create duplicates. - "How do you pick an embedding model?" — Test two or three on 50–100 of your own labelled questions; public leaderboards do not know your domain.