Course Content
RAG Systems
12 sections · 66 lessons
What are the key components needed to build and query a vector store?
What you need to know
The components
- Chunks with metadata — what you retrieve and filter on.
- Embedding function — pinned model and version, same for writes and queries.
- The store — where vectors, text and metadata live.
- Distance metric — cosine for most text models (set it explicitly: Chroma, for example, defaults to L2).
- Index and its parameters — see below.
- Write path — upserts by stable ID, deletes for removed documents.
- Query path — k, metadata filters, search type.
- Rebuild script — because the embedding model will change one day.
HNSW: a graph you can walk
Its three knobs:
M— links per node (often 16). Higher means better recall and more memory.ef_construction— how hard it searches when inserting (often 64 to 200). Higher means a better graph and slower builds.ef_search— how many candidates it keeps while searching (often 40 to 200). Higher means better recall and slower queries. It can be changed per query, which makes it the main tuning knob.
In a Chroma 1.x collection, the defaults are max_neighbors (its name for M) 16, ef_construction 100 and ef_search 100. In pgvector the defaults are m = 16, ef_construction = 64 and hnsw.ef_search = 40:
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64);SET hnsw.ef_search = 100; -- per session: trade speed for recallHNSW gives excellent recall and speed but keeps the graph and usually the vectors in memory, and deletes leave gaps that some engines must clean up.
IVF: search a few buckets
IVF (inverted file index) clusters all vectors into nlist groups using k-means. A query compares itself with the cluster centres, then searches only the closest nprobe clusters. With 1 million vectors, nlist = 1000 and nprobe = 10, a query scans about 1% of the vectors. It builds faster and uses less memory than HNSW, but recall drops for vectors near cluster borders unless nprobe rises, and clusters must be retrained if the data drifts.
Quantisation: smaller vectors
Memory for 1 million vectors of 1,024 dimensions:
| Storage | Bytes per vector | 1M vectors |
|---|---|---|
| float32 | 4,096 | ~4.1 GB |
| float16 | 2,048 | ~2.0 GB |
| int8 (scalar quantisation) | 1,024 | ~1.0 GB |
| Product quantisation, 64 codes | 64 | ~64 MB |
| Binary (1 bit per dimension) | 128 | ~128 MB |
Scalar quantisation maps each float to a small integer. Product quantisation (PQ) splits the vector into parts and replaces each part with the ID of its nearest code in a learned codebook. Binary quantisation keeps only the sign of each number. Each step loses some accuracy, so the usual pattern is: search the compressed vectors for, say, the top 100, then re-score those 100 with full-precision vectors to get the final top 10.
1import numpy as np2Db = np.packbits(D > 0, axis=1) # 384 floats (1,536 bytes) -> 48 bytes3qb = np.packbits(q > 0)4hamming = np.unpackbits(Db ^ qb, axis=1).sum(1) # fewer differing bits = closer5cand = np.argsort(hamming)[:100] # fast, rough shortlist6top10 = cand[np.argsort(-(D[cand] @ q))[:10]] # exact re-score on the shortlistA real-life example
An e-commerce company indexes 40 million review and Q&A chunks at 1,024 dimensions for product Q&A. As float32 that is about 164 GB of vectors before any graph, which would need several large memory nodes.
They choose HNSW over int8 vectors (about 41 GB) with full-precision vectors kept on disk for re-scoring the top 100. Search is filtered by product_id first, so each query searches a small slice. They tune ef_search on a labelled set: raising it until recall@10 against exact search stops improving, then stopping, because latency keeps rising after that point.
Follow-up questions to expect
- "HNSW or IVF?" — HNSW for the best recall at low latency when memory allows; IVF (often with PQ) when memory is tight or the collection is very large and mostly static.
- "How do you measure ANN recall?" — Run exact brute-force search on a sample of queries and count what fraction of the true top-k the index returns.
- "Does filtering interact with HNSW?" — Yes. Heavy filters can cut the graph into pieces the search cannot reach; good engines handle this with filter-aware traversal or fall back to exact search on small filtered sets.