Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
How would you choose a vector database for 20 million chunks with metadata filtering, hybrid search and a small platform team?
What you need to know
Size first: the one number that decides a lot
20,000,000 vectors x 768 dimensions x 4 bytes (float32) = 61.4 GB of raw vectors+ HNSW graph links (a few GB at m = 16)+ metadata payloads and working headroom≈ plan for about 80 GB of RAMThat tells you a single small VM is out, and that memory-saving options matter. It also tells you the cost of a managed service at this scale before any sales call.
The main options
| Option | Choose it when | Watch out for |
|---|---|---|
| pgvector (plus pgvectorscale for larger sets) | You already run Postgres; you need joins, transactions and SQL filters | Heavy filtered ANN at tens of millions needs careful tuning |
| Qdrant | Heavy metadata filtering plus vector search; payload indexes; native sparse vectors for hybrid | Another system to run, unless managed |
| Milvus | Very large scale, many collections, GPU indexes | More moving parts to operate |
| Pinecone or a managed Qdrant/Weaviate | No infra team; cost is acceptable | Vendor lock-in; data-residency terms |
With a small platform team, "who is on call for it?" is a real criterion. A managed service or the Postgres you already run often beats a technically better system nobody knows how to operate.
Filtering is where products differ
A naive vector search finds the top 10 nearest vectors and then applies the filter. If only 2% of chunks match tenant = acme AND year = 2026, most of those 10 are thrown away and you return one or two results. Good engines apply the filter during the graph search. Qdrant's filterable HNSW does this; pgvector 0.8 added iterative scans to reduce the problem. Test your real filter selectivity, not just unfiltered speed.
The tuning dials
- Build parameters —
m(links per node, 16 to 32) andef_construction(128 to 256). Higher means better recall, a slower build and more memory. - Search parameter —
ef_searchis the live recall-versus-latency dial. Raise it until recall@10 meets the target, then stop. - Quantization — int8 scalar quantization cuts vector memory about 4x (61 GB to about 15 GB). Keep full vectors on disk and rescore the top candidates to recover most of the lost recall.
- Hybrid search — add a sparse or BM25 leg for exact terms such as SKUs and error codes, fused with reciprocal rank fusion.
1client.create_collection(2 "chunks",3 vectors_config=VectorParams(size=768, distance=Distance.COSINE),4 hnsw_config=HnswConfigDiff(m=16, ef_construct=200),5 quantization_config=ScalarQuantization(6 scalar=ScalarQuantizationConfig(type=ScalarType.INT8, always_ram=True)),7)This is the Qdrant Python client: an HNSW index with int8 quantized vectors kept in RAM, while the full-precision originals can live on disk for rescoring.
The failure mode to call out
Recall degrades quietly as the index grows and as you delete and re-insert. Nothing errors. So ship a nightly job: take 1,000 real queries, compute exact top-10 by brute force, compare with the index's top-10, and alert if recall@10 drops below the target.
A real-life example
Scenario (illustrative numbers). An online-learning company indexes 20 million chunks of course transcripts and notes. Every query filters by course and language. Their first build on an unfiltered-benchmark favourite returns only three results for many Tamil-language queries, because post-filtering throws away most candidates.
They move to an engine that filters inside the graph, create a payload index on course_id and language, and turn on int8 quantization with rescoring. RAM falls from about 80 GB to about 30 GB including the graph and payloads. The nightly brute-force check shows recall@10 at 0.96 with ef_search = 128, and p95 search latency stays under 40 ms.
Follow-up questions to expect
- "Why not just use pgvector for everything?" — Often you should, below about 10 million vectors and when you already run Postgres. At larger scale with selective filters, a purpose-built engine is easier to tune.
- "How do you choose embedding dimensions?" — Test smaller dimensions on your eval set; many current models support shortened embeddings. Halving dimensions halves memory, often with a small recall loss.
- "What happens when you change the embedding model?" — Every vector must be rebuilt. Build a new index alongside the old one, backfill, compare on the eval set, then switch reads.