Course Content
Vector Databases: A Deep Dive
3 sections · 5 lessons
Choosing the Right Vector Database
A team picks a vector database off a benchmark chart. The chart says 12 ms at 0.98 recall on 10 million vectors. They build on it. In production they measure 340 ms at the 95th percentile, the service is OOM-killed twice a day, and search costs more than the language model calls it feeds.
Nothing in the benchmark was dishonest. Four things were different:
- The benchmark used SIFT1M — one million 128-dimension vectors, 512 MB in total. Production had ten million 1,536-dimension vectors: 61.4 GB, or 120× the data. Distance computations scale with dimension, and memory does too.
- The benchmark issued queries in batches of 1,000 across all 16 cores. Production issues one query per request. Batching hides per-query overhead and inflates throughput numbers by 5 to 10×.
- The benchmark applied no filters. Production filters by tenant and permission on every single query.
- The benchmark ran against a static, fully-built index. Production writes 200,000 new chunks a day, so the background optimiser competes for the same cores.
Choosing a vector database is capacity planning, not shopping. This lesson gives you the four questions that actually determine the answer, the memory formula that tells you what you will pay, a way to read benchmark claims without being misled, and the mechanics of moving once you have chosen wrong — because you eventually will.
A four-layer selection framework
Answer these in order. Each layer eliminates options, and answering them out of order is how teams end up with a system that is technically excellent and organisationally impossible.
Layer 1 — Hosting model
This is first because it is usually decided by people who are not engineers, and it removes whole categories.
| Constraint | Implication |
|---|---|
| Data cannot leave your infrastructure (regulation, contract, classification) | Self-hosted only. Managed SaaS is eliminated regardless of merit. |
| No infrastructure or on-call team | Managed. Running a stateful distributed database with no one to page is a slow-motion outage. |
| Existing Kubernetes platform and database expertise | Self-hosted is genuinely cheap for you; managed premium buys less. |
| Air-gapped or on-premise deployment | Self-hosted, and check the licence permits it commercially. |
| Spiky, unpredictable traffic with long idle periods | Serverless managed. Paying for a provisioned cluster to sit idle is pure waste. |
Layer 2 — Scale
Not "how much data do you have" but "how much will you have in eighteen months, and how much of it must be in RAM". Scale bands behave qualitatively differently:
| Vectors | What is true at this size | Reasonable choices |
|---|---|---|
| Under 100k | Everything fits in RAM on a laptop; exact search is under 50 ms | pgvector, or an in-process index; a dedicated database is overhead |
| 100k – 1M | One node, no sharding, no quantisation needed | Anything. Choose on features and ops, not performance |
| 1M – 10M | One large node still works; quantisation starts paying for itself | Qdrant, Weaviate, Pinecone, Milvus |
| 10M – 100M | Sharding or aggressive quantisation becomes mandatory | Qdrant, Milvus, Pinecone; self-hosting needs a real operator |
| 100M+ | Distributed by necessity; disk-based indexes and tiering matter | Milvus, Vespa, DiskANN-based systems, Pinecone at high tiers |
Layer 3 — Features you genuinely need
Distinguish the features that are hard to add from the ones you can build yourself in a week.
| Feature | Build it yourself? | Verdict |
|---|---|---|
| Filtered vector search at low selectivity | No — it is an index-internals problem | Must come from the engine |
| Multi-tenancy with isolation | Painfully | Strongly prefer native support |
| Hybrid keyword + vector search | Yes, with a separate BM25 store and rank fusion | Nice to have natively; not decisive |
| Reranking with a cross-encoder | Yes, trivially, in your application | Do not choose a database for this |
| Embedding generation inside the database | Yes | Convenient, but it couples your pipeline to the vendor |
| Quantisation with rescoring | No | Must come from the engine at scale |
| Point-in-time snapshots and restore | No | Must come from the engine, and test it |
The features worth choosing a database for are the ones that live inside the index. Everything that lives in your application layer, you can add later and swap freely.
Layer 4 — Traffic pattern
Four numbers describe your workload, and every capacity conversation needs all four:
- Peak QPS, not average. A search API that averages 20 QPS but peaks at 400 during business hours must be sized for 400.
- Read/write ratio. A read-mostly corpus rebuilt nightly is a different system from one absorbing 3,000 upserts per second.
- Freshness requirement. "Searchable within 60 seconds" and "searchable within 24 hours" allow entirely different architectures — the second lets you rebuild an IVF index offline and swap it atomically.
- Latency objective, stated as a percentile. "p95 under 100 ms" is a requirement; "fast" is not.
Why utilisation, not throughput, sets your node count
This is the part teams consistently get wrong. Suppose one query occupies a core for 3 ms and you have 8 cores. Peak theoretical throughput is 8 / 0.003 = 2,667 QPS. So an 8-core node handles 2,600 QPS, correct?
No — because latency depends on how full the queue is. For a simple queueing model the mean response time is service time divided by the idle fraction: W=S/(1−ρ), where ρ is utilisation. With S = 3 ms:
| Utilisation | QPS on 8 cores | Mean response time | Rough p95 |
|---|---|---|---|
| 30% | 800 | 4.3 ms | ~13 ms |
| 50% | 1,333 | 6.0 ms | ~18 ms |
| 70% | 1,867 | 10.0 ms | ~30 ms |
| 90% | 2,400 | 30.0 ms | ~90 ms |
| 95% | 2,533 | 60.0 ms | ~180 ms |
Going from 70% to 95% utilisation buys 36% more throughput and costs 6× the latency. This is why capacity planning targets roughly 60–70% peak utilisation, and why "our benchmark hit 2,500 QPS" is a claim about a machine running at a latency nobody would ship.
The memory formula you can trust
Vector databases are memory-bound. Get this number right and everything else — instance class, shard count, monthly bill — follows.
where:
- N = number of vectors
- d = dimensions
- b = bytes per component: 4 for float32, 2 for float16, 1 for int8, 0.125 for binary
- g = graph overhead per vector. For HNSW this is about 2M×4 bytes for the bottom layer plus 8% for upper layers, so roughly 138 bytes at M = 16 and 276 bytes at M = 32. IVF has almost none — just nlist×d×4 bytes of centroids in total.
- p = payload/metadata bytes per vector, including the inverted indexes over filterable fields
- h = headroom multiplier, 1.4 to 1.6, covering compaction, query buffers, replication state and the operating system
Worked, for 50 million vectors at 1,024 dimensions with 300 bytes of payload, HNSW at M = 16:
| Configuration | Vectors | Graph | Payload | Subtotal | ×1.5 headroom |
|---|---|---|---|---|---|
| float32 | 204.8 GB | 6.9 GB | 15.0 GB | 226.7 GB | 340 GB |
| float16 | 102.4 GB | 6.9 GB | 15.0 GB | 124.3 GB | 187 GB |
| int8 scalar | 51.2 GB | 6.9 GB | 15.0 GB | 73.1 GB | 110 GB |
| binary (1 bit) | 6.4 GB | 6.9 GB | 15.0 GB | 28.3 GB | 42 GB |
Read that table as a plan, not as trivia. At float32 you need six or seven 64 GB nodes and a sharding strategy. At int8 you need one 128 GB machine and no distributed system at all — and int8 with rescoring typically costs one to two recall points, which you win back by oversampling candidates. Quantisation is not a micro-optimisation; it is the difference between operating a cluster and operating a server.
The binary row comes with a condition: you must keep the full-precision vectors somewhere for rescoring, usually on SSD. 50M × 1024 × 4 = 204.8 GB of disk, read only for the few hundred candidates per query, which an NVMe drive serves comfortably.
Compute your memory number before you shortlist anything. Half of all vector database decisions are settled by that one line of arithmetic.
Reading scalability claims critically
Vendor benchmarks are not lies; they are answers to questions you did not ask. Translate every claim:
| The claim | What is usually unstated | What to measure instead |
|---|---|---|
| "Billions of vectors" | On how many nodes, with what quantisation, at what recall | Cost per million vectors per month at your recall target |
| "Sub-millisecond latency" | Batched queries, all cores, no filters, warm cache, tiny dimensions | p95 for a single query, with filters, under concurrent write load |
| "10,000 QPS" | Machine at 95%+ utilisation; p99 unreported | QPS at which p95 crosses your objective |
| "99% recall" | At which ef / nprobe, on which dataset, against what ground truth | recall@10 on your vectors versus your own flat index |
| "Scales horizontally" | Rebalancing time, replication lag, behaviour during node loss | Kill a node under load and time the recovery |
| "Real-time updates" | Delay before new vectors are actually in the graph | Insert, then poll until the vector is retrievable; record the delay |
Two dataset details silently invalidate most public comparisons. SIFT1M is 128 dimensions; GloVe-100 is 100. Modern text embeddings are 768 to 3,072. A result at 128 dimensions tells you almost nothing about behaviour at 1,536, because both memory traffic and distance cost scale linearly with dimension while cache behaviour degrades non-linearly. And most benchmark datasets have no metadata at all, so they cannot measure the thing that will actually decide your architecture.
A cost framework that survives price changes
Never compare two vendors on the numbers printed on their pricing pages, because those pages meter different things. Convert everything into two normalised figures:
- Cost per million vectors stored per month
- Cost per million queries served
And include the line most comparisons omit. A fully-loaded senior engineer costs on the order of 150,000 USD a year, or about 12,500 a month. Even a quarter of that person's time spent operating a database is 3,125 USD a month — which will dwarf the infrastructure line for most systems under 50 million vectors.
Worked comparison for 10 million vectors at 768 dimensions, int8 quantised, serving 20 million queries a month:
| Line item | Self-hosted | Managed |
|---|---|---|
| Memory needed (formula above) | ~22 GB → 32 GB node | same data, vendor sizes it |
| Compute | 2 × 0.25 USD/hr × 730 = 365 USD | — |
| Storage and backups | ~40 USD | included |
| Managed service fee | — | ~700–1,100 USD |
| Engineering time | 0.25 FTE = 3,125 USD | 0.05 FTE = 625 USD |
| Total per month | ~3,530 USD | ~1,600 USD |
| Cost per million queries | 176 USD | 80 USD |
| Cost per million vectors stored | 353 USD | 160 USD |
At this scale managed wins, and it is not close. The crossover comes from the fixed engineering line: it barely grows with data volume, so once infrastructure dominates it — somewhere above roughly 100 million vectors, or when you are already running the same infrastructure for other stateful services — self-hosting flips to being clearly cheaper. That relationship holds no matter what any vendor charges next year, which is why it is worth internalising instead of a price table.
Model the shape of the cost, not the price. Fixed labour plus variable infrastructure means small systems should be managed and large systems should not.
Decision matrix by scenario
| Scenario | Typical shape | Recommended | Why, and what to avoid |
|---|---|---|---|
| Startup / MVP | <1M vectors, <10 QPS, one engineer | pgvector if PostgreSQL exists; otherwise managed serverless | Optimise for changing your mind cheaply. Avoid a cluster you must operate. |
| Scaling, 10M–100M | Real traffic, filters everywhere, cost now visible | Qdrant or Milvus, self-hosted or their cloud | Quantisation and filtered search become decisive. Avoid per-query metered pricing at high QPS. |
| Enterprise, 100M+ | Multi-region, SLAs, platform team, compliance | Milvus, Vespa, or a high-tier managed contract | Operational maturity and disaster recovery outweigh benchmark deltas. |
| Structured-data-heavy | Vectors are one column among many; joins and transactions matter | pgvector, or a database with strong typed filtering | Keeping vectors beside the source rows removes an entire class of sync bug. |
| Spiky / bursty | Idle for hours, then 500 QPS for twenty minutes | Serverless managed | Provisioned capacity for the peak is idle 95% of the time. |
| Regulated / air-gapped | Data cannot leave the perimeter | Self-hosted open source | Layer 1 already decided this; budget for the operator. |
| Offline / batch analytics | No live queries; nightly similarity jobs | FAISS on a big machine | You need an index, not a database. Do not pay for a server that is idle all day. |
Migration, in practice
You will move at least once. Vectors are portable — they are just floats — so migration is never about the data model. It is about ids, filters, and proving equivalence.
FAISS to a vector database
The trap is identity. A bare FAISS index returns positional indices, so row 4,182 means "the 4,182nd vector you added". If your ingestion order ever changes, or you delete anything, every stored id silently points at the wrong document.
1import json, uuid2import faiss3from qdrant_client import QdrantClient, models45index = faiss.read_index("corpus.faiss") # IndexFlatIP over normalised vectors6doc_ids = json.load(open("doc_ids.json")) # your positional -> id map78client = QdrantClient(url="http://localhost:6333")9if not client.collection_exists("corpus"):10 client.create_collection(11 "corpus",12 # IndexFlatIP + faiss.normalize_L2 corresponds to COSINE, not DOT13 vectors_config=models.VectorParams(size=index.d,14 distance=models.Distance.COSINE),15 )1617BATCH = 204818for start in range(0, index.ntotal, BATCH):19 n = min(BATCH, index.ntotal - start)20 vecs = index.reconstruct_n(start, n) # pull raw vectors back out21 points = [22 models.PointStruct(23 id=str(uuid.uuid5(uuid.NAMESPACE_URL, doc_ids[start + i])),24 vector=vecs[i].tolist(),25 payload={"doc_id": doc_ids[start + i]},26 )27 for i in range(n)28 ]29 client.upsert("corpus", points=points, wait=True)Three details carry the migration. reconstruct_n works on flat and IVF-flat indexes but returns approximations from a PQ index — if your source is quantised you must re-embed or restore from the original vectors. The metric mapping matters: a FAISS IndexFlatIP fed with faiss.normalize_L2 vectors is cosine, and declaring it as dot product will change your rankings. And uuid5 gives you a deterministic id from your string document id, so re-running the migration is idempotent rather than duplicating everything.
Between two hosted databases
Here the trap is that there is often no bulk export. You paginate through ids and fetch in batches, which for 10 million vectors at 100 per request is 100,000 API calls — budget hours, not minutes, and expect to pay read charges for the privilege.
The other trap is metadata typing. Systems that store metadata as JSON tend to return every number as a float, so an integer timestamp comes back as 1737590400.0. Write it into a system with typed fields and your integer range filters stop matching. Coerce explicitly on the way in.
The cutover protocol
Do not switch on a Friday and hope. This sequence has no downtime and a working escape hatch:
- Dual-write. Send every insert, update and delete to both systems. Run this for at least a full data-refresh cycle so you know deletions propagate too.
- Backfill the history into the new system while dual-writing continues.
- Shadow-read. Query both on live traffic, serve the old results, and log the overlap. Compute overlap@10 between the two result sets over a few thousand real queries. Anything below about 0.9 means a configuration difference — usually the metric, the normalisation, or a filter that translated incorrectly — and you find it here rather than in an incident.
- Shift traffic gradually: 1%, 10%, 50%, 100%, watching p95 latency and result-quality metrics at each step.
- Keep the old system writable for two weeks after cutover. Rollback should be a configuration flag, not a restore.
Where people get this wrong
Choosing for a scale you do not have. Picking a distributed system at 200,000 vectors because you might reach a billion means paying operational cost every day for a capability you may never use — and by the time you get there, the landscape will have changed anyway. Choose for eighteen months out, not five years.
Ignoring the write path. Evaluations focus entirely on query latency. Then production discovers that HNSW insertion is expensive, that a bulk load competes with queries for cores, and that deletes are tombstones which do not free memory until compaction runs. Benchmark with your real write rate running concurrently.
Sizing on average traffic. Average QPS is the wrong number twice over: peaks are what you must serve, and the queueing table above shows that running near capacity destroys latency long before you run out of throughput.
Forgetting the headroom multiplier. The index needs room to compact, to build a new segment before dropping the old one, and to hold query buffers. Provisioning exactly the computed subtotal is how you get OOM-killed during an ordinary background merge, typically at 3 a.m.
Assuming recall transfers. An ef or nprobe setting that gives 0.98 recall on someone's benchmark dataset can give 0.91 on yours, because recall depends on how clustered your embeddings are. Recall is a property of the data plus the parameters, never of the parameters alone. Re-measure on your own corpus every time.
No exit plan. If your filter construction, id scheme and metric are scattered across the codebase, migration becomes a rewrite. Confine them to one module from day one and the cost of being wrong drops from months to days.
Running the evaluation
Here is a protocol that fits in two days and settles the question with evidence.
- Write the requirement down as numbers: vectors at 18 months, dimensions, peak QPS, write rate, p95 objective, recall objective, filter shapes and their selectivities, and any hosting constraint. Most disagreements dissolve the moment this exists.
- Compute the memory formula for float32 and for int8. If int8 moves you from a cluster to a single node, your shortlist is now "engines with quantisation and rescoring".
- Take a real sample — at least a million of your own vectors with your own metadata. Synthetic random vectors are actively misleading, because random data has no cluster structure and ANN indexes behave differently on it.
- Build a flat index as ground truth and hold 200 real queries against it. Every recall number you quote from now on is measured against this.
- For each candidate, measure recall@10, p50/p95/p99 latency for single queries at your real concurrency, the same with your real filters applied, index build time, and steady-state memory. Six numbers per candidate, on one table.
- Convert costs into cost per million vectors per month and cost per million queries, with the engineering line included.
- Write a one-page decision record stating what you chose, the numbers that decided it, and the conditions that would change the answer — a scale threshold, a cost threshold, a missing feature. When someone reopens the debate in a year, that page is worth more than the database.
The teams that get this right are not the ones who picked the best engine. They are the ones whose answer to "why this one?" is a table of measurements rather than a preference, and whose migration path is a single module rather than a rewrite.