Vector Databases: A Deep Dive

Course Content

Vector Databases: A Deep Dive

3 sections · 5 lessons

Choosing the Right Vector Database


A team picks a vector database off a benchmark chart. The chart says 12 ms at 0.98 recall on 10 million vectors. They build on it. In production they measure 340 ms at the 95th percentile, the service is OOM-killed twice a day, and search costs more than the language model calls it feeds.

Nothing in the benchmark was dishonest. Four things were different:

  • The benchmark used SIFT1M — one million 128-dimension vectors, 512 MB in total. Production had ten million 1,536-dimension vectors: 61.4 GB, or 120× the data. Distance computations scale with dimension, and memory does too.
  • The benchmark issued queries in batches of 1,000 across all 16 cores. Production issues one query per request. Batching hides per-query overhead and inflates throughput numbers by 5 to 10×.
  • The benchmark applied no filters. Production filters by tenant and permission on every single query.
  • The benchmark ran against a static, fully-built index. Production writes 200,000 new chunks a day, so the background optimiser competes for the same cores.

Choosing a vector database is capacity planning, not shopping. This lesson gives you the four questions that actually determine the answer, the memory formula that tells you what you will pay, a way to read benchmark claims without being misled, and the mechanics of moving once you have chosen wrong — because you eventually will.

Four layers, applied in this orderHosting: managed, self-hosted, or embeddedScale: vectors, dimensions, RAM the formula givesFeatures you genuinely need, not the feature listTraffic pattern: peak QPS, write rate, burstiness
A benchmark chart answers layer two only, which is how '12 ms at 0.98 recall' becomes 340 ms at the 95th percentile with two OOM kills a day.

A four-layer selection framework

Answer these in order. Each layer eliminates options, and answering them out of order is how teams end up with a system that is technically excellent and organisationally impossible.

Layer 1 — Hosting model

This is first because it is usually decided by people who are not engineers, and it removes whole categories.

ConstraintImplication
Data cannot leave your infrastructure (regulation, contract, classification)Self-hosted only. Managed SaaS is eliminated regardless of merit.
No infrastructure or on-call teamManaged. Running a stateful distributed database with no one to page is a slow-motion outage.
Existing Kubernetes platform and database expertiseSelf-hosted is genuinely cheap for you; managed premium buys less.
Air-gapped or on-premise deploymentSelf-hosted, and check the licence permits it commercially.
Spiky, unpredictable traffic with long idle periodsServerless managed. Paying for a provisioned cluster to sit idle is pure waste.

Layer 2 — Scale

Not "how much data do you have" but "how much will you have in eighteen months, and how much of it must be in RAM". Scale bands behave qualitatively differently:

VectorsWhat is true at this sizeReasonable choices
Under 100kEverything fits in RAM on a laptop; exact search is under 50 mspgvector, or an in-process index; a dedicated database is overhead
100k – 1MOne node, no sharding, no quantisation neededAnything. Choose on features and ops, not performance
1M – 10MOne large node still works; quantisation starts paying for itselfQdrant, Weaviate, Pinecone, Milvus
10M – 100MSharding or aggressive quantisation becomes mandatoryQdrant, Milvus, Pinecone; self-hosting needs a real operator
100M+Distributed by necessity; disk-based indexes and tiering matterMilvus, Vespa, DiskANN-based systems, Pinecone at high tiers

Layer 3 — Features you genuinely need

Distinguish the features that are hard to add from the ones you can build yourself in a week.

FeatureBuild it yourself?Verdict
Filtered vector search at low selectivityNo — it is an index-internals problemMust come from the engine
Multi-tenancy with isolationPainfullyStrongly prefer native support
Hybrid keyword + vector searchYes, with a separate BM25 store and rank fusionNice to have natively; not decisive
Reranking with a cross-encoderYes, trivially, in your applicationDo not choose a database for this
Embedding generation inside the databaseYesConvenient, but it couples your pipeline to the vendor
Quantisation with rescoringNoMust come from the engine at scale
Point-in-time snapshots and restoreNoMust come from the engine, and test it

The features worth choosing a database for are the ones that live inside the index. Everything that lives in your application layer, you can add later and swap freely.

Layer 4 — Traffic pattern

Four numbers describe your workload, and every capacity conversation needs all four:

  • Peak QPS, not average. A search API that averages 20 QPS but peaks at 400 during business hours must be sized for 400.
  • Read/write ratio. A read-mostly corpus rebuilt nightly is a different system from one absorbing 3,000 upserts per second.
  • Freshness requirement. "Searchable within 60 seconds" and "searchable within 24 hours" allow entirely different architectures — the second lets you rebuild an IVF index offline and swap it atomically.
  • Latency objective, stated as a percentile. "p95 under 100 ms" is a requirement; "fast" is not.

Why utilisation, not throughput, sets your node count

This is the part teams consistently get wrong. Suppose one query occupies a core for 3 ms and you have 8 cores. Peak theoretical throughput is 8 / 0.003 = 2,667 QPS. So an 8-core node handles 2,600 QPS, correct?

No — because latency depends on how full the queue is. For a simple queueing model the mean response time is service time divided by the idle fraction: W=S/(1−ρ)W = S / (1-\rho), where ρ\rho is utilisation. With S = 3 ms:

UtilisationQPS on 8 coresMean response timeRough p95
30%8004.3 ms~13 ms
50%1,3336.0 ms~18 ms
70%1,86710.0 ms~30 ms
90%2,40030.0 ms~90 ms
95%2,53360.0 ms~180 ms

Going from 70% to 95% utilisation buys 36% more throughput and costs 6× the latency. This is why capacity planning targets roughly 60–70% peak utilisation, and why "our benchmark hit 2,500 QPS" is a claim about a machine running at a latency nobody would ship.

The memory formula you can trust

Vector databases are memory-bound. Get this number right and everything else — instance class, shard count, monthly bill — follows.

RAM≈N×(d×b  +  g  +  p)×h\text{RAM} \approx N \times \big(d \times b \;+\; g \;+\; p\big) \times h

where:

  • NN = number of vectors
  • dd = dimensions
  • bb = bytes per component: 4 for float32, 2 for float16, 1 for int8, 0.125 for binary
  • gg = graph overhead per vector. For HNSW this is about 2M×42M \times 4 bytes for the bottom layer plus 8% for upper layers, so roughly 138 bytes at M = 16 and 276 bytes at M = 32. IVF has almost none — just nlist×d×4\text{nlist} \times d \times 4 bytes of centroids in total.
  • pp = payload/metadata bytes per vector, including the inverted indexes over filterable fields
  • hh = headroom multiplier, 1.4 to 1.6, covering compaction, query buffers, replication state and the operating system

Worked, for 50 million vectors at 1,024 dimensions with 300 bytes of payload, HNSW at M = 16:

ConfigurationVectorsGraphPayloadSubtotal×1.5 headroom
float32204.8 GB6.9 GB15.0 GB226.7 GB340 GB
float16102.4 GB6.9 GB15.0 GB124.3 GB187 GB
int8 scalar51.2 GB6.9 GB15.0 GB73.1 GB110 GB
binary (1 bit)6.4 GB6.9 GB15.0 GB28.3 GB42 GB

Read that table as a plan, not as trivia. At float32 you need six or seven 64 GB nodes and a sharding strategy. At int8 you need one 128 GB machine and no distributed system at all — and int8 with rescoring typically costs one to two recall points, which you win back by oversampling candidates. Quantisation is not a micro-optimisation; it is the difference between operating a cluster and operating a server.

The binary row comes with a condition: you must keep the full-precision vectors somewhere for rescoring, usually on SSD. 50M × 1024 × 4 = 204.8 GB of disk, read only for the few hundred candidates per query, which an NVMe drive serves comfortably.

Compute your memory number before you shortlist anything. Half of all vector database decisions are settled by that one line of arithmetic.

Reading scalability claims critically

Vendor benchmarks are not lies; they are answers to questions you did not ask. Translate every claim:

The claimWhat is usually unstatedWhat to measure instead
"Billions of vectors"On how many nodes, with what quantisation, at what recallCost per million vectors per month at your recall target
"Sub-millisecond latency"Batched queries, all cores, no filters, warm cache, tiny dimensionsp95 for a single query, with filters, under concurrent write load
"10,000 QPS"Machine at 95%+ utilisation; p99 unreportedQPS at which p95 crosses your objective
"99% recall"At which ef / nprobe, on which dataset, against what ground truthrecall@10 on your vectors versus your own flat index
"Scales horizontally"Rebalancing time, replication lag, behaviour during node lossKill a node under load and time the recovery
"Real-time updates"Delay before new vectors are actually in the graphInsert, then poll until the vector is retrievable; record the delay

Two dataset details silently invalidate most public comparisons. SIFT1M is 128 dimensions; GloVe-100 is 100. Modern text embeddings are 768 to 3,072. A result at 128 dimensions tells you almost nothing about behaviour at 1,536, because both memory traffic and distance cost scale linearly with dimension while cache behaviour degrades non-linearly. And most benchmark datasets have no metadata at all, so they cannot measure the thing that will actually decide your architecture.

A cost framework that survives price changes

Never compare two vendors on the numbers printed on their pricing pages, because those pages meter different things. Convert everything into two normalised figures:

  • Cost per million vectors stored per month
  • Cost per million queries served

And include the line most comparisons omit. A fully-loaded senior engineer costs on the order of 150,000 USD a year, or about 12,500 a month. Even a quarter of that person's time spent operating a database is 3,125 USD a month — which will dwarf the infrastructure line for most systems under 50 million vectors.

Worked comparison for 10 million vectors at 768 dimensions, int8 quantised, serving 20 million queries a month:

Line itemSelf-hostedManaged
Memory needed (formula above)~22 GB → 32 GB nodesame data, vendor sizes it
Compute2 × 0.25 USD/hr × 730 = 365 USD—
Storage and backups~40 USDincluded
Managed service fee—~700–1,100 USD
Engineering time0.25 FTE = 3,125 USD0.05 FTE = 625 USD
Total per month~3,530 USD~1,600 USD
Cost per million queries176 USD80 USD
Cost per million vectors stored353 USD160 USD

At this scale managed wins, and it is not close. The crossover comes from the fixed engineering line: it barely grows with data volume, so once infrastructure dominates it — somewhere above roughly 100 million vectors, or when you are already running the same infrastructure for other stateful services — self-hosting flips to being clearly cheaper. That relationship holds no matter what any vendor charges next year, which is why it is worth internalising instead of a price table.

Model the shape of the cost, not the price. Fixed labour plus variable infrastructure means small systems should be managed and large systems should not.

Decision matrix by scenario

ScenarioTypical shapeRecommendedWhy, and what to avoid
Startup / MVP<1M vectors, <10 QPS, one engineerpgvector if PostgreSQL exists; otherwise managed serverlessOptimise for changing your mind cheaply. Avoid a cluster you must operate.
Scaling, 10M–100MReal traffic, filters everywhere, cost now visibleQdrant or Milvus, self-hosted or their cloudQuantisation and filtered search become decisive. Avoid per-query metered pricing at high QPS.
Enterprise, 100M+Multi-region, SLAs, platform team, complianceMilvus, Vespa, or a high-tier managed contractOperational maturity and disaster recovery outweigh benchmark deltas.
Structured-data-heavyVectors are one column among many; joins and transactions matterpgvector, or a database with strong typed filteringKeeping vectors beside the source rows removes an entire class of sync bug.
Spiky / burstyIdle for hours, then 500 QPS for twenty minutesServerless managedProvisioned capacity for the peak is idle 95% of the time.
Regulated / air-gappedData cannot leave the perimeterSelf-hosted open sourceLayer 1 already decided this; budget for the operator.
Offline / batch analyticsNo live queries; nightly similarity jobsFAISS on a big machineYou need an index, not a database. Do not pay for a server that is idle all day.

Migration, in practice

You will move at least once. Vectors are portable — they are just floats — so migration is never about the data model. It is about ids, filters, and proving equivalence.

FAISS to a vector database

The trap is identity. A bare FAISS index returns positional indices, so row 4,182 means "the 4,182nd vector you added". If your ingestion order ever changes, or you delete anything, every stored id silently points at the wrong document.

Python
import json, uuidimport faissfrom qdrant_client import QdrantClient, modelsindex = faiss.read_index("corpus.faiss")          # IndexFlatIP over normalised vectorsdoc_ids = json.load(open("doc_ids.json"))         # your positional -> id mapclient = QdrantClient(url="http://localhost:6333")if not client.collection_exists("corpus"):    client.create_collection(        "corpus",        # IndexFlatIP + faiss.normalize_L2 corresponds to COSINE, not DOT        vectors_config=models.VectorParams(size=index.d,                                           distance=models.Distance.COSINE),    )BATCH = 2048for start in range(0, index.ntotal, BATCH):    n = min(BATCH, index.ntotal - start)    vecs = index.reconstruct_n(start, n)          # pull raw vectors back out    points = [        models.PointStruct(            id=str(uuid.uuid5(uuid.NAMESPACE_URL, doc_ids[start + i])),            vector=vecs[i].tolist(),            payload={"doc_id": doc_ids[start + i]},        )        for i in range(n)    ]    client.upsert("corpus", points=points, wait=True)

Three details carry the migration. reconstruct_n works on flat and IVF-flat indexes but returns approximations from a PQ index — if your source is quantised you must re-embed or restore from the original vectors. The metric mapping matters: a FAISS IndexFlatIP fed with faiss.normalize_L2 vectors is cosine, and declaring it as dot product will change your rankings. And uuid5 gives you a deterministic id from your string document id, so re-running the migration is idempotent rather than duplicating everything.

Between two hosted databases

Here the trap is that there is often no bulk export. You paginate through ids and fetch in batches, which for 10 million vectors at 100 per request is 100,000 API calls — budget hours, not minutes, and expect to pay read charges for the privilege.

The other trap is metadata typing. Systems that store metadata as JSON tend to return every number as a float, so an integer timestamp comes back as 1737590400.0. Write it into a system with typed fields and your integer range filters stop matching. Coerce explicitly on the way in.

The cutover protocol

Do not switch on a Friday and hope. This sequence has no downtime and a working escape hatch:

  1. Dual-write. Send every insert, update and delete to both systems. Run this for at least a full data-refresh cycle so you know deletions propagate too.
  2. Backfill the history into the new system while dual-writing continues.
  3. Shadow-read. Query both on live traffic, serve the old results, and log the overlap. Compute overlap@10 between the two result sets over a few thousand real queries. Anything below about 0.9 means a configuration difference — usually the metric, the normalisation, or a filter that translated incorrectly — and you find it here rather than in an incident.
  4. Shift traffic gradually: 1%, 10%, 50%, 100%, watching p95 latency and result-quality metrics at each step.
  5. Keep the old system writable for two weeks after cutover. Rollback should be a configuration flag, not a restore.

Where people get this wrong

Choosing for a scale you do not have. Picking a distributed system at 200,000 vectors because you might reach a billion means paying operational cost every day for a capability you may never use — and by the time you get there, the landscape will have changed anyway. Choose for eighteen months out, not five years.

Ignoring the write path. Evaluations focus entirely on query latency. Then production discovers that HNSW insertion is expensive, that a bulk load competes with queries for cores, and that deletes are tombstones which do not free memory until compaction runs. Benchmark with your real write rate running concurrently.

Sizing on average traffic. Average QPS is the wrong number twice over: peaks are what you must serve, and the queueing table above shows that running near capacity destroys latency long before you run out of throughput.

Forgetting the headroom multiplier. The index needs room to compact, to build a new segment before dropping the old one, and to hold query buffers. Provisioning exactly the computed subtotal is how you get OOM-killed during an ordinary background merge, typically at 3 a.m.

Assuming recall transfers. An ef or nprobe setting that gives 0.98 recall on someone's benchmark dataset can give 0.91 on yours, because recall depends on how clustered your embeddings are. Recall is a property of the data plus the parameters, never of the parameters alone. Re-measure on your own corpus every time.

No exit plan. If your filter construction, id scheme and metric are scattered across the codebase, migration becomes a rewrite. Confine them to one module from day one and the cost of being wrong drops from months to days.

Running the evaluation

Here is a protocol that fits in two days and settles the question with evidence.

  1. Write the requirement down as numbers: vectors at 18 months, dimensions, peak QPS, write rate, p95 objective, recall objective, filter shapes and their selectivities, and any hosting constraint. Most disagreements dissolve the moment this exists.
  2. Compute the memory formula for float32 and for int8. If int8 moves you from a cluster to a single node, your shortlist is now "engines with quantisation and rescoring".
  3. Take a real sample — at least a million of your own vectors with your own metadata. Synthetic random vectors are actively misleading, because random data has no cluster structure and ANN indexes behave differently on it.
  4. Build a flat index as ground truth and hold 200 real queries against it. Every recall number you quote from now on is measured against this.
  5. For each candidate, measure recall@10, p50/p95/p99 latency for single queries at your real concurrency, the same with your real filters applied, index build time, and steady-state memory. Six numbers per candidate, on one table.
  6. Convert costs into cost per million vectors per month and cost per million queries, with the engineering line included.
  7. Write a one-page decision record stating what you chose, the numbers that decided it, and the conditions that would change the answer — a scale threshold, a cost threshold, a missing feature. When someone reopens the debate in a year, that page is worth more than the database.

The teams that get this right are not the ones who picked the best engine. They are the ones whose answer to "why this one?" is a table of measurements rather than a preference, and whose migration path is a single module rather than a rewrite.