Vector Databases: A Deep Dive

Course Content

Vector Databases: A Deep Dive

3 sections · 5 lessons

Pinecone, Qdrant, and Weaviate Overview


Your retrieval prototype works. 40,000 document chunks, a FAISS flat index held in a Python process, 12 ms queries, and answers that impress everyone in the demo. Then three requirements arrive in the same week.

First: results must be restricted to documents the requesting user is allowed to see, and permissions change hourly. Second: the sales team wants to add 5,000 new chunks a day without a nightly rebuild. Third: the service needs to survive a pod restart without a twenty-minute reload from S3.

None of those are search problems. They are database problems — filtering, incremental writes, durability, concurrency, replication. FAISS is a vector index: a superb data structure with no answer to any of them. What you need is a vector database: an index wrapped in a storage engine, a query layer, an API, and an operational story. This lesson takes apart the three systems teams most often choose between — Pinecone, Qdrant and Weaviate — including the code you actually write against each, and the filtering behaviour that decides whether your permissions requirement works or quietly returns three results when you asked for twenty.

What a database adds on top of an indexIndexversus databaseFiltered searchwith honest recallUpserts anddeletes, not a rebuildPersistence andcrash recoveryReplication andhorizontal shardsPayloads storedbeside the vectorsAuth, quotas, multi-tenancy
A FAISS index in a Python process does the arithmetic; everything on the rim is what you would otherwise have to write and operate yourself.

Index versus database, and why the gap costs so much

CapabilityVector index (FAISS, hnswlib, ScaNN)Vector database (Pinecone, Qdrant, Weaviate, Milvus)
Approximate nearest neighbour searchYes, usually fasterYes, built on the same algorithms
Metadata storage and filteringNo — you maintain a side tableNative, with filter-aware search
Incremental insert / update / deletePartial; many index types need rebuildingYes, with background compaction
Durability across restartsYou serialise to disk yourselfWrite-ahead log and snapshots
Concurrent readers and writersYour problemHandled
Horizontal scaling, replicationYour problemSharding and replicas
Access control, multi-tenancyNoneAPI keys, namespaces, tenants
Hybrid keyword + vector searchNoUsually yes

Choosing a vector database is rarely a choice about search quality. All the serious options run HNSW and land within a couple of recall points of each other. You are choosing an operations model.

The filtering problem, with arithmetic

Before looking at any specific product, understand the single technical difference that separates them, because vendors describe it in marketing language and the consequences are severe.

You have 1,000,000 chunks. A user can see 0.1% of them — 1,000 documents. They search, and you want 20 results.

Post-filtering means: run the vector search, then discard results the user cannot see. Ask the index for the top 1,000 by similarity and filter afterwards. How many survive? The visible documents are scattered through the corpus, so of your 1,000 retrieved, the expected number that pass the filter is 1,000 × 0.001 = 1. You asked for 20 and got one. The usual reaction is to retrieve more — top 10,000, top 50,000 — until you are effectively doing brute-force search and your latency has gone from 2 ms to 400 ms.

Pre-filtering means: restrict the candidate set to the 1,000 visible documents first, then search within them. With only 1,000 candidates, exact brute force takes about 0.1 ms. Perfect recall, faster than the approximate search would have been.

The awkward middle case is a filter that matches, say, 30% of the corpus. Pre-filtering by materialising 300,000 ids is expensive; post-filtering wastes maybe 70% of results but still returns enough. The engines differ in how well they handle this middle.

Filter selectivityMatching docs (of 1M)Correct strategyPost-filter outcome if you get it wrong
0.1%1,000Brute force the subset~1 result out of 20 requested
1%10,000Brute force or filtered graph~10 of 20; visibly thin results
10%100,000Filtered graph traversal~100 of 1,000 fetched; usable but wasteful
60%+600,000Plain ANN, filter afterwardsFine

Every engine below claims to filter. What varies is whether the engine estimates selectivity and switches strategy, or always does one thing.

Pinecone

What it is

A fully managed, closed-source vector database, and the one that made the category mainstream. There is no open-source core and no way to run it yourself in production: you get an HTTPS endpoint, an API key, and no servers. (Pinecone Local, a Docker image, is an in-memory emulator for tests; it keeps nothing after it stops.) The serverless architecture separates storage from compute, so an idle index costs storage only and scales reads elastically.

Creating an index and writing to it

Python
from pinecone import Pinecone, ServerlessSpecpc = Pinecone(api_key=PINECONE_API_KEY)if not pc.has_index("support-docs"):    pc.create_index(        name="support-docs",        dimension=768,        metric="cosine",        spec=ServerlessSpec(cloud="aws", region="us-east-1"),    )index = pc.Index("support-docs")    # SDK v10 also offers pc.index(name=...)index.upsert(    vectors=[        {"id": "doc-1#c0",         "values": embed("Resolving payment authorisation failures"),         "metadata": {"team": "billing", "lang": "en", "updated": 1737590400}},        {"id": "doc-1#c1",         "values": embed("If the issuing bank declines the charge..."),         "metadata": {"team": "billing", "lang": "en", "updated": 1737590400}},    ],    namespace="tenant-acme",)

Three design decisions are visible in that snippet. Ids are strings you choose — encoding the parent document and chunk number, as above, makes "delete every chunk of document 1" a prefix operation rather than a lookup table. Metadata is a flat map of strings, numbers, booleans or string lists; there is no nesting, so flatten at write time. Namespaces partition an index; a query touches exactly one namespace, which makes them the natural unit of tenancy.

Querying

Python
res = index.query(    vector=embed("my card got declined"),    top_k=20,    namespace="tenant-acme",    filter={"team": {"$eq": "billing"}, "updated": {"$gte": 1735689600}},    include_metadata=True,)for match in res["matches"]:    print(round(match["score"], 4), match["id"], match["metadata"]["team"])

The filter syntax is MongoDB-flavoured, with $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $and and $or. Pinecone applies filters as part of the search rather than after it, so the 0.1% case above returns a full page of results.

Pros, cons, and how it charges

StrengthsWeaknesses
Zero operational burden; no cluster to size or patchClosed source — you cannot inspect, fork, or self-host it (the local emulator is for tests only)
Serverless idle cost is storage onlyData lives in the vendor's account; a hard stop in some regulated environments
Filtering integrated into searchLeast control over index internals; few tuning dials exposed
Namespaces make per-tenant isolation trivialCosts are usage-driven and therefore hard to forecast for spiky traffic
Mature managed hybrid search and rerankingMigration away requires re-exporting every vector by id

Serverless pricing is metered on three axes: storage (per GB-month of vectors plus metadata), read units (proportional to how much data a query scans, so wide filters and large top_k cost more) and write units (per upsert). The structural consequence outlives any price list: a low-traffic index with lots of data is cheap, and a high-QPS index over a big corpus is where the bill concentrates. Model your cost as roughly storage_GB × rate + monthly_queries × read_units_per_query × rate, and measure read units per query on real traffic rather than guessing, because top_k and filter breadth move that number by an order of magnitude.

Qdrant

What it is

An open-source engine written in Rust (Apache 2.0), available as a single Docker container, a Kubernetes deployment, or a managed cloud service. The same binary runs on your laptop and in production, which makes local development and CI genuinely straightforward. Its distinguishing feature is filterable HNSW: it builds extra graph links for indexed payload fields and estimates filter cardinality at query time, choosing between filtered graph traversal and brute-forcing the matching subset.

Setup and collection creation

Bash
docker run -p 6333:6333 -p 6334:6334 \  -v "$(pwd)/qdrant_storage:/qdrant/storage" \  qdrant/qdrant
Python
from qdrant_client import QdrantClient, modelsclient = QdrantClient(url="http://localhost:6333")   # or url=..., api_key=... for cloud# recreate_collection is deprecated: it silently DROPS existing data.# Check and create instead.if not client.collection_exists("support_docs"):    client.create_collection(        collection_name="support_docs",        vectors_config=models.VectorParams(            size=768,            distance=models.Distance.COSINE,            on_disk=False,        ),        hnsw_config=models.HnswConfigDiff(m=16, ef_construct=200),        quantization_config=models.ScalarQuantization(            scalar=models.ScalarQuantizationConfig(                type=models.ScalarType.INT8,                quantile=0.99,                always_ram=True,            )        ),    )# Payload indexes are what make filtering fast. Without them,# a filter is a linear scan over payloads.client.create_payload_index("support_docs", "team",                            field_schema=models.PayloadSchemaType.KEYWORD)client.create_payload_index("support_docs", "updated",                            field_schema=models.PayloadSchemaType.INTEGER)

Two things there deserve emphasis. recreate_collection is deprecated and destructive — it deletes the collection if it exists and makes a new empty one. It appears in a great deal of older example code, and running that example against production is a data-loss incident. Use collection_exists plus create_collection.

And payload indexes are not optional if you filter. Qdrant will happily accept a filter on an unindexed field and evaluate it by scanning payloads, which turns a 2 ms query into a 200 ms one. Create an index for every field you filter on.

Writing and querying

Python
client.upsert(    collection_name="support_docs",    points=[        models.PointStruct(            id=1,                       # int or UUID, not an arbitrary string            vector=embed("Resolving payment authorisation failures"),            payload={"team": "billing", "lang": "en", "updated": 1737590400,                     "doc_id": "doc-1", "chunk": 0},        ),    ],    wait=True,   # block until the write is applied (the Python client's default))hits = client.query_points(    collection_name="support_docs",    query=embed("my card got declined"),    limit=20,    query_filter=models.Filter(        must=[            models.FieldCondition(key="team", match=models.MatchValue(value="billing")),            models.FieldCondition(key="updated", range=models.Range(gte=1735689600)),        ]    ),    search_params=models.SearchParams(        hnsw_ef=128,        quantization=models.QuantizationSearchParams(rescore=True, oversampling=2.0),    ),    with_payload=True,).points

That search_params block is Qdrant's character in one object. hnsw_ef is the recall/latency dial at query time — no rebuild needed. And with scalar quantisation enabled, oversampling=2.0 fetches twice as many candidates from the int8 index and rescore=True re-ranks them against the original float32 vectors, recovering nearly all the recall the compression cost. Four times less RAM for roughly a millisecond of rescoring.

One gotcha in the write path: wait. The Python client waits by default, but bulk loaders often pass wait=False for speed, and then an upsert returns as soon as it is queued. A test that writes that way and immediately searches will find nothing and look like a bug in the search. In correctness tests, always wait.

StrengthsWeaknesses
Best-in-class filtered search with strategy selectionYou operate it, unless you buy the cloud version
Apache 2.0; identical engine locally and in productionPoint ids must be integers or UUIDs, not free strings
Scalar, product and binary quantisation built inSmaller ecosystem than the largest vendors
Rust engine, low and predictable memory useDistributed mode needs more care than a single node
Rich payload types, nested filters, geo filtersMulti-tenancy is a payload field (flagged is_tenant in its index), not a separate namespace object

Because you can self-host, cost is a capacity question rather than a metering question. Size the RAM (there is a formula later in this lesson), pick an instance, multiply by hours. Qdrant Cloud charges for a cluster by its RAM, vCPU and disk, so the same sizing arithmetic gives you both numbers and lets you compare them honestly.

Weaviate

What it is

An open-source database (BSD-3) written in Go, with a managed cloud offering. Its distinguishing idea is that it wants to be the whole retrieval layer, not just the vector store: it can call an embedding model for you at write and query time (vectoriser modules), run BM25 and vector search together as a first-class hybrid query, and store typed objects with cross-references between collections.

Setup and schema

Python
import weaviatefrom weaviate.classes.config import Configure, Property, DataTypefrom weaviate.classes.query import Filter, MetadataQueryclient = weaviate.connect_to_local()      # or connect_to_weaviate_cloud(...)if not client.collections.exists("SupportDoc"):    client.collections.create(        name="SupportDoc",        # bring your own vectors; use Configure.Vectors.text2vec_*        # only if you want Weaviate to embed for you        vector_config=Configure.Vectors.self_provided(            vector_index_config=Configure.VectorIndex.hnsw(                ef_construction=200,                max_connections=16,                quantizer=Configure.VectorIndex.Quantizer.sq(),            ),        ),        properties=[            Property(name="text", data_type=DataType.TEXT),            Property(name="team", data_type=DataType.TEXT),            Property(name="updated", data_type=DataType.INT),        ],    )

Older v4 examples pass vectorizer_config= and vector_index_config= separately. Recent client versions mark both as deprecated in favour of the single vector_config= shown here.

Note the explicit schema. Weaviate is the most database-like of the three: properties have declared types, and you get typed filters and aggregations as a result. The cost is ceremony — adding a filterable field means a schema change.

Inserting and querying

Python
docs = client.collections.get("SupportDoc")with docs.batch.dynamic() as batch:          # batching is the fast path    batch.add_object(        properties={"text": "Resolving payment authorisation failures",                    "team": "billing", "updated": 1737590400},        vector=embed("Resolving payment authorisation failures"),    )response = docs.query.near_vector(    near_vector=embed("my card got declined"),    limit=20,    filters=(Filter.by_property("team").equal("billing")             & Filter.by_property("updated").greater_or_equal(1735689600)),    return_metadata=MetadataQuery(distance=True),)for obj in response.objects:    print(round(obj.metadata.distance, 4), obj.properties["team"])client.close()      # the v4 client holds a gRPC connection; close it

Hybrid search is where Weaviate is genuinely differentiated — one call runs BM25 and vector search and fuses the rankings:

Python
response = docs.query.hybrid(    query="card declined",              # used for the BM25 half    vector=embed("my card got declined"),    alpha=0.6,                          # 1.0 = pure vector, 0.0 = pure keyword    limit=20,)

That alpha is worth a moment. At 0.6 the fused score leans on the vector side while still rewarding literal term matches, which is exactly what rescues the error-code and product-identifier queries that pure embedding search handles badly. Tuning alpha on a labelled query set usually gains more than any index parameter.

StrengthsWeaknesses
First-class hybrid (BM25 + vector) search with one knobHeavier: Go service, more memory per object than Qdrant
Vectoriser modules can remove your embedding pipeline entirelyThose modules couple you to Weaviate's model integrations
Typed schema, cross-references, aggregationsSchema changes are real migrations
Native multi-tenancy with per-tenant index isolationThe v3 to v4 client rewrite invalidated most tutorial code online
Generative search built inMore concepts to learn before the first query runs

Managed Weaviate charges by stored dimensions with a service-level tier multiplier — the number to compute is objects × dimensions, not gigabytes. Self-hosted, it is again an instance-sizing exercise, with the caveat that its per-object overhead is higher than Qdrant's because of the object store and inverted indexes alongside the vectors.

Side by side

PineconeQdrantWeaviate
LicenceClosed, SaaS onlyApache 2.0BSD-3
Run locallyNoYes (Docker, one command)Yes (Docker)
Language—RustGo
IndexProprietary, graph-basedHNSW, filter-awareHNSW, plus flat and dynamic
FilteringIntegrated single-stageCardinality-aware strategy switchInverted index + roaring bitmaps
QuantisationManaged, not exposedScalar, product, binaryScalar, product, binary
Hybrid searchYes (sparse-dense)Yes (sparse vectors + fusion)Yes, native BM25 fusion with alpha
Multi-tenancyNamespacesPayload field convention, or per-collectionNative tenants with isolation
Embedding generationOptional managed modelsNo, bring your ownYes, vectoriser modules
Ops burdenNoneLow to moderateModerate
Cost shapeMetered usageCapacity (RAM-driven)Stored dimensions or capacity

Sizing and cost, without a price table

Prices change; the arithmetic does not. Work out how much RAM your data needs, then price that.

For 5,000,000 chunks at 768 dimensions with roughly 500 bytes of metadata each:

ComponentCalculationfloat32int8 quantised
Vectors5,000,000 × 768 × bytes15.36 GB3.84 GB
HNSW graph (M = 16)5,000,000 × ~138 B0.69 GB0.69 GB
Payload / metadata5,000,000 × 500 B2.50 GB2.50 GB
Subtotal18.55 GB7.03 GB
With 1.5× operating headroomcompaction, query buffers, OS27.8 GB10.5 GB
Instance class needed32 GB RAM16 GB RAM

A 32 GB / 4 vCPU cloud instance runs around 0.25 USD per hour on demand, so 0.25 × 730 ≈ 183 USD per month, doubled to about 366 for a replica. The 16 GB option halves it. Add engineering time — realistically a few hours a month once steady, considerably more during the first quarter — and compare against the managed quote for the same capacity. Managed services typically land at 1.5–3× the raw infrastructure cost, which is cheap if it saves you a day a month and expensive if you already run a database team.

Quantisation is usually the largest single lever on your bill: int8 with rescoring cuts vector memory by 75% for one to two points of recall you can win back by oversampling.

Using them through LangChain

All three implement the same retriever interface, so application code is portable even though the setup is not.

Python
from langchain_pinecone import PineconeVectorStorefrom langchain_qdrant import QdrantVectorStorefrom langchain_weaviate import WeaviateVectorStore# Pineconestore = PineconeVectorStore(index=index, embedding=embeddings,                            namespace="tenant-acme")# Qdrantstore = QdrantVectorStore(client=client, collection_name="support_docs",                          embedding=embeddings)# Weaviatestore = WeaviateVectorStore(client=client, index_name="SupportDoc",                            text_key="text", embedding=embeddings)# Identical from here onretriever = store.as_retriever(search_kwargs={"k": 20})docs = retriever.invoke("my card got declined")

Where the abstraction leaks is filtering. The filter value inside search_kwargs is passed through to the underlying client untranslated, so it is a Mongo-style dict for Pinecone, a models.Filter object for Qdrant, and a Filter builder expression for Weaviate. Any code that filters is engine-specific. If you expect to switch engines, put filter construction behind one function of your own — that function is the entire migration surface, and it is usually under a hundred lines.

Where people get this wrong

Benchmarking without filters. Almost every published comparison measures unfiltered top-k on a static corpus. Real workloads filter on tenant, permission, date and language simultaneously, and that is precisely where the engines diverge. Benchmark with your real filter shapes or your benchmark is fiction.

Filtering on an unindexed field. In Qdrant this silently becomes a payload scan; in Weaviate an unindexed property cannot be filtered efficiently either. The symptom is a query that is fast in development with 10,000 rows and 100× slower in production, with no error anywhere.

Copying recreate_collection from a tutorial. It drops your data. The same trap exists in spirit elsewhere: any "create or reset" convenience call belongs in tests only.

Querying before the index is built. Qdrant returns results immediately after an upsert, but the HNSW graph is built by a background optimiser; queries during that window fall back to exact scan. Your latency measurements will be wrong in one direction and your recall wrong in the other. Wait for the collection status to go green before measuring anything.

Treating namespaces, tenants and collections as interchangeable. They differ in isolation and cost. A namespace or tenant per customer is cheap and scales to thousands; a collection per customer means a separate HNSW graph each, and a few thousand of those will exhaust memory on any machine.

Leaving the Weaviate v4 client open. It holds a gRPC connection; forgetting client.close() in a request handler leaks connections until the service falls over. Use it as a context manager where you can.

How to actually decide

Work down this list and stop at the first line that describes you.

  1. Under 100,000 vectors, and you already run PostgreSQL. Use pgvector. One system, transactional consistency with your application data, and joins for free. Revisit at a million.
  2. No infrastructure team and no data-residency constraint. Pinecone. You will pay a premium and get all of your engineering time back.
  3. Heavy metadata filtering, per-tenant permissions, or cost sensitivity at scale. Qdrant. The filtering behaviour and the quantisation options are the differentiators, and self-hosting caps your bill.
  4. You want hybrid keyword-plus-vector search out of the box, or you want the database to do the embedding. Weaviate.
  5. Billions of vectors, dedicated platform team. Look at Milvus and the DiskANN family; the three here are all viable at that scale but the operational calculus changes.

Then do the thing that actually settles it: load one million of your vectors into your top two candidates, with your real metadata and your real filter shapes, and measure recall@10 against a flat index plus p95 latency under concurrent load. It takes about a day. Every team that skips it makes the decision on marketing pages and discovers the filtering behaviour in production, which costs considerably more than a day.