Course Content
Vector Databases: A Deep Dive
3 sections · 5 lessons
Pinecone, Qdrant, and Weaviate Overview
Your retrieval prototype works. 40,000 document chunks, a FAISS flat index held in a Python process, 12 ms queries, and answers that impress everyone in the demo. Then three requirements arrive in the same week.
First: results must be restricted to documents the requesting user is allowed to see, and permissions change hourly. Second: the sales team wants to add 5,000 new chunks a day without a nightly rebuild. Third: the service needs to survive a pod restart without a twenty-minute reload from S3.
None of those are search problems. They are database problems — filtering, incremental writes, durability, concurrency, replication. FAISS is a vector index: a superb data structure with no answer to any of them. What you need is a vector database: an index wrapped in a storage engine, a query layer, an API, and an operational story. This lesson takes apart the three systems teams most often choose between — Pinecone, Qdrant and Weaviate — including the code you actually write against each, and the filtering behaviour that decides whether your permissions requirement works or quietly returns three results when you asked for twenty.
Index versus database, and why the gap costs so much
| Capability | Vector index (FAISS, hnswlib, ScaNN) | Vector database (Pinecone, Qdrant, Weaviate, Milvus) |
|---|---|---|
| Approximate nearest neighbour search | Yes, usually faster | Yes, built on the same algorithms |
| Metadata storage and filtering | No — you maintain a side table | Native, with filter-aware search |
| Incremental insert / update / delete | Partial; many index types need rebuilding | Yes, with background compaction |
| Durability across restarts | You serialise to disk yourself | Write-ahead log and snapshots |
| Concurrent readers and writers | Your problem | Handled |
| Horizontal scaling, replication | Your problem | Sharding and replicas |
| Access control, multi-tenancy | None | API keys, namespaces, tenants |
| Hybrid keyword + vector search | No | Usually yes |
Choosing a vector database is rarely a choice about search quality. All the serious options run HNSW and land within a couple of recall points of each other. You are choosing an operations model.
The filtering problem, with arithmetic
Before looking at any specific product, understand the single technical difference that separates them, because vendors describe it in marketing language and the consequences are severe.
You have 1,000,000 chunks. A user can see 0.1% of them — 1,000 documents. They search, and you want 20 results.
Post-filtering means: run the vector search, then discard results the user cannot see. Ask the index for the top 1,000 by similarity and filter afterwards. How many survive? The visible documents are scattered through the corpus, so of your 1,000 retrieved, the expected number that pass the filter is 1,000 × 0.001 = 1. You asked for 20 and got one. The usual reaction is to retrieve more — top 10,000, top 50,000 — until you are effectively doing brute-force search and your latency has gone from 2 ms to 400 ms.
Pre-filtering means: restrict the candidate set to the 1,000 visible documents first, then search within them. With only 1,000 candidates, exact brute force takes about 0.1 ms. Perfect recall, faster than the approximate search would have been.
The awkward middle case is a filter that matches, say, 30% of the corpus. Pre-filtering by materialising 300,000 ids is expensive; post-filtering wastes maybe 70% of results but still returns enough. The engines differ in how well they handle this middle.
| Filter selectivity | Matching docs (of 1M) | Correct strategy | Post-filter outcome if you get it wrong |
|---|---|---|---|
| 0.1% | 1,000 | Brute force the subset | ~1 result out of 20 requested |
| 1% | 10,000 | Brute force or filtered graph | ~10 of 20; visibly thin results |
| 10% | 100,000 | Filtered graph traversal | ~100 of 1,000 fetched; usable but wasteful |
| 60%+ | 600,000 | Plain ANN, filter afterwards | Fine |
Every engine below claims to filter. What varies is whether the engine estimates selectivity and switches strategy, or always does one thing.
Pinecone
What it is
A fully managed, closed-source vector database, and the one that made the category mainstream. There is no open-source core and no way to run it yourself in production: you get an HTTPS endpoint, an API key, and no servers. (Pinecone Local, a Docker image, is an in-memory emulator for tests; it keeps nothing after it stops.) The serverless architecture separates storage from compute, so an idle index costs storage only and scales reads elastically.
Creating an index and writing to it
1from pinecone import Pinecone, ServerlessSpec23pc = Pinecone(api_key=PINECONE_API_KEY)45if not pc.has_index("support-docs"):6 pc.create_index(7 name="support-docs",8 dimension=768,9 metric="cosine",10 spec=ServerlessSpec(cloud="aws", region="us-east-1"),11 )1213index = pc.Index("support-docs") # SDK v10 also offers pc.index(name=...)1415index.upsert(16 vectors=[17 {"id": "doc-1#c0",18 "values": embed("Resolving payment authorisation failures"),19 "metadata": {"team": "billing", "lang": "en", "updated": 1737590400}},20 {"id": "doc-1#c1",21 "values": embed("If the issuing bank declines the charge..."),22 "metadata": {"team": "billing", "lang": "en", "updated": 1737590400}},23 ],24 namespace="tenant-acme",25)Three design decisions are visible in that snippet. Ids are strings you choose — encoding the parent document and chunk number, as above, makes "delete every chunk of document 1" a prefix operation rather than a lookup table. Metadata is a flat map of strings, numbers, booleans or string lists; there is no nesting, so flatten at write time. Namespaces partition an index; a query touches exactly one namespace, which makes them the natural unit of tenancy.
Querying
1res = index.query(2 vector=embed("my card got declined"),3 top_k=20,4 namespace="tenant-acme",5 filter={"team": {"$eq": "billing"}, "updated": {"$gte": 1735689600}},6 include_metadata=True,7)89for match in res["matches"]:10 print(round(match["score"], 4), match["id"], match["metadata"]["team"])The filter syntax is MongoDB-flavoured, with $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $and and $or. Pinecone applies filters as part of the search rather than after it, so the 0.1% case above returns a full page of results.
Pros, cons, and how it charges
| Strengths | Weaknesses |
|---|---|
| Zero operational burden; no cluster to size or patch | Closed source — you cannot inspect, fork, or self-host it (the local emulator is for tests only) |
| Serverless idle cost is storage only | Data lives in the vendor's account; a hard stop in some regulated environments |
| Filtering integrated into search | Least control over index internals; few tuning dials exposed |
| Namespaces make per-tenant isolation trivial | Costs are usage-driven and therefore hard to forecast for spiky traffic |
| Mature managed hybrid search and reranking | Migration away requires re-exporting every vector by id |
Serverless pricing is metered on three axes: storage (per GB-month of vectors plus metadata), read units (proportional to how much data a query scans, so wide filters and large top_k cost more) and write units (per upsert). The structural consequence outlives any price list: a low-traffic index with lots of data is cheap, and a high-QPS index over a big corpus is where the bill concentrates. Model your cost as roughly storage_GB × rate + monthly_queries × read_units_per_query × rate, and measure read units per query on real traffic rather than guessing, because top_k and filter breadth move that number by an order of magnitude.
Qdrant
What it is
An open-source engine written in Rust (Apache 2.0), available as a single Docker container, a Kubernetes deployment, or a managed cloud service. The same binary runs on your laptop and in production, which makes local development and CI genuinely straightforward. Its distinguishing feature is filterable HNSW: it builds extra graph links for indexed payload fields and estimates filter cardinality at query time, choosing between filtered graph traversal and brute-forcing the matching subset.
Setup and collection creation
docker run -p 6333:6333 -p 6334:6334 \ -v "$(pwd)/qdrant_storage:/qdrant/storage" \ qdrant/qdrant1from qdrant_client import QdrantClient, models23client = QdrantClient(url="http://localhost:6333") # or url=..., api_key=... for cloud45# recreate_collection is deprecated: it silently DROPS existing data.6# Check and create instead.7if not client.collection_exists("support_docs"):8 client.create_collection(9 collection_name="support_docs",10 vectors_config=models.VectorParams(11 size=768,12 distance=models.Distance.COSINE,13 on_disk=False,14 ),15 hnsw_config=models.HnswConfigDiff(m=16, ef_construct=200),16 quantization_config=models.ScalarQuantization(17 scalar=models.ScalarQuantizationConfig(18 type=models.ScalarType.INT8,19 quantile=0.99,20 always_ram=True,21 )22 ),23 )2425# Payload indexes are what make filtering fast. Without them,26# a filter is a linear scan over payloads.27client.create_payload_index("support_docs", "team",28 field_schema=models.PayloadSchemaType.KEYWORD)29client.create_payload_index("support_docs", "updated",30 field_schema=models.PayloadSchemaType.INTEGER)Two things there deserve emphasis. recreate_collection is deprecated and destructive — it deletes the collection if it exists and makes a new empty one. It appears in a great deal of older example code, and running that example against production is a data-loss incident. Use collection_exists plus create_collection.
And payload indexes are not optional if you filter. Qdrant will happily accept a filter on an unindexed field and evaluate it by scanning payloads, which turns a 2 ms query into a 200 ms one. Create an index for every field you filter on.
Writing and querying
1client.upsert(2 collection_name="support_docs",3 points=[4 models.PointStruct(5 id=1, # int or UUID, not an arbitrary string6 vector=embed("Resolving payment authorisation failures"),7 payload={"team": "billing", "lang": "en", "updated": 1737590400,8 "doc_id": "doc-1", "chunk": 0},9 ),10 ],11 wait=True, # block until the write is applied (the Python client's default)12)1314hits = client.query_points(15 collection_name="support_docs",16 query=embed("my card got declined"),17 limit=20,18 query_filter=models.Filter(19 must=[20 models.FieldCondition(key="team", match=models.MatchValue(value="billing")),21 models.FieldCondition(key="updated", range=models.Range(gte=1735689600)),22 ]23 ),24 search_params=models.SearchParams(25 hnsw_ef=128,26 quantization=models.QuantizationSearchParams(rescore=True, oversampling=2.0),27 ),28 with_payload=True,29).pointsThat search_params block is Qdrant's character in one object. hnsw_ef is the recall/latency dial at query time — no rebuild needed. And with scalar quantisation enabled, oversampling=2.0 fetches twice as many candidates from the int8 index and rescore=True re-ranks them against the original float32 vectors, recovering nearly all the recall the compression cost. Four times less RAM for roughly a millisecond of rescoring.
One gotcha in the write path: wait. The Python client waits by default, but bulk loaders often pass wait=False for speed, and then an upsert returns as soon as it is queued. A test that writes that way and immediately searches will find nothing and look like a bug in the search. In correctness tests, always wait.
| Strengths | Weaknesses |
|---|---|
| Best-in-class filtered search with strategy selection | You operate it, unless you buy the cloud version |
| Apache 2.0; identical engine locally and in production | Point ids must be integers or UUIDs, not free strings |
| Scalar, product and binary quantisation built in | Smaller ecosystem than the largest vendors |
| Rust engine, low and predictable memory use | Distributed mode needs more care than a single node |
| Rich payload types, nested filters, geo filters | Multi-tenancy is a payload field (flagged is_tenant in its index), not a separate namespace object |
Because you can self-host, cost is a capacity question rather than a metering question. Size the RAM (there is a formula later in this lesson), pick an instance, multiply by hours. Qdrant Cloud charges for a cluster by its RAM, vCPU and disk, so the same sizing arithmetic gives you both numbers and lets you compare them honestly.
Weaviate
What it is
An open-source database (BSD-3) written in Go, with a managed cloud offering. Its distinguishing idea is that it wants to be the whole retrieval layer, not just the vector store: it can call an embedding model for you at write and query time (vectoriser modules), run BM25 and vector search together as a first-class hybrid query, and store typed objects with cross-references between collections.
Setup and schema
1import weaviate2from weaviate.classes.config import Configure, Property, DataType3from weaviate.classes.query import Filter, MetadataQuery45client = weaviate.connect_to_local() # or connect_to_weaviate_cloud(...)67if not client.collections.exists("SupportDoc"):8 client.collections.create(9 name="SupportDoc",10 # bring your own vectors; use Configure.Vectors.text2vec_*11 # only if you want Weaviate to embed for you12 vector_config=Configure.Vectors.self_provided(13 vector_index_config=Configure.VectorIndex.hnsw(14 ef_construction=200,15 max_connections=16,16 quantizer=Configure.VectorIndex.Quantizer.sq(),17 ),18 ),19 properties=[20 Property(name="text", data_type=DataType.TEXT),21 Property(name="team", data_type=DataType.TEXT),22 Property(name="updated", data_type=DataType.INT),23 ],24 )Older v4 examples pass vectorizer_config= and vector_index_config= separately. Recent client versions mark both as deprecated in favour of the single vector_config= shown here.
Note the explicit schema. Weaviate is the most database-like of the three: properties have declared types, and you get typed filters and aggregations as a result. The cost is ceremony — adding a filterable field means a schema change.
Inserting and querying
1docs = client.collections.get("SupportDoc")23with docs.batch.dynamic() as batch: # batching is the fast path4 batch.add_object(5 properties={"text": "Resolving payment authorisation failures",6 "team": "billing", "updated": 1737590400},7 vector=embed("Resolving payment authorisation failures"),8 )910response = docs.query.near_vector(11 near_vector=embed("my card got declined"),12 limit=20,13 filters=(Filter.by_property("team").equal("billing")14 & Filter.by_property("updated").greater_or_equal(1735689600)),15 return_metadata=MetadataQuery(distance=True),16)1718for obj in response.objects:19 print(round(obj.metadata.distance, 4), obj.properties["team"])2021client.close() # the v4 client holds a gRPC connection; close itHybrid search is where Weaviate is genuinely differentiated — one call runs BM25 and vector search and fuses the rankings:
1response = docs.query.hybrid(2 query="card declined", # used for the BM25 half3 vector=embed("my card got declined"),4 alpha=0.6, # 1.0 = pure vector, 0.0 = pure keyword5 limit=20,6)That alpha is worth a moment. At 0.6 the fused score leans on the vector side while still rewarding literal term matches, which is exactly what rescues the error-code and product-identifier queries that pure embedding search handles badly. Tuning alpha on a labelled query set usually gains more than any index parameter.
| Strengths | Weaknesses |
|---|---|
| First-class hybrid (BM25 + vector) search with one knob | Heavier: Go service, more memory per object than Qdrant |
| Vectoriser modules can remove your embedding pipeline entirely | Those modules couple you to Weaviate's model integrations |
| Typed schema, cross-references, aggregations | Schema changes are real migrations |
| Native multi-tenancy with per-tenant index isolation | The v3 to v4 client rewrite invalidated most tutorial code online |
| Generative search built in | More concepts to learn before the first query runs |
Managed Weaviate charges by stored dimensions with a service-level tier multiplier — the number to compute is objects × dimensions, not gigabytes. Self-hosted, it is again an instance-sizing exercise, with the caveat that its per-object overhead is higher than Qdrant's because of the object store and inverted indexes alongside the vectors.
Side by side
| Pinecone | Qdrant | Weaviate | |
|---|---|---|---|
| Licence | Closed, SaaS only | Apache 2.0 | BSD-3 |
| Run locally | No | Yes (Docker, one command) | Yes (Docker) |
| Language | — | Rust | Go |
| Index | Proprietary, graph-based | HNSW, filter-aware | HNSW, plus flat and dynamic |
| Filtering | Integrated single-stage | Cardinality-aware strategy switch | Inverted index + roaring bitmaps |
| Quantisation | Managed, not exposed | Scalar, product, binary | Scalar, product, binary |
| Hybrid search | Yes (sparse-dense) | Yes (sparse vectors + fusion) | Yes, native BM25 fusion with alpha |
| Multi-tenancy | Namespaces | Payload field convention, or per-collection | Native tenants with isolation |
| Embedding generation | Optional managed models | No, bring your own | Yes, vectoriser modules |
| Ops burden | None | Low to moderate | Moderate |
| Cost shape | Metered usage | Capacity (RAM-driven) | Stored dimensions or capacity |
Sizing and cost, without a price table
Prices change; the arithmetic does not. Work out how much RAM your data needs, then price that.
For 5,000,000 chunks at 768 dimensions with roughly 500 bytes of metadata each:
| Component | Calculation | float32 | int8 quantised |
|---|---|---|---|
| Vectors | 5,000,000 × 768 × bytes | 15.36 GB | 3.84 GB |
| HNSW graph (M = 16) | 5,000,000 × ~138 B | 0.69 GB | 0.69 GB |
| Payload / metadata | 5,000,000 × 500 B | 2.50 GB | 2.50 GB |
| Subtotal | 18.55 GB | 7.03 GB | |
| With 1.5× operating headroom | compaction, query buffers, OS | 27.8 GB | 10.5 GB |
| Instance class needed | 32 GB RAM | 16 GB RAM |
A 32 GB / 4 vCPU cloud instance runs around 0.25 USD per hour on demand, so 0.25 × 730 ≈ 183 USD per month, doubled to about 366 for a replica. The 16 GB option halves it. Add engineering time — realistically a few hours a month once steady, considerably more during the first quarter — and compare against the managed quote for the same capacity. Managed services typically land at 1.5–3× the raw infrastructure cost, which is cheap if it saves you a day a month and expensive if you already run a database team.
Quantisation is usually the largest single lever on your bill: int8 with rescoring cuts vector memory by 75% for one to two points of recall you can win back by oversampling.
Using them through LangChain
All three implement the same retriever interface, so application code is portable even though the setup is not.
1from langchain_pinecone import PineconeVectorStore2from langchain_qdrant import QdrantVectorStore3from langchain_weaviate import WeaviateVectorStore45# Pinecone6store = PineconeVectorStore(index=index, embedding=embeddings,7 namespace="tenant-acme")89# Qdrant10store = QdrantVectorStore(client=client, collection_name="support_docs",11 embedding=embeddings)1213# Weaviate14store = WeaviateVectorStore(client=client, index_name="SupportDoc",15 text_key="text", embedding=embeddings)1617# Identical from here on18retriever = store.as_retriever(search_kwargs={"k": 20})19docs = retriever.invoke("my card got declined")Where the abstraction leaks is filtering. The filter value inside search_kwargs is passed through to the underlying client untranslated, so it is a Mongo-style dict for Pinecone, a models.Filter object for Qdrant, and a Filter builder expression for Weaviate. Any code that filters is engine-specific. If you expect to switch engines, put filter construction behind one function of your own — that function is the entire migration surface, and it is usually under a hundred lines.
Where people get this wrong
Benchmarking without filters. Almost every published comparison measures unfiltered top-k on a static corpus. Real workloads filter on tenant, permission, date and language simultaneously, and that is precisely where the engines diverge. Benchmark with your real filter shapes or your benchmark is fiction.
Filtering on an unindexed field. In Qdrant this silently becomes a payload scan; in Weaviate an unindexed property cannot be filtered efficiently either. The symptom is a query that is fast in development with 10,000 rows and 100× slower in production, with no error anywhere.
Copying recreate_collection from a tutorial. It drops your data. The same trap exists in spirit elsewhere: any "create or reset" convenience call belongs in tests only.
Querying before the index is built. Qdrant returns results immediately after an upsert, but the HNSW graph is built by a background optimiser; queries during that window fall back to exact scan. Your latency measurements will be wrong in one direction and your recall wrong in the other. Wait for the collection status to go green before measuring anything.
Treating namespaces, tenants and collections as interchangeable. They differ in isolation and cost. A namespace or tenant per customer is cheap and scales to thousands; a collection per customer means a separate HNSW graph each, and a few thousand of those will exhaust memory on any machine.
Leaving the Weaviate v4 client open. It holds a gRPC connection; forgetting client.close() in a request handler leaks connections until the service falls over. Use it as a context manager where you can.
How to actually decide
Work down this list and stop at the first line that describes you.
- Under 100,000 vectors, and you already run PostgreSQL. Use pgvector. One system, transactional consistency with your application data, and joins for free. Revisit at a million.
- No infrastructure team and no data-residency constraint. Pinecone. You will pay a premium and get all of your engineering time back.
- Heavy metadata filtering, per-tenant permissions, or cost sensitivity at scale. Qdrant. The filtering behaviour and the quantisation options are the differentiators, and self-hosting caps your bill.
- You want hybrid keyword-plus-vector search out of the box, or you want the database to do the embedding. Weaviate.
- Billions of vectors, dedicated platform team. Look at Milvus and the DiskANN family; the three here are all viable at that scale but the operational calculus changes.
Then do the thing that actually settles it: load one million of your vectors into your top two candidates, with your real metadata and your real filter shapes, and measure recall@10 against a flat index plus p95 latency under concurrent load. It takes about a day. Every team that skips it makes the decision on marketing pages and discovers the filtering behaviour in production, which costs considerably more than a day.