LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do vector databases differ from traditional databases in LLM systems?


Hybrid retrieval for "status of form 12-B" in TamilKeywordsearch (BM25)Vector search,filtered to TamilFuse thetwo rankingsCross-encoderrerankTop 5 tothe promptRecall at 5 rose from 78 to 91 percent on 500 questions.
Vectors find meaning, keywords find the exact form number — fusing both is what put 12-B above its look-alikes.

What you need to know

Side by side

Traditional database

  • Exact match and range queries
  • B-tree and hash indexes
  • Same query, same rows, every time
  • Cheap inserts, updates and deletes

Vector database

  • Similarity (cosine, dot product) queries
  • HNSW graph or IVF cluster indexes
  • Ranked top-k, approximate by design
  • Index updates cost more; memory-heavy

What that means in practice

  • Tunable recall. HNSW's ef_search (or IVF's nprobe) trades speed for recall. Higher values find more true neighbours, more slowly.
  • Memory. One million 1,536-dimension float32 vectors take about 6.1 GB before index overhead; HNSW wants that in RAM. Half-precision, int8 or binary quantization of vectors, and smaller embedding dimensions cut this a lot.
  • Filtering is the hard part. "Nearest chunks where language = Tamil and scheme = pension" — filtering after the vector search can return too few results; filtering first can be slow. Databases differ most here.
  • Updates and deletes are more expensive on graph indexes; plan re-indexing.

Hybrid search

Embeddings capture meaning but are weak at exact tokens: form numbers, product codes, names, acronyms. Hybrid search runs keyword search (BM25) and vector search together, merges the results (for example with reciprocal rank fusion) and reranks the top candidates with a cross-encoder. It reliably beats vector-only retrieval on real queries.

Choosing

  • pgvector — the right first choice if you run Postgres: one system, transactions, joins with your metadata, HNSW indexes, and support for iterative index scans that help filtered queries in recent versions.
  • Qdrant, Weaviate, Milvus — dedicated engines with strong filtering, quantization and scale-out.
  • Managed services (Pinecone and cloud-native offerings) — least operations.
SQL
CREATE EXTENSION IF NOT EXISTS vector;CREATE TABLE chunks (  id bigserial PRIMARY KEY,  scheme text, lang text, body text,  embedding vector(1024));CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);SET hnsw.ef_search = 100;SELECT id, bodyFROM chunksWHERE lang = 'ta' AND scheme = 'pension'ORDER BY embedding <=> $1      -- cosine distance to the query vectorLIMIT 5;

<=> is pgvector's cosine-distance operator. The WHERE filter is applied alongside the approximate index scan, which is why filtered queries need testing for recall.

A real-life example

A state government chatbot retrieves from 400,000 chunks of scheme documents in 12 languages, with 1,024-dimension embeddings. That is about 1.6 GB of float32 vectors (0.8 GB in half precision) — small enough for pgvector on the Postgres server they already run, with scheme metadata in the same tables.

Two problems appear in testing. First, questions like "status of form 12-B" return chunks about similar forms; adding BM25 keyword search with rank fusion puts the exact form first. Second, Tamil-only queries sometimes return only two results instead of five, because the language filter removed most of the top candidates of the approximate search; enabling iterative index scans and raising ef_search fixes it. Retrieval recall@5 on their 500-question eval rises from 78% to 91%, and the system still runs on one familiar database.

Follow-up questions to expect

  • "When would you leave pgvector?" — At tens or hundreds of millions of vectors, heavy filtered traffic, or when vector search load starts hurting the transactional database.
  • "How do you measure retrieval quality?" — Recall@k and MRR on a labelled set of questions with known correct chunks, tracked per language and per filter.
  • "What happens when you change the embedding model?" — All vectors must be re-embedded and re-indexed; run old and new side by side and switch after evals.