RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is the purpose of embedding text and storing it in a vector store?


What you need to know

What "close in meaning" looks like

Using bge-small-en-v1.5 (384 dimensions), the query "how do I reset my password" scored these cosine similarities:

PassageCosine
"Account recovery steps: open the login page, choose 'Forgot password'…"0.771
"To change your delivery address, open My Orders…"0.491
"Our cafeteria is open from 8 am to 8 pm on weekdays."0.378

The right passage wins clearly. Notice that even the cafeteria sentence scores 0.378, not 0. Scores are only meaningful relative to each other and to the same model; each model has its own typical range.

What the vector store adds

A vector store is more than a list of vectors. It provides:

  • Storage of vector, text and metadata together, under a stable ID.
  • An index for fast approximate nearest-neighbour (ANN) search.
  • Filters on metadata (tenant, document type, date) applied during search.
  • Updates and deletes as documents change.

Why an index is needed

Brute-force search compares the query with every vector. For 5 million chunks of 1,024 dimensions, that is about 5 billion multiply-adds per query, and the vectors take about 20 GB of memory as 32-bit floats. That can work for a batch job, but not for many users per second. An HNSW index typically visits only a few thousand vectors per query. The next lessons explain how HNSW and IVF do that.

Where dense embeddings are weak

Embeddings capture general meaning and smooth over exact strings. "Error E-4021" and "Error E-4012" produce almost the same vector. So do two SKUs, two policy numbers or two surnames. BM25 keyword search matches those exactly, which is why hybrid search (covered in the retrieval section) is the production default.

A real-life example

An e-commerce product Q&A system gets the question "is this safe to put in the dishwasher?". The seller's listing says "Top-rack washable. Avoid abrasive cleaners." Keyword search finds nothing: no shared words except "the". Dense search ranks that sentence first, because the embedding model learned that "dishwasher safe" and "top-rack washable" mean nearly the same.

The same week, a shopper asks about model "WX-220B". Dense search returns chunks about "WX-220A" and "WX-230B" with nearly equal scores, because the model sees them as the same kind of thing. BM25 finds the exact "WX-220B" chunk. The team runs both and fuses the results.

Follow-up questions to expect

  • "Can I store embeddings in Postgres?" — Yes, with the pgvector extension, which supports HNSW and IVFFlat indexes. For up to a few million vectors it is often the simplest option.
  • "Why not just use a Python list and numpy?" — For a few thousand chunks that is fine and exact. Beyond that you need an index, persistence, filters and concurrent updates.
  • "Do embeddings work across languages?" — Multilingual models map, for example, Hindi and English text with the same meaning to nearby vectors, so an English query can find a Hindi passage.