RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How does the choice of embedding model affect RAG performance?


What you need to know

What to compare

FactorWhy it matters
Domain fitGeneral models can struggle with legal, medical or code text, and with Indian names and terms.
Max input256 or 512 tokens forces small chunks; 8,192 allows larger ones.
Dimensions384, 768, 1,024, 1,536, 3,072. More dimensions often separate meanings slightly better, but memory and search time grow in proportion.
Matryoshka supportSome models are trained so you can keep only the first 256 or 512 numbers of each vector with a small quality loss, cutting memory.
PrefixesE5 models expect query: and passage: ; BGE English models suggest an instruction before queries. Getting this wrong can quietly lower recall.
LanguagesNeeded if users ask in Hindi, Tamil or mixed "Hinglish".
Hosting and costAn API is simple; a self-hosted open model keeps data in-house and has no per-token fee.

Examples in common use in 2026 include hosted models from OpenAI (text-embedding-3-small and -large), Cohere, Voyage and Google, and open models such as BGE-M3, the Qwen3-Embedding family, E5 and Nomic. Rankings change often, so test rather than trust a list.

Beyond one vector per chunk

  • Sparse and multi-vector outputs. BGE-M3 can produce a dense vector, sparse keyword weights and per-token vectors from one pass, which helps hybrid search.
  • Late-interaction models such as ColBERT keep one small vector per token and score a query by matching each query token to its best document token. They are more accurate on many tasks than single-vector models, at the cost of a much larger index.

Test on your own data

Python
for name, prefix in [("sentence-transformers/all-MiniLM-L6-v2", ""),                     ("BAAI/bge-small-en-v1.5",                      "Represent this sentence for searching relevant passages: ")]:    model = SentenceTransformer(name)    D = model.encode(chunks, normalize_embeddings=True)    Q = model.encode([prefix + q for q, _ in evalset], normalize_embeddings=True)    top = np.argsort(-(Q @ D.T), axis=1)[:, :5]    recall5 = np.mean([any(g in chunks[i] for i in row)                       for row, (_, g) in zip(top, evalset)])    print(name, recall5)

This is the same harness as the chunk-size sweep: labelled questions, the text that must be found, and recall at the k you will use. Keep chunking fixed while you compare models, or you will not know which change helped.

A real-life example

A legal-contract search tool starts with a popular general-purpose hosted model. On 100 labelled questions from associates, recall@10 is acceptable for plain-English questions but weak for questions using legal terms of art ("carve-outs from the indemnity", "MAC clause").

The team shortlists three models and runs the harness above on the same 100 questions and the same clause-level chunks. One open model does noticeably better on the legal-term questions and is self-hostable, which the firm's clients prefer. Its vectors are 1,024-dimensional instead of 1,536, which also cuts index memory by a third. The test took two days; re-embedding 700,000 clauses took an afternoon on one GPU.

Follow-up questions to expect

  • "Should I fine-tune the embedding model?" — When a strong general model still misses domain terms and you have a few thousand query-passage pairs (real or LLM-generated then checked), fine-tuning often gives a clear lift.
  • "Does a bigger model always win?" — No. On your data a small model can match a larger one, at a fraction of the cost and latency. Only your eval set can tell.
  • "How do you handle Hinglish queries?" — Test multilingual models on real mixed-language queries; also consider normalising transliterated text before embedding.