Course Content
RAG Systems
12 sections · 66 lessons
How does the choice of embedding model affect RAG performance?
What you need to know
What to compare
| Factor | Why it matters |
|---|---|
| Domain fit | General models can struggle with legal, medical or code text, and with Indian names and terms. |
| Max input | 256 or 512 tokens forces small chunks; 8,192 allows larger ones. |
| Dimensions | 384, 768, 1,024, 1,536, 3,072. More dimensions often separate meanings slightly better, but memory and search time grow in proportion. |
| Matryoshka support | Some models are trained so you can keep only the first 256 or 512 numbers of each vector with a small quality loss, cutting memory. |
| Prefixes | E5 models expect query: and passage: ; BGE English models suggest an instruction before queries. Getting this wrong can quietly lower recall. |
| Languages | Needed if users ask in Hindi, Tamil or mixed "Hinglish". |
| Hosting and cost | An API is simple; a self-hosted open model keeps data in-house and has no per-token fee. |
Examples in common use in 2026 include hosted models from OpenAI (text-embedding-3-small and -large), Cohere, Voyage and Google, and open models such as BGE-M3, the Qwen3-Embedding family, E5 and Nomic. Rankings change often, so test rather than trust a list.
Beyond one vector per chunk
- Sparse and multi-vector outputs. BGE-M3 can produce a dense vector, sparse keyword weights and per-token vectors from one pass, which helps hybrid search.
- Late-interaction models such as ColBERT keep one small vector per token and score a query by matching each query token to its best document token. They are more accurate on many tasks than single-vector models, at the cost of a much larger index.
Test on your own data
1for name, prefix in [("sentence-transformers/all-MiniLM-L6-v2", ""),2 ("BAAI/bge-small-en-v1.5",3 "Represent this sentence for searching relevant passages: ")]:4 model = SentenceTransformer(name)5 D = model.encode(chunks, normalize_embeddings=True)6 Q = model.encode([prefix + q for q, _ in evalset], normalize_embeddings=True)7 top = np.argsort(-(Q @ D.T), axis=1)[:, :5]8 recall5 = np.mean([any(g in chunks[i] for i in row)9 for row, (_, g) in zip(top, evalset)])10 print(name, recall5)This is the same harness as the chunk-size sweep: labelled questions, the text that must be found, and recall at the k you will use. Keep chunking fixed while you compare models, or you will not know which change helped.
A real-life example
A legal-contract search tool starts with a popular general-purpose hosted model. On 100 labelled questions from associates, recall@10 is acceptable for plain-English questions but weak for questions using legal terms of art ("carve-outs from the indemnity", "MAC clause").
The team shortlists three models and runs the harness above on the same 100 questions and the same clause-level chunks. One open model does noticeably better on the legal-term questions and is self-hostable, which the firm's clients prefer. Its vectors are 1,024-dimensional instead of 1,536, which also cuts index memory by a third. The test took two days; re-embedding 700,000 clauses took an afternoon on one GPU.
Follow-up questions to expect
- "Should I fine-tune the embedding model?" — When a strong general model still misses domain terms and you have a few thousand query-passage pairs (real or LLM-generated then checked), fine-tuning often gives a clear lift.
- "Does a bigger model always win?" — No. On your data a small model can match a larger one, at a fraction of the cost and latency. Only your eval set can tell.
- "How do you handle Hinglish queries?" — Test multilingual models on real mixed-language queries; also consider normalising transliterated text before embedding.