Embeddings and Semantic Search

Sentence Transformers - Turning Text into Meaningful Vectors


A recipe site ships a search box. A user types "quick meals for kids". The database contains a recipe titled "5-Minute Snacks Children Will Love" — precisely what they want. It does not appear. Not at rank 1, not at rank 50, not at all.

Line up the words and the reason is obvious. "quick" is not "5-minute", "meals" is not "snacks", "kids" is not "children". Zero tokens overlap. To an engine that compares strings, these two texts have nothing in common, so the recipe is not merely ranked low — it is never a candidate.

The instinct is to patch it with a synonym list: kids→children, quick→fast→5-minute, meals→snacks. That works for exactly the phrases you thought of. Then someone types "something my toddler will eat in a hurry" and you are back to zero. The number of ways to phrase one idea is unbounded; you cannot enumerate your way out.

What you want is a system where those two phrases end up as points sitting next to each other, and "the fall of the Western Roman Empire" ends up far away — with nobody writing a single synonym rule. That is what an embedding is: a list of numbers, produced by a trained model, positioned so that closeness in the numbers means closeness in meaning.

From a sentence to one 384-dimensional vector"quickmeals for kids"Tokenise: 6subword tokensTransformer: 6contextual vectorsMean-poolover themask, not the padNormalise tounit lengthPool over the attention mask; averaging in the padding drags every short sentence toward the same point.
One-hot vectors make every pair of distinct words equally unrelated — pooling is the step that turns a variable-length sentence into a single point where distance finally means something.

Why literal representations cannot be rescued

One-hot encoding: every word equally unrelated to every other word

The simplest way to turn words into numbers is one-hot encoding: give each word its own slot, and mark it with a 1.

Text
Vocabulary: ["dog", "puppy", "car", "seven"]"dog"   -> [1, 0, 0, 0]"puppy" -> [0, 1, 0, 0]"car"   -> [0, 0, 1, 0]"seven" -> [0, 0, 0, 1]

Measure relatedness with a dot product — multiply position by position, add up the results:

Text
dog . puppy = (1x0) + (0x1) + (0x0) + (0x0) = 0dog . car   = (1x0) + (0x0) + (0x1) + (0x0) = 0dog . seven = (1x0) + (0x0) + (0x0) + (0x1) = 0

All three answers are 0. "dog" is exactly as related to "puppy" as to "seven". This is structural, not a tuning problem: two distinct one-hot vectors never share a non-zero position, so their dot product is always zero. Every pair of distinct words is perpendicular, and the geometry has no room to express "these two mean similar things".

It also scales appallingly: a 100,000-word vocabulary needs 100,000-dimensional vectors of which 99,999 entries are zero — 400 KB per word to encode one integer's worth of information.

Word embeddings fixed half the problem

Word2Vec and GloVe replaced that sparse slot with a learned dense vector — perhaps 300 non-zero numbers — trained so that words appearing in similar contexts get similar vectors. Now "dog" and "puppy" have a high dot product and "dog" and "seven" do not, and famously king - man + woman lands near queen: the vectors encode relationships, not just identity.

But a word vector represents a word. To represent a sentence, the obvious move is to average the word vectors — and that move throws away everything that word order does:

Text
"The movie was not good, it was great""The movie was not great, it was good"Identical word multiset -> identical average -> identical vector.Opposite readings.

Averaging also flattens negation and cannot disambiguate: "bank" gets one fixed vector whether the sentence is about a river or a mortgage.

What a sentence embedding is

A sentence embedding is a single dense vector — typically 384 to 3072 numbers — representing the meaning of an entire sentence or short document, produced by a model that reads the whole sequence at once. Two texts that mean similar things land close together even with no shared vocabulary.

An embedding does not store words. It stores a position. The whole design bet is that a well-chosen position in a few hundred dimensions can carry more useful information about a sentence than the sentence's own words can.

Inside a sentence transformer: four steps and some arithmetic

Sentence-BERT (SBERT), and the sentence-transformers library built around it, takes a BERT-style transformer — which naturally outputs one vector per token — and adds two steps that collapse those into one vector per text.

Text
"The cat sat"  [1] Tokenizer     -> [CLS] the cat sat [SEP]  [2] Transformer   -> one 384-number vector PER TOKEN (context-aware)  [3] Pooling       -> combine token vectors into ONE vector  [4] L2 normalise  -> rescale that vector to length exactly 1  =>  a single 384-number vector, unit length

Step 1, tokenisation splits text into subword pieces and adds two markers, [CLS] at the front and [SEP] at the end. Step 2, the transformer, produces a vector per token that depends on the surrounding tokens — which is why "bank" in "river bank" and "bank" in "savings bank" come out different. The ambiguity that killed word embeddings is resolved here.

Step 3, pooling, is where the collapse happens, and it is worth doing by hand. Take these five token vectors, using 4 dimensions instead of 384 so the arithmetic fits on a page:

Text
[CLS] = [ 0.02,  0.11, -0.05,  0.20]the   = [ 0.10,  0.20, -0.10,  0.05]cat   = [ 0.80,  0.10,  0.30,  0.40]sat   = [ 0.20,  0.70, -0.20,  0.10][SEP] = [ 0.03,  0.09,  0.00,  0.15]

Mean pooling averages each position across all five tokens:

Text
dim 0: (0.02 + 0.10 + 0.80 + 0.20 + 0.03) / 5 = 1.15 / 5 =  0.230dim 1: (0.11 + 0.20 + 0.10 + 0.70 + 0.09) / 5 = 1.20 / 5 =  0.240dim 2: (-0.05 - 0.10 + 0.30 - 0.20 + 0.00) / 5 = -0.05 / 5 = -0.010dim 3: (0.20 + 0.05 + 0.40 + 0.10 + 0.15) / 5 = 0.90 / 5 =  0.180pooled = [0.230, 0.240, -0.010, 0.180]

Step 4, L2 normalisation, divides that vector by its own length. The length is the square root of the sum of the squares:

Text
||pooled|| = sqrt(0.230^2 + 0.240^2 + (-0.010)^2 + 0.180^2)           = sqrt(0.0529 + 0.0576 + 0.0001 + 0.0324)           = sqrt(0.1430)           = 0.37815normalised = [0.230/0.37815, 0.240/0.37815, -0.010/0.37815, 0.180/0.37815]           = [0.6082, 0.6347, -0.0264, 0.4760]check: 0.6082^2 + 0.6347^2 + 0.0264^2 + 0.4760^2     = 0.3699 + 0.4028 + 0.0007 + 0.2266 = 1.0000  ✓

Every embedding now sits on a sphere of radius 1. Only direction carries information; magnitude is gone by construction. That is deliberate: comparing two embeddings reduces to one dot product, and a long document cannot outrank a short one just for being long.

Pooling strategies, compared on the same numbers

Mean pooling is the default; running each strategy on the five token vectors above shows why the choice matters.

StrategyResult on the example aboveWhat it doesWhen it is right
mean[0.230, 0.240, -0.010, 0.180]Averages every token, so every word contributesDefault for similarity and retrieval; almost always the best starting point
max[0.800, 0.700, 0.300, 0.400]Takes the largest value per dimension — here entirely dominated by "cat" and "sat"Rarely; useful when one salient keyword should dominate the representation
cls[0.020, 0.110, -0.050, 0.200]Uses only the [CLS] token vector, ignoring all real wordsOnly when the model was fine-tuned to make [CLS] a summary; on an untuned model it is near-arbitrary

Notice how far apart those three results are. Max pooling produced a value in dimension 0 that is 3.5x larger than mean pooling's, because a single token carried it; [CLS] pooling produced something almost unrelated to either.

The common misconception is that "[CLS] is the sentence embedding". That is true for BERT models fine-tuned on classification, where [CLS] was explicitly trained to summarise. For similarity embeddings, mean pooling wins because it draws signal from every word rather than betting everything on one special token.

You can override the default when you need to:

Python
from sentence_transformers import SentenceTransformerfrom sentence_transformers.sentence_transformer.modules import Transformer, Pooling, Normalizebackbone = Transformer('bert-base-uncased')pooling = Pooling(backbone.get_embedding_dimension(),                  pooling_mode='mean')          # 'mean' | 'max' | 'cls'model = SentenceTransformer(modules=[backbone, pooling, Normalize()])print(model.encode(["This is a sentence."]).shape)   # (1, 768)

This is the sentence-transformers 6 layout. Older releases import the same classes from sentence_transformers.models and call get_word_embedding_dimension(); version 6 still accepts those, with a deprecation warning.

Your first embeddings

Bash
pip install sentence-transformers torch numpy scikit-learn
Python
from sentence_transformers import SentenceTransformermodel = SentenceTransformer('all-MiniLM-L6-v2')   # downloads on first use (~90 MB)sentences = [    "The cat sits on the mat.",    "A feline rests on the rug.",    "The weather is sunny today.",]embeddings = model.encode(sentences, normalize_embeddings=True)print(embeddings.shape)          # (3, 384)print(embeddings[0][:8])         # first 8 of 384 numbersprint((embeddings[0] ** 2).sum())  # 1.0000001 -- unit length, as designed

(3, 384) means three sentences, each turned into 384 floating-point numbers. Sentences 0 and 1 share only the word "the" yet describe the same scene, so a working model puts them close together. The numbers are individually meaningless — no dimension is "the cat dimension". Only relative positions mean anything.

Choosing a model, and the gotchas that come with each

ModelDimsRunsNotes and quirks
all-MiniLM-L6-v2384Local, CPU-friendlyThe sensible default. Six layers, ~22M parameters, max input 256 tokens.
all-mpnet-base-v2768LocalBetter quality; roughly 4-5x slower and 2x the storage.
multi-qa-mpnet-base-dot-v1768LocalTuned for asymmetric use: short question against long passage.
BAAI/bge-large-en-v1.51024Local, GPU preferredStrong retrieval. An instruction prefix on short queries helps a little (optional for v1.5); never on documents.
intfloat/e5-large-v21024Local, GPU preferredRequires "query: " / "passage: " prefixes — degrades silently without them.
nomic-embed-text-v1.5up to 768Local8,192-token context, so it handles long documents. Trained for truncation (down to 64). Requires "search_query: " / "search_document: " prefixes, and loads custom code (trust_remote_code=True).
text-embedding-3-smallup to 1536Hosted APIOpenAI's current small model (it replaced text-embedding-ada-002). No GPU needed, pay per token, 8,192-token input; supports a dimensions parameter.
text-embedding-3-largeup to 3072Hosted APIHighest quality in that family; 8x MiniLM's storage.
paraphrase-multilingual-mpnet-base-v2768Local50+ languages share one space, so English and German text compare directly.

Six things drive the choice.

  1. Latency and storage budget. 384 dimensions at 4 bytes each is 1,536 bytes per document; 3072 dimensions is 12,288 bytes — eight times the disk and eight times the arithmetic per comparison.
  2. Symmetric or asymmetric. Symmetric compares like with like (duplicate tickets). Asymmetric compares a short query against a long passage. Models tuned for one are measurably worse at the other.
  3. Input length. Every model has a hard max_seq_length; MiniLM stops at 256 tokens and discards the rest silently.
  4. Language. A monolingual English model given German text produces a vector — a confident, meaningless one.
  5. Prefix requirements. The most common silent failure in this whole area.
  6. Hardware. A 1024-dimension model on a CPU-only box can be an order of magnitude slower per document.
Python
model = SentenceTransformer('intfloat/e5-large-v2')# RIGHT -- prefixes as the model was trained to expectq = model.encode("query: What is the capital of France?")p = model.encode("passage: Paris is the capital and largest city of France.")# WRONG -- no error, no warning, just quietly worse rankings foreverq_bad = model.encode("What is the capital of France?")

Nothing crashes in the wrong version. Search returns worse results, with no signal telling you why.

Recent versions of sentence-transformers also have model.encode_query() and model.encode_document(), which add the right prefix automatically, but only when the model's saved config defines one. E5, BGE and Nomic ship no prefixes in their configs, so for them those methods add nothing. Write the prefix yourself, or pass it explicitly with model.encode(text, prompt="query: ").

To shortlist candidates objectively, use MTEB (Massive Text Embedding Benchmark), which scores models across dozens of retrieval, clustering, and classification tasks. But the ranking that matters is the one measured on your documents and your queries.

Matryoshka embeddings: shrinking a vector without retraining

Work out the storage bill before picking a dimensionality.

DimensionsBytes per vector (float32)1 million documents100 million documents
2561,0241.02 GB102 GB
3841,5361.54 GB154 GB
7683,0723.07 GB307 GB
10244,0964.10 GB410 GB
307212,28812.29 GB1,229 GB

At 100 million documents, the gap between 1024 and 256 dimensions is 308 GB of RAM — the difference between one machine and a cluster. And since every comparison costs one multiply-add per dimension, 256 dimensions is also 4x cheaper to search.

The obvious fix, using a smaller model, usually costs accuracy across the board. Matryoshka Representation Learning (MRL), named after Russian nesting dolls, offers something better: train one model so that prefixes of its output are themselves valid embeddings.

Text
[ x1 ... x256 | x257 ... x512 | x513 ... x1024 ]  \___________/  usable 256-dim embedding on its own  \__________________________/  usable 512-dim embedding  \_____________________________________________/  full 1024-dim

MRL achieves this by changing the training objective: the loss is computed on the full vector and on several truncated prefixes at once, forcing the model to pack the most important information into the earliest dimensions. Ordinary training applies no such pressure, so information spreads arbitrarily across all dimensions.

Python
import numpy as npfrom sentence_transformers import SentenceTransformermodel = SentenceTransformer('mixedbread-ai/mxbai-embed-large-v1')   # MRL-trained, 1024-dfull = model.encode("Matryoshka embeddings can be shrunk without retraining.",                    normalize_embeddings=True)print(full.shape)                       # (1024,)truncated = full[:256]                  # keep the first 256 numbers# Slicing changes the vector's length, so unit-norm is broken -- restore itprint(np.linalg.norm(truncated))        # about 0.48, no longer 1.0truncated = truncated / np.linalg.norm(truncated)print(np.linalg.norm(truncated))        # 1.0

SentenceTransformer(name, truncate_dim=256) does the slicing for you on every encode() call; add normalize_embeddings=True and it re-normalises as well.

That re-normalisation is not optional. Truncating a unit vector leaves something shorter than 1 — how much shorter depends on how much energy lived in the discarded tail. Skip it, and your dot-product scores are scaled by a different amount for every document. Ranking becomes nonsense.

Truncation only works on models trained with MRL. Chopping all-MiniLM-L6-v2 from 384 to 128 dimensions is not a compression — it is deleting two-thirds of the information at random, and retrieval quality collapses accordingly.

Encoding a large collection without melting your laptop

Embedding three sentences is instant. Embedding 500,000 support tickets needs thought.

Python
from sentence_transformers import SentenceTransformermodel = SentenceTransformer('all-MiniLM-L6-v2')embeddings = model.encode(    documents,    batch_size=64,              # texts pushed through the model at once    normalize_embeddings=True,  # do the L2 step here, once, for everything    show_progress_bar=True,    device='cuda',              # omit to auto-select; 'cpu' is fine for MiniLM)

Batch size is the main lever, and its effect has a consistent shape:

Batch sizeRelative throughputWhy
11.0x (baseline)Per-call overhead dominates the actual maths
8~4xOverhead amortised; matrix multiplications get usefully wide
32~6xHardware near saturation
64~6.5xDiminishing returns begin
512+may run out of GPU memoryActivations scale with batch size x sequence length

The lesson is not "use 512": going from 1 to 32 buys almost everything, and going far beyond buys little speed and, on a small GPU, a crash. Sorting documents by length first helps too, because each batch is padded to its longest member — one 250-token document among 63 ten-token ones wastes most of the batch on padding.

The silent truncation trap

This one costs people weeks.

Python
print(model.max_seq_length)   # 256 for all-MiniLM-L6-v2long_article = open('support_article.txt').read()   # ~900 words ~= 1,200 tokens# 256 / 1200 = 21% kept. The other 79% is dropped -- no exception, no warning.emb = model.encode(long_article)

Your embedding represents the first fifth of the article. If the answer lives in the conclusion, it is not in the vector, and no amount of search tuning will find it. The fix is a long-context model (nomic-embed-text-v1.5 handles 8,192 tokens) or splitting the document into overlapping chunks of a few hundred tokens and embedding each separately.

Judging whether an embedding model is actually good

Four properties matter, and they pull against each other:

  1. Semantic preservation — texts that mean the same thing land close together.
  2. Discriminativeness — texts that mean different things land clearly apart. A model mapping everything to nearly one point scores perfectly on the first property and is useless.
  3. Efficiency — encoding throughput and storage cost you can afford.
  4. Robustness — quality holds on your domain's vocabulary, not just Wikipedia prose.

The standard measurement is Spearman rank correlation between the model's scores and human judgements on the same pairs: when people rank 200 sentence pairs from most to least similar, does the model produce the same ordering? 1.0 is perfect agreement, 0.0 is none.

Python
from sentence_transformers.sentence_transformer.evaluation import EmbeddingSimilarityEvaluator# Pairs from YOUR domain, scored 0-1 by people who know the domains1   = ["Card declined at checkout", "How do I reset my password"]      # ...200 of theses2   = ["Payment failed during purchase", "Where can I change my email"]gold = [0.90, 0.35]results = EmbeddingSimilarityEvaluator(s1, s2, gold)(model)   # a dict of metricsprint(results["spearman_cosine"])                             # higher is better

Fifty to two hundred hand-scored pairs from your own data tell you more than any leaderboard. Benchmarks show which models are generally competent; only your pairs show which one understands "chargeback" the way your users mean it.

What people build with these vectors

Ranking documents against a query

Python
import numpy as npfrom sentence_transformers import SentenceTransformermodel = SentenceTransformer('all-MiniLM-L6-v2')documents = ["Python is a programming language",             "Java is also a programming language",             "The weather is nice today",             "I like to code in Python"]doc_emb = model.encode(documents, normalize_embeddings=True)q_emb   = model.encode("Programming languages like Python", normalize_embeddings=True)# Both sides are unit length, so a plain dot product IS cosine similarityscores = doc_emb @ q_embfor rank, i in enumerate(np.argsort(scores)[::-1][:3], start=1):    print(f"{rank}. {scores[i]:.3f}  {documents[i]}")

One matrix multiply scores every document at once; for 100,000 documents NumPy handles it in milliseconds.

Grouping documents with no labels

Python
from sklearn.cluster import KMeanslabels = KMeans(n_clusters=5, n_init=10, random_state=42).fit_predict(doc_emb)

Because unit-length vectors make Euclidean distance a monotonic function of cosine similarity, k-means on normalised embeddings groups by meaning — which is how a support team discovers that 30% of last month's tickets were one billing complaint phrased forty ways.

Finding near-duplicates

Python
from sklearn.metrics.pairwise import cosine_similaritydocs = ["The quick brown fox jumps over the lazy dog",        "A fast brown fox leaps over a sleepy dog",        "Python is great for data science"]sim = cosine_similarity(model.encode(docs))for i in range(len(docs)):    for j in range(i + 1, len(docs)):        if sim[i][j] > 0.70:            print(f"{sim[i][j]:.2f}  {docs[i]!r} / {docs[j]!r}")

Documents 0 and 1 share four words out of nine yet describe an identical event. Keyword overlap scores them around 0.4; an embedding model scores them near 0.85.

Fine-tuning when the general model is not enough

Pre-trained models learned general English. They do not know that "SEV-2" and "priority incident" are the same thing at your company, or that in cardiology "MI" means myocardial infarction, not Michigan. Fine-tuning teaches an existing model your vocabulary without starting over.

Bash
pip install sentence-transformers datasets accelerate
Python
from datasets import Datasetfrom sentence_transformers import SentenceTransformer, SentenceTransformerTrainerfrom sentence_transformers.sentence_transformer.losses import CosineSimilarityLossmodel = SentenceTransformer('all-MiniLM-L6-v2')# Option A: explicit similarity labels, 0.0 to 1.0labelled = Dataset.from_dict({    "sentence1": ["myocardial infarction", "myocardial infarction"],    "sentence2": ["heart attack", "bone fracture"],    "score":     [0.95, 0.05],})trainer = SentenceTransformerTrainer(model=model, train_dataset=labelled,                                     loss=CosineSimilarityLoss(model))trainer.train()model.save_pretrained('cardiology-embeddings')

The trainer reads the columns in order (two texts, then the label) and uses Hugging Face's training loop underneath. Older tutorials call model.fit(...) with InputExample objects; that still exists as a legacy wrapper, but it now needs datasets installed too and new code should use the trainer.

Hand-labelling scores is slow. A cheaper option usually works better:

Python
from sentence_transformers import SentenceTransformerTrainingArgumentsfrom sentence_transformers.sentence_transformer.losses import MultipleNegativesRankingLoss# Option B: positive PAIRS only -- no scores neededpairs = Dataset.from_dict({    "anchor":   ["myocardial infarction", "how do I reset my password"],    "positive": ["heart attack", "password reset instructions"],})args = SentenceTransformerTrainingArguments(    output_dir="cardiology-embeddings", num_train_epochs=1,    per_device_train_batch_size=32,)trainer = SentenceTransformerTrainer(model=model, args=args, train_dataset=pairs,                                     loss=MultipleNegativesRankingLoss(model))trainer.train()

MultipleNegativesRankingLoss treats the other items in the same batch as negatives. With a batch size of 32, each positive pair is trained against 31 automatically generated negatives — you supplied 32 labels and the loss manufactured 992 comparisons from them. That is why larger batches help with this loss specifically, and why a few thousand naturally-occurring pairs (a question and its accepted answer, a ticket and its resolution article) beat a few hundred painstakingly scored ones.

What this means when you build something

The failures below all show up in real projects, none raise an exception, and every one degrades results quietly. That is what makes them expensive.

SymptomLikely causeFix
Search results are plausible but consistently mediocre, and always have beenModel needs a prefix (E5, BGE) and you never added oneRead the model card; add "query: " / "passage: " exactly as specified
Long documents never match queries about their later contentSilent truncation at max_seq_lengthPrint model.max_seq_length; chunk documents or switch to a long-context model
Similarity scores look random after a storage optimisationTruncated a non-MRL model, or forgot to re-normalise after slicingOnly truncate MRL-trained models, and always divide by the new norm
Quality collapsed after "upgrading" the modelQueries encoded with model B against an index built with model ARe-embed the entire corpus on any model change; there is no partial migration
Scores hover around 0.9 for everything, relevant or notPoor discriminativeness on your domainJudge the score distribution, not absolute values; consider fine-tuning
Encoding 100k documents takes hoursbatch_size=1, or unsorted inputs padding every batchSet batch_size=32+, sort by length, use a GPU if available

Two operational rules follow, and both are worth writing down before you start.

Pin the model name in configuration, never in scattered code. Every embedding belongs to exactly one model's coordinate system. Two models can both output 768 numbers and be completely incomparable, because there is no reason dimension 412 means the same thing in both. The day someone changes the model name in one file and not another, your search starts comparing coordinates from two different universes — and it returns results, ranked confidently, and wrong.

Store the model name alongside the vectors. A vector on disk with no record of what produced it is unusable data. Write the model identifier, the dimensionality, and whether the vectors are normalised into the same store as the embeddings. Six months later, when someone adds a million documents to your index, that one line of metadata is the difference between a five-minute job and re-embedding everything.