Course Content
Embeddings and Semantic Search
3 sections · 5 lessons
Sentence Transformers - Turning Text into Meaningful Vectors
A recipe site ships a search box. A user types "quick meals for kids". The database contains a recipe titled "5-Minute Snacks Children Will Love" — precisely what they want. It does not appear. Not at rank 1, not at rank 50, not at all.
Line up the words and the reason is obvious. "quick" is not "5-minute", "meals" is not "snacks", "kids" is not "children". Zero tokens overlap. To an engine that compares strings, these two texts have nothing in common, so the recipe is not merely ranked low — it is never a candidate.
The instinct is to patch it with a synonym list: kids→children, quick→fast→5-minute, meals→snacks. That works for exactly the phrases you thought of. Then someone types "something my toddler will eat in a hurry" and you are back to zero. The number of ways to phrase one idea is unbounded; you cannot enumerate your way out.
What you want is a system where those two phrases end up as points sitting next to each other, and "the fall of the Western Roman Empire" ends up far away — with nobody writing a single synonym rule. That is what an embedding is: a list of numbers, produced by a trained model, positioned so that closeness in the numbers means closeness in meaning.
Why literal representations cannot be rescued
One-hot encoding: every word equally unrelated to every other word
The simplest way to turn words into numbers is one-hot encoding: give each word its own slot, and mark it with a 1.
Vocabulary: ["dog", "puppy", "car", "seven"]"dog" -> [1, 0, 0, 0]"puppy" -> [0, 1, 0, 0]"car" -> [0, 0, 1, 0]"seven" -> [0, 0, 0, 1]Measure relatedness with a dot product — multiply position by position, add up the results:
dog . puppy = (1x0) + (0x1) + (0x0) + (0x0) = 0dog . car = (1x0) + (0x0) + (0x1) + (0x0) = 0dog . seven = (1x0) + (0x0) + (0x0) + (0x1) = 0All three answers are 0. "dog" is exactly as related to "puppy" as to "seven". This is structural, not a tuning problem: two distinct one-hot vectors never share a non-zero position, so their dot product is always zero. Every pair of distinct words is perpendicular, and the geometry has no room to express "these two mean similar things".
It also scales appallingly: a 100,000-word vocabulary needs 100,000-dimensional vectors of which 99,999 entries are zero — 400 KB per word to encode one integer's worth of information.
Word embeddings fixed half the problem
Word2Vec and GloVe replaced that sparse slot with a learned dense vector — perhaps 300 non-zero numbers — trained so that words appearing in similar contexts get similar vectors. Now "dog" and "puppy" have a high dot product and "dog" and "seven" do not, and famously king - man + woman lands near queen: the vectors encode relationships, not just identity.
But a word vector represents a word. To represent a sentence, the obvious move is to average the word vectors — and that move throws away everything that word order does:
"The movie was not good, it was great""The movie was not great, it was good"Identical word multiset -> identical average -> identical vector.Opposite readings.Averaging also flattens negation and cannot disambiguate: "bank" gets one fixed vector whether the sentence is about a river or a mortgage.
What a sentence embedding is
A sentence embedding is a single dense vector — typically 384 to 3072 numbers — representing the meaning of an entire sentence or short document, produced by a model that reads the whole sequence at once. Two texts that mean similar things land close together even with no shared vocabulary.
An embedding does not store words. It stores a position. The whole design bet is that a well-chosen position in a few hundred dimensions can carry more useful information about a sentence than the sentence's own words can.
Inside a sentence transformer: four steps and some arithmetic
Sentence-BERT (SBERT), and the sentence-transformers library built around it, takes a BERT-style transformer — which naturally outputs one vector per token — and adds two steps that collapse those into one vector per text.
"The cat sat" [1] Tokenizer -> [CLS] the cat sat [SEP] [2] Transformer -> one 384-number vector PER TOKEN (context-aware) [3] Pooling -> combine token vectors into ONE vector [4] L2 normalise -> rescale that vector to length exactly 1 => a single 384-number vector, unit lengthStep 1, tokenisation splits text into subword pieces and adds two markers, [CLS] at the front and [SEP] at the end. Step 2, the transformer, produces a vector per token that depends on the surrounding tokens — which is why "bank" in "river bank" and "bank" in "savings bank" come out different. The ambiguity that killed word embeddings is resolved here.
Step 3, pooling, is where the collapse happens, and it is worth doing by hand. Take these five token vectors, using 4 dimensions instead of 384 so the arithmetic fits on a page:
[CLS] = [ 0.02, 0.11, -0.05, 0.20]the = [ 0.10, 0.20, -0.10, 0.05]cat = [ 0.80, 0.10, 0.30, 0.40]sat = [ 0.20, 0.70, -0.20, 0.10][SEP] = [ 0.03, 0.09, 0.00, 0.15]Mean pooling averages each position across all five tokens:
dim 0: (0.02 + 0.10 + 0.80 + 0.20 + 0.03) / 5 = 1.15 / 5 = 0.230dim 1: (0.11 + 0.20 + 0.10 + 0.70 + 0.09) / 5 = 1.20 / 5 = 0.240dim 2: (-0.05 - 0.10 + 0.30 - 0.20 + 0.00) / 5 = -0.05 / 5 = -0.010dim 3: (0.20 + 0.05 + 0.40 + 0.10 + 0.15) / 5 = 0.90 / 5 = 0.180pooled = [0.230, 0.240, -0.010, 0.180]Step 4, L2 normalisation, divides that vector by its own length. The length is the square root of the sum of the squares:
||pooled|| = sqrt(0.230^2 + 0.240^2 + (-0.010)^2 + 0.180^2) = sqrt(0.0529 + 0.0576 + 0.0001 + 0.0324) = sqrt(0.1430) = 0.37815normalised = [0.230/0.37815, 0.240/0.37815, -0.010/0.37815, 0.180/0.37815] = [0.6082, 0.6347, -0.0264, 0.4760]check: 0.6082^2 + 0.6347^2 + 0.0264^2 + 0.4760^2 = 0.3699 + 0.4028 + 0.0007 + 0.2266 = 1.0000 ✓Every embedding now sits on a sphere of radius 1. Only direction carries information; magnitude is gone by construction. That is deliberate: comparing two embeddings reduces to one dot product, and a long document cannot outrank a short one just for being long.
Pooling strategies, compared on the same numbers
Mean pooling is the default; running each strategy on the five token vectors above shows why the choice matters.
| Strategy | Result on the example above | What it does | When it is right |
|---|---|---|---|
mean | [0.230, 0.240, -0.010, 0.180] | Averages every token, so every word contributes | Default for similarity and retrieval; almost always the best starting point |
max | [0.800, 0.700, 0.300, 0.400] | Takes the largest value per dimension — here entirely dominated by "cat" and "sat" | Rarely; useful when one salient keyword should dominate the representation |
cls | [0.020, 0.110, -0.050, 0.200] | Uses only the [CLS] token vector, ignoring all real words | Only when the model was fine-tuned to make [CLS] a summary; on an untuned model it is near-arbitrary |
Notice how far apart those three results are. Max pooling produced a value in dimension 0 that is 3.5x larger than mean pooling's, because a single token carried it; [CLS] pooling produced something almost unrelated to either.
The common misconception is that "
[CLS]is the sentence embedding". That is true for BERT models fine-tuned on classification, where[CLS]was explicitly trained to summarise. For similarity embeddings, mean pooling wins because it draws signal from every word rather than betting everything on one special token.
You can override the default when you need to:
1from sentence_transformers import SentenceTransformer2from sentence_transformers.sentence_transformer.modules import Transformer, Pooling, Normalize34backbone = Transformer('bert-base-uncased')5pooling = Pooling(backbone.get_embedding_dimension(),6 pooling_mode='mean') # 'mean' | 'max' | 'cls'7model = SentenceTransformer(modules=[backbone, pooling, Normalize()])8print(model.encode(["This is a sentence."]).shape) # (1, 768)This is the sentence-transformers 6 layout. Older releases import the same classes from sentence_transformers.models and call get_word_embedding_dimension(); version 6 still accepts those, with a deprecation warning.
Your first embeddings
pip install sentence-transformers torch numpy scikit-learn1from sentence_transformers import SentenceTransformer23model = SentenceTransformer('all-MiniLM-L6-v2') # downloads on first use (~90 MB)45sentences = [6 "The cat sits on the mat.",7 "A feline rests on the rug.",8 "The weather is sunny today.",9]1011embeddings = model.encode(sentences, normalize_embeddings=True)1213print(embeddings.shape) # (3, 384)14print(embeddings[0][:8]) # first 8 of 384 numbers15print((embeddings[0] ** 2).sum()) # 1.0000001 -- unit length, as designed(3, 384) means three sentences, each turned into 384 floating-point numbers. Sentences 0 and 1 share only the word "the" yet describe the same scene, so a working model puts them close together. The numbers are individually meaningless — no dimension is "the cat dimension". Only relative positions mean anything.
Choosing a model, and the gotchas that come with each
| Model | Dims | Runs | Notes and quirks |
|---|---|---|---|
all-MiniLM-L6-v2 | 384 | Local, CPU-friendly | The sensible default. Six layers, ~22M parameters, max input 256 tokens. |
all-mpnet-base-v2 | 768 | Local | Better quality; roughly 4-5x slower and 2x the storage. |
multi-qa-mpnet-base-dot-v1 | 768 | Local | Tuned for asymmetric use: short question against long passage. |
BAAI/bge-large-en-v1.5 | 1024 | Local, GPU preferred | Strong retrieval. An instruction prefix on short queries helps a little (optional for v1.5); never on documents. |
intfloat/e5-large-v2 | 1024 | Local, GPU preferred | Requires "query: " / "passage: " prefixes — degrades silently without them. |
nomic-embed-text-v1.5 | up to 768 | Local | 8,192-token context, so it handles long documents. Trained for truncation (down to 64). Requires "search_query: " / "search_document: " prefixes, and loads custom code (trust_remote_code=True). |
text-embedding-3-small | up to 1536 | Hosted API | OpenAI's current small model (it replaced text-embedding-ada-002). No GPU needed, pay per token, 8,192-token input; supports a dimensions parameter. |
text-embedding-3-large | up to 3072 | Hosted API | Highest quality in that family; 8x MiniLM's storage. |
paraphrase-multilingual-mpnet-base-v2 | 768 | Local | 50+ languages share one space, so English and German text compare directly. |
Six things drive the choice.
- Latency and storage budget. 384 dimensions at 4 bytes each is 1,536 bytes per document; 3072 dimensions is 12,288 bytes — eight times the disk and eight times the arithmetic per comparison.
- Symmetric or asymmetric. Symmetric compares like with like (duplicate tickets). Asymmetric compares a short query against a long passage. Models tuned for one are measurably worse at the other.
- Input length. Every model has a hard
max_seq_length; MiniLM stops at 256 tokens and discards the rest silently. - Language. A monolingual English model given German text produces a vector — a confident, meaningless one.
- Prefix requirements. The most common silent failure in this whole area.
- Hardware. A 1024-dimension model on a CPU-only box can be an order of magnitude slower per document.
1model = SentenceTransformer('intfloat/e5-large-v2')23# RIGHT -- prefixes as the model was trained to expect4q = model.encode("query: What is the capital of France?")5p = model.encode("passage: Paris is the capital and largest city of France.")67# WRONG -- no error, no warning, just quietly worse rankings forever8q_bad = model.encode("What is the capital of France?")Nothing crashes in the wrong version. Search returns worse results, with no signal telling you why.
Recent versions of sentence-transformers also have model.encode_query() and model.encode_document(), which add the right prefix automatically, but only when the model's saved config defines one. E5, BGE and Nomic ship no prefixes in their configs, so for them those methods add nothing. Write the prefix yourself, or pass it explicitly with model.encode(text, prompt="query: ").
To shortlist candidates objectively, use MTEB (Massive Text Embedding Benchmark), which scores models across dozens of retrieval, clustering, and classification tasks. But the ranking that matters is the one measured on your documents and your queries.
Matryoshka embeddings: shrinking a vector without retraining
Work out the storage bill before picking a dimensionality.
| Dimensions | Bytes per vector (float32) | 1 million documents | 100 million documents |
|---|---|---|---|
| 256 | 1,024 | 1.02 GB | 102 GB |
| 384 | 1,536 | 1.54 GB | 154 GB |
| 768 | 3,072 | 3.07 GB | 307 GB |
| 1024 | 4,096 | 4.10 GB | 410 GB |
| 3072 | 12,288 | 12.29 GB | 1,229 GB |
At 100 million documents, the gap between 1024 and 256 dimensions is 308 GB of RAM — the difference between one machine and a cluster. And since every comparison costs one multiply-add per dimension, 256 dimensions is also 4x cheaper to search.
The obvious fix, using a smaller model, usually costs accuracy across the board. Matryoshka Representation Learning (MRL), named after Russian nesting dolls, offers something better: train one model so that prefixes of its output are themselves valid embeddings.
[ x1 ... x256 | x257 ... x512 | x513 ... x1024 ] \___________/ usable 256-dim embedding on its own \__________________________/ usable 512-dim embedding \_____________________________________________/ full 1024-dimMRL achieves this by changing the training objective: the loss is computed on the full vector and on several truncated prefixes at once, forcing the model to pack the most important information into the earliest dimensions. Ordinary training applies no such pressure, so information spreads arbitrarily across all dimensions.
1import numpy as np2from sentence_transformers import SentenceTransformer34model = SentenceTransformer('mixedbread-ai/mxbai-embed-large-v1') # MRL-trained, 1024-d56full = model.encode("Matryoshka embeddings can be shrunk without retraining.",7 normalize_embeddings=True)8print(full.shape) # (1024,)910truncated = full[:256] # keep the first 256 numbers1112# Slicing changes the vector's length, so unit-norm is broken -- restore it13print(np.linalg.norm(truncated)) # about 0.48, no longer 1.014truncated = truncated / np.linalg.norm(truncated)15print(np.linalg.norm(truncated)) # 1.0SentenceTransformer(name, truncate_dim=256) does the slicing for you on every encode() call; add normalize_embeddings=True and it re-normalises as well.
That re-normalisation is not optional. Truncating a unit vector leaves something shorter than 1 — how much shorter depends on how much energy lived in the discarded tail. Skip it, and your dot-product scores are scaled by a different amount for every document. Ranking becomes nonsense.
Truncation only works on models trained with MRL. Chopping
all-MiniLM-L6-v2from 384 to 128 dimensions is not a compression — it is deleting two-thirds of the information at random, and retrieval quality collapses accordingly.
Encoding a large collection without melting your laptop
Embedding three sentences is instant. Embedding 500,000 support tickets needs thought.
1from sentence_transformers import SentenceTransformer23model = SentenceTransformer('all-MiniLM-L6-v2')45embeddings = model.encode(6 documents,7 batch_size=64, # texts pushed through the model at once8 normalize_embeddings=True, # do the L2 step here, once, for everything9 show_progress_bar=True,10 device='cuda', # omit to auto-select; 'cpu' is fine for MiniLM11)Batch size is the main lever, and its effect has a consistent shape:
| Batch size | Relative throughput | Why |
|---|---|---|
| 1 | 1.0x (baseline) | Per-call overhead dominates the actual maths |
| 8 | ~4x | Overhead amortised; matrix multiplications get usefully wide |
| 32 | ~6x | Hardware near saturation |
| 64 | ~6.5x | Diminishing returns begin |
| 512+ | may run out of GPU memory | Activations scale with batch size x sequence length |
The lesson is not "use 512": going from 1 to 32 buys almost everything, and going far beyond buys little speed and, on a small GPU, a crash. Sorting documents by length first helps too, because each batch is padded to its longest member — one 250-token document among 63 ten-token ones wastes most of the batch on padding.
The silent truncation trap
This one costs people weeks.
1print(model.max_seq_length) # 256 for all-MiniLM-L6-v223long_article = open('support_article.txt').read() # ~900 words ~= 1,200 tokens4# 256 / 1200 = 21% kept. The other 79% is dropped -- no exception, no warning.5emb = model.encode(long_article)Your embedding represents the first fifth of the article. If the answer lives in the conclusion, it is not in the vector, and no amount of search tuning will find it. The fix is a long-context model (nomic-embed-text-v1.5 handles 8,192 tokens) or splitting the document into overlapping chunks of a few hundred tokens and embedding each separately.
Judging whether an embedding model is actually good
Four properties matter, and they pull against each other:
- Semantic preservation — texts that mean the same thing land close together.
- Discriminativeness — texts that mean different things land clearly apart. A model mapping everything to nearly one point scores perfectly on the first property and is useless.
- Efficiency — encoding throughput and storage cost you can afford.
- Robustness — quality holds on your domain's vocabulary, not just Wikipedia prose.
The standard measurement is Spearman rank correlation between the model's scores and human judgements on the same pairs: when people rank 200 sentence pairs from most to least similar, does the model produce the same ordering? 1.0 is perfect agreement, 0.0 is none.
1from sentence_transformers.sentence_transformer.evaluation import EmbeddingSimilarityEvaluator23# Pairs from YOUR domain, scored 0-1 by people who know the domain4s1 = ["Card declined at checkout", "How do I reset my password"] # ...200 of these5s2 = ["Payment failed during purchase", "Where can I change my email"]6gold = [0.90, 0.35]78results = EmbeddingSimilarityEvaluator(s1, s2, gold)(model) # a dict of metrics9print(results["spearman_cosine"]) # higher is betterFifty to two hundred hand-scored pairs from your own data tell you more than any leaderboard. Benchmarks show which models are generally competent; only your pairs show which one understands "chargeback" the way your users mean it.
What people build with these vectors
Ranking documents against a query
1import numpy as np2from sentence_transformers import SentenceTransformer34model = SentenceTransformer('all-MiniLM-L6-v2')56documents = ["Python is a programming language",7 "Java is also a programming language",8 "The weather is nice today",9 "I like to code in Python"]1011doc_emb = model.encode(documents, normalize_embeddings=True)12q_emb = model.encode("Programming languages like Python", normalize_embeddings=True)1314# Both sides are unit length, so a plain dot product IS cosine similarity15scores = doc_emb @ q_emb16for rank, i in enumerate(np.argsort(scores)[::-1][:3], start=1):17 print(f"{rank}. {scores[i]:.3f} {documents[i]}")One matrix multiply scores every document at once; for 100,000 documents NumPy handles it in milliseconds.
Grouping documents with no labels
from sklearn.cluster import KMeanslabels = KMeans(n_clusters=5, n_init=10, random_state=42).fit_predict(doc_emb)Because unit-length vectors make Euclidean distance a monotonic function of cosine similarity, k-means on normalised embeddings groups by meaning — which is how a support team discovers that 30% of last month's tickets were one billing complaint phrased forty ways.
Finding near-duplicates
1from sklearn.metrics.pairwise import cosine_similarity23docs = ["The quick brown fox jumps over the lazy dog",4 "A fast brown fox leaps over a sleepy dog",5 "Python is great for data science"]67sim = cosine_similarity(model.encode(docs))8for i in range(len(docs)):9 for j in range(i + 1, len(docs)):10 if sim[i][j] > 0.70:11 print(f"{sim[i][j]:.2f} {docs[i]!r} / {docs[j]!r}")Documents 0 and 1 share four words out of nine yet describe an identical event. Keyword overlap scores them around 0.4; an embedding model scores them near 0.85.
Fine-tuning when the general model is not enough
Pre-trained models learned general English. They do not know that "SEV-2" and "priority incident" are the same thing at your company, or that in cardiology "MI" means myocardial infarction, not Michigan. Fine-tuning teaches an existing model your vocabulary without starting over.
pip install sentence-transformers datasets accelerate1from datasets import Dataset2from sentence_transformers import SentenceTransformer, SentenceTransformerTrainer3from sentence_transformers.sentence_transformer.losses import CosineSimilarityLoss45model = SentenceTransformer('all-MiniLM-L6-v2')67# Option A: explicit similarity labels, 0.0 to 1.08labelled = Dataset.from_dict({9 "sentence1": ["myocardial infarction", "myocardial infarction"],10 "sentence2": ["heart attack", "bone fracture"],11 "score": [0.95, 0.05],12})13trainer = SentenceTransformerTrainer(model=model, train_dataset=labelled,14 loss=CosineSimilarityLoss(model))15trainer.train()16model.save_pretrained('cardiology-embeddings')The trainer reads the columns in order (two texts, then the label) and uses Hugging Face's training loop underneath. Older tutorials call model.fit(...) with InputExample objects; that still exists as a legacy wrapper, but it now needs datasets installed too and new code should use the trainer.
Hand-labelling scores is slow. A cheaper option usually works better:
1from sentence_transformers import SentenceTransformerTrainingArguments2from sentence_transformers.sentence_transformer.losses import MultipleNegativesRankingLoss34# Option B: positive PAIRS only -- no scores needed5pairs = Dataset.from_dict({6 "anchor": ["myocardial infarction", "how do I reset my password"],7 "positive": ["heart attack", "password reset instructions"],8})9args = SentenceTransformerTrainingArguments(10 output_dir="cardiology-embeddings", num_train_epochs=1,11 per_device_train_batch_size=32,12)13trainer = SentenceTransformerTrainer(model=model, args=args, train_dataset=pairs,14 loss=MultipleNegativesRankingLoss(model))15trainer.train()MultipleNegativesRankingLoss treats the other items in the same batch as negatives. With a batch size of 32, each positive pair is trained against 31 automatically generated negatives — you supplied 32 labels and the loss manufactured 992 comparisons from them. That is why larger batches help with this loss specifically, and why a few thousand naturally-occurring pairs (a question and its accepted answer, a ticket and its resolution article) beat a few hundred painstakingly scored ones.
What this means when you build something
The failures below all show up in real projects, none raise an exception, and every one degrades results quietly. That is what makes them expensive.
| Symptom | Likely cause | Fix |
|---|---|---|
| Search results are plausible but consistently mediocre, and always have been | Model needs a prefix (E5, BGE) and you never added one | Read the model card; add "query: " / "passage: " exactly as specified |
| Long documents never match queries about their later content | Silent truncation at max_seq_length | Print model.max_seq_length; chunk documents or switch to a long-context model |
| Similarity scores look random after a storage optimisation | Truncated a non-MRL model, or forgot to re-normalise after slicing | Only truncate MRL-trained models, and always divide by the new norm |
| Quality collapsed after "upgrading" the model | Queries encoded with model B against an index built with model A | Re-embed the entire corpus on any model change; there is no partial migration |
| Scores hover around 0.9 for everything, relevant or not | Poor discriminativeness on your domain | Judge the score distribution, not absolute values; consider fine-tuning |
| Encoding 100k documents takes hours | batch_size=1, or unsorted inputs padding every batch | Set batch_size=32+, sort by length, use a GPU if available |
Two operational rules follow, and both are worth writing down before you start.
Pin the model name in configuration, never in scattered code. Every embedding belongs to exactly one model's coordinate system. Two models can both output 768 numbers and be completely incomparable, because there is no reason dimension 412 means the same thing in both. The day someone changes the model name in one file and not another, your search starts comparing coordinates from two different universes — and it returns results, ranked confidently, and wrong.
Store the model name alongside the vectors. A vector on disk with no record of what produced it is unusable data. Write the model identifier, the dimensionality, and whether the vectors are normalised into the same store as the embeddings. Six months later, when someone adds a million documents to your index, that one line of metadata is the difference between a five-minute job and re-embedding everything.