Embeddings and Semantic Search

Cosine Similarity & Vector Arithmetic


A support team builds a duplicate-ticket detector. Every incoming ticket is embedded, compared against the last 30 days of tickets, and anything "close enough" gets merged. It works in testing. In production it merges almost nothing, and the tickets it does merge are wrong.

The engineer digs in. Ticket A reads "App crashes on login" — five words. Ticket B is the same complaint from a different customer: three paragraphs, a stack trace, and an apology for the length. The system scored them 12.7 apart and called them unrelated, while merging two short tickets about entirely different features because both were short.

The bug was one line. The comparison used Euclidean distance on raw, unnormalised vectors. Euclidean distance measures how far apart two points are, and a long document produces a vector further from the origin than a short one. The detector was not measuring topic. It was measuring length.

A vector on its own says nothing. It becomes useful only when you compare it to another vector, and the choice of comparison decides what "similar" means. Get it wrong and you build a length detector, and you will not notice until production.

Direction or position — they answer different questionsCosine similarity• Angle only; length is divided out• A long guide and a short note can match• Range minus 1 to 1,comparable across pairs• Free afternormalisation — it is a dot productEuclidean distance• Straight-line gap, so length counts• Long documents driftaway from short queries• Range 0 upward, no natural cut-off• On unit vectors it ranks identically
Once both vectors are normalised the two metrics give the same ordering, so the real decision is whether you normalised — and a 0.7 threshold tuned on one model means nothing on the next.

What we actually want to measure: direction, not position

Think of an embedding as an arrow from the origin. It has a direction (which way it points) and a magnitude (how long it is). For text embeddings, direction carries the meaning; magnitude mostly carries artefacts — document length, token count, how strongly the model activated.

Here is the clearest possible demonstration. Take a toy three-dimensional space whose axes count occurrences of python, learn and code:

  • Doc A — a short tutorial: [2, 1, 1]
  • Doc B — the same tutorial expanded five-fold: [10, 5, 5]
  • Doc C — a different article, mostly about learning: [1, 2, 1]

Now compute Euclidean distance, which is the straight-line gap between the two points:

d(a,b)=∑i(ai−bi)2d(\mathbf{a},\mathbf{b}) = \sqrt{\sum_i (a_i - b_i)^2}

For A and B: differences are (2−10,1−5,1−5)=(−8,−4,−4)(2-10, 1-5, 1-5) = (-8, -4, -4), so d=64+16+16=96=9.80d = \sqrt{64 + 16 + 16} = \sqrt{96} = 9.80.

For A and C: differences are (1,−1,0)(1, -1, 0), so d=1+1+0=2=1.41d = \sqrt{1 + 1 + 0} = \sqrt{2} = 1.41.

Euclidean distance says Doc C is seven times closer to Doc A than Doc B is — even though Doc B is literally the same document, just longer. That is the ticket bug, reproduced in three dimensions.

Cosine similarity, computed by hand

Cosine similarity throws magnitude away and keeps only the angle:

cos⁡(θ)=a⋅b∥a∥ ∥b∥=∑iaibi∑iai2  ∑ibi2\cos(\theta) = \frac{\mathbf{a} \cdot \mathbf{b}}{\lVert \mathbf{a} \rVert \, \lVert \mathbf{b} \rVert} = \frac{\sum_i a_i b_i}{\sqrt{\sum_i a_i^2}\;\sqrt{\sum_i b_i^2}}

The numerator is the dot product — multiply the vectors element by element and add up the results. The denominator is the product of the two L2 norms (each vector's length). Dividing by the lengths is what cancels magnitude out.

Run it on the same three documents, slowly.

A versus B

  • Dot product: (2)(10)+(1)(5)+(1)(5)=20+5+5=30(2)(10) + (1)(5) + (1)(5) = 20 + 5 + 5 = 30
  • ∥A∥=4+1+1=6=2.4495\lVert A \rVert = \sqrt{4 + 1 + 1} = \sqrt{6} = 2.4495
  • ∥B∥=100+25+25=150=12.2474\lVert B \rVert = \sqrt{100 + 25 + 25} = \sqrt{150} = 12.2474
  • cos⁡=30/(2.4495×12.2474)=30/30.0000=1.000\cos = 30 / (2.4495 \times 12.2474) = 30 / 30.0000 = \mathbf{1.000}

Exactly 1.0, and that is the point: [10, 5, 5] is precisely 5 times [2, 1, 1], so the two arrows point in identical directions. Cosine recognises them as the same document; Euclidean distance did not.

A versus C

  • Dot product: (2)(1)+(1)(2)+(1)(1)=2+2+1=5(2)(1) + (1)(2) + (1)(1) = 2 + 2 + 1 = 5
  • ∥A∥=6=2.4495\lVert A \rVert = \sqrt{6} = 2.4495, ∥C∥=1+4+1=6=2.4495\lVert C \rVert = \sqrt{1 + 4 + 1} = \sqrt{6} = 2.4495
  • cos⁡=5/(2.4495×2.4495)=5/6=0.8333\cos = 5 / (2.4495 \times 2.4495) = 5 / 6 = \mathbf{0.8333}

So the cosine ranking is B (1.000) above C (0.833) — the opposite of the Euclidean ranking, and the one a human would give.

Euclidean distance asks "where are these two points?". Cosine similarity asks "which way are these two arrows pointing?". For text, the answer you want is almost always the second one.

The full range, with worked cases

VectorsDot productNormsCosineAngleMeaning
[1,0] and [1,0]11 × 11.0000°Identical direction
[1,0] and [0.866, 0.5]0.8661 × 10.86630°Closely related
[1,0] and [0.707, 0.707]0.7071 × 10.70745°Moderately related
[1,0] and [0,1]01 × 10.00090°Unrelated (orthogonal)
[1,1] and [-1,-1]−21.414 × 1.414−1.000180°Exactly opposed

The formal range is −1 to +1, but with a well-trained sentence encoder you will almost never see a negative score. Trained embeddings occupy a narrow cone of the space, so real scores cluster between 0.0 and 1.0, with genuinely unrelated pairs landing around 0.05 to 0.20 rather than at zero. If you see −0.6 between two English sentences, suspect a bug before concluding they are semantic opposites.

Normalisation, and why it makes the dot product free

To normalise a vector means to divide it by its own length so the result has length exactly 1. Normalise Doc A:

A^=[2,1,1]2.4495=[0.8165,  0.4082,  0.4082]\hat{A} = \frac{[2, 1, 1]}{2.4495} = [0.8165,\; 0.4082,\; 0.4082]

Check it: 0.81652+0.40822+0.40822=0.6667+0.1666+0.1666=0.9999≈10.8165^2 + 0.4082^2 + 0.4082^2 = 0.6667 + 0.1666 + 0.1666 = 0.9999 \approx 1. Good.

Normalise Doc C the same way: C^=[0.4082,  0.8165,  0.4082]\hat{C} = [0.4082,\; 0.8165,\; 0.4082].

Now take the plain dot product of the two normalised vectors, with no division at all:

A^⋅C^=(0.8165)(0.4082)+(0.4082)(0.8165)+(0.4082)(0.4082)=0.3333+0.3333+0.1666=0.8332\hat{A} \cdot \hat{C} = (0.8165)(0.4082) + (0.4082)(0.8165) + (0.4082)(0.4082) = 0.3333 + 0.3333 + 0.1666 = 0.8332

That is 0.8333 again — the cosine we computed the long way. Not a coincidence: when both norms are 1, the cosine denominator is 1×1=11 \times 1 = 1, so cosine is the dot product. That single fact drives most production design decisions. Normalise once at write time and every later comparison is one multiply-add per dimension, with no square roots.

On unit-length vectors, dot product and cosine similarity are the same number. Normalise when you store, and you never pay for a norm again.

And Euclidean distance collapses too

Expand the squared distance between two unit vectors:

d2=∥a∥2+∥b∥2−2(a⋅b)=1+1−2cos⁡θ=2(1−cos⁡θ)d^2 = \lVert a \rVert^2 + \lVert b \rVert^2 - 2(\mathbf{a} \cdot \mathbf{b}) = 1 + 1 - 2\cos\theta = 2(1 - \cos\theta)

Test it on A^\hat{A} and C^\hat{C}, where cos⁡=0.8333\cos = 0.8333: predicted d=2×0.1667=0.3333=0.5774d = \sqrt{2 \times 0.1667} = \sqrt{0.3333} = 0.5774. Compute it directly instead: A^−C^=[0.4083,−0.4083,0]\hat{A} - \hat{C} = [0.4083, -0.4083, 0], giving d=0.1667+0.1667=0.3334=0.5774d = \sqrt{0.1667 + 0.1667} = \sqrt{0.3334} = 0.5774. Identical.

Because dd shrinks monotonically as cos⁡\cos grows, on normalised vectors Euclidean distance and cosine similarity produce exactly the same ranking. So when a library offers only L2 distance, normalise first and the metric choice stops mattering. It also re-frames the opening bug: Euclidean was not the wrong metric, it was applied to vectors nobody had normalised.

Three implementations, and when each earns its place

Python
import numpy as npdef cosine_by_hand(a, b):    """Educational version - shows every step of the formula."""    na, nb = np.linalg.norm(a), np.linalg.norm(b)    if na == 0 or nb == 0:          # guard: see the NaN trap below        return 0.0    return float(np.dot(a, b) / (na * nb))A = np.array([2.0, 1.0, 1.0])C = np.array([1.0, 2.0, 1.0])print(f"{cosine_by_hand(A, C):.4f}")   # 0.8333

Beyond one pair, use the library version — a whole matrix in one BLAS call:

Python
from sklearn.metrics.pairwise import cosine_similarityimport numpy as npvectors = np.array([    [2.0, 1.0, 1.0],    # Doc A    [10.0, 5.0, 5.0],   # Doc B - same doc, 5x longer    [1.0, 2.0, 1.0],    # Doc C])print(np.round(cosine_similarity(vectors), 4))# [[1.     1.     0.8333]#  [1.     1.     0.8333]#  [0.8333 0.8333 1.    ]]

cosine_similarity normalises internally, which is why Doc A and Doc B come out at 1.0 even though neither was normalised beforehand. That built-in safety is exactly what makes people forget normalisation matters — and then they move to a vector index that does not normalise for them.

On GPU, PyTorch keeps everything on the device and avoids a round trip to host memory:

Python
import torchimport torch.nn.functional as Fdevice = "cuda" if torch.cuda.is_available() else "cpu"q = F.normalize(torch.randn(8, 384, device=device), p=2, dim=1)d = F.normalize(torch.randn(50_000, 384, device=device), p=2, dim=1)scores = q @ d.T          # (8, 50000) cosine scores in one matmulprint(scores.topk(k=5, dim=1).values[0])
ApproachBest forCost per query over 10k docs (384-d)Watch out for
NumPy, hand-written loopLearning the formula~0.5 s (10,000 Python iterations)Unusably slow at any real size
sklearn.cosine_similarityEveryday CPU work~1–3 ms (single matmul)Materialises the full matrix in RAM
PyTorch on GPUBatched queries, large corpora< 1 ms once data is residentTransfer cost dominates for small jobs

Why the loop is 500 times slower

The arithmetic is identical in both cases: one query against 10,000 documents at 384 dimensions is 10,000×384=3.8410{,}000 \times 384 = 3.84 million multiply-adds, roughly 7.7 MFLOP, which a laptop's BLAS library runs in well under a millisecond. The loop performs the same 7.7 MFLOP but wraps every 384-element chunk in a Python call, an array allocation and a library dispatch — call it 50 microseconds each. Ten thousand of those is half a second of pure overhead around a millisecond of real work.

Scale up and even the fast path strains. One million documents at 768 dimensions is 768 million multiply-adds per query — about 1.5 GFLOP, or 30 to 50 ms of pure compute — and the matrix occupies 1,000,000×768×41{,}000{,}000 \times 768 \times 4 bytes = 3.07 GB in float32. At that point exhaustive comparison stops being viable. But be honest about the crossover: under roughly 100,000 documents, a plain matrix multiply is fast enough, simpler, and exact.

Vector arithmetic: adding and subtracting meaning

Something strange falls out of representing meaning as coordinates: you can do algebra on it. The classic demonstration is king−man+woman≈queen\text{king} - \text{man} + \text{woman} \approx \text{queen}. To see why, build a toy space with four interpretable axes — royalty, maleness, femaleness, humanness:

Wordroyaltymalefemalehuman
king0.90.80.00.7
man0.10.90.00.9
woman0.10.00.90.9
queen0.90.00.850.7

Compute the analogy component by component:

king−man+woman=[0.9−0.1+0.1,  0.8−0.9+0.0,  0.0−0.0+0.9,  0.7−0.9+0.9]=[0.9,  −0.1,  0.9,  0.7]\text{king} - \text{man} + \text{woman} = [0.9 - 0.1 + 0.1,\; 0.8 - 0.9 + 0.0,\; 0.0 - 0.0 + 0.9,\; 0.7 - 0.9 + 0.9] = [0.9,\; -0.1,\; 0.9,\; 0.7]

Subtracting man stripped out the maleness (0.8 became −0.1) while leaving royalty almost untouched, because man barely has any. Adding woman installed femaleness. Score the result against queen:

  • Dot product: (0.9)(0.9)+(−0.1)(0.0)+(0.9)(0.85)+(0.7)(0.7)=0.81+0+0.765+0.49=2.065(0.9)(0.9) + (-0.1)(0.0) + (0.9)(0.85) + (0.7)(0.7) = 0.81 + 0 + 0.765 + 0.49 = 2.065
  • ∥result∥=0.81+0.01+0.81+0.49=2.12=1.4560\lVert \text{result} \rVert = \sqrt{0.81 + 0.01 + 0.81 + 0.49} = \sqrt{2.12} = 1.4560
  • ∥queen∥=0.81+0+0.7225+0.49=2.0225=1.4221\lVert \text{queen} \rVert = \sqrt{0.81 + 0 + 0.7225 + 0.49} = \sqrt{2.0225} = 1.4221
  • cos⁡=2.065/(1.4560×1.4221)=2.065/2.0707=0.9973\cos = 2.065 / (1.4560 \times 1.4221) = 2.065 / 2.0707 = \mathbf{0.9973}

Against king, the same arithmetic gives 1.22/(1.4560×1.3928)=1.22/2.0280=0.60161.22 / (1.4560 \times 1.3928) = 1.22 / 2.0280 = 0.6016. Against man: 0.63/(1.4560×1.2767)=0.63/1.8589=0.33890.63 / (1.4560 \times 1.2767) = 0.63 / 1.8589 = 0.3389. So queen wins decisively at 0.997.

Where this breaks, and what every real implementation does about it

My toy axes are cleanly disentangled — one axis is maleness. Real embedding dimensions are nothing like that; meaning is smeared across hundreds of coordinates, and no single dimension corresponds to a human concept. The consequence is severe: in a real space, king - man + woman usually lands closest to king itself, because subtracting and re-adding a near-parallel pair barely moves you. Every published analogy result quietly excludes the input words from the candidate list. Fail to exclude them and your analogy solver returns its own input.

Python
import numpy as npfrom sentence_transformers import SentenceTransformermodel = SentenceTransformer("all-MiniLM-L6-v2")inputs = ["king", "man", "woman"]candidates = ["queen", "princess", "monarch", "boy", "girl"] + inputsemb = model.encode(inputs + candidates, normalize_embeddings=True)k, m, w = emb[0], emb[1], emb[2]target = k - m + wtarget = target / np.linalg.norm(target)   # the sum is NOT unit lengthscores = emb[3:] @ targetfor name, s in sorted(zip(candidates, scores), key=lambda p: -p[1]):    if name not in inputs:          # exclude the inputs, or "king" wins        print(f"{s:.3f}  {name}")

Two details there matter more than the analogy. normalize_embeddings=True makes the final @ a genuine cosine, and the sum of unit vectors is not a unit vector, so the result is renormalised before comparison. Skip that and your scores are wrong in a way that still looks plausible.

Blending two concepts

Averaging is the well-behaved cousin of analogy arithmetic, and it is how you build a "profile" vector from several examples. Take two unit vectors: smart=[0.8,0.6]\text{smart} = [0.8, 0.6] and friendly=[0.6,0.8]\text{friendly} = [0.6, 0.8], whose cosine is (0.8)(0.6)+(0.6)(0.8)=0.96(0.8)(0.6) + (0.6)(0.8) = 0.96 — about 16.3° apart. Their blend is [0.7,0.7][0.7, 0.7], with length 0.98=0.98995\sqrt{0.98} = 0.98995, shorter than 1. Its cosine to smart is 0.98/0.98995=0.989950.98 / 0.98995 = 0.98995, and by symmetry the same to friendly: the blend sits on the bisector, 8.13° from each, exactly half of 16.3°.

Weighted interpolation, (1−t) a+t b(1-t)\,\mathbf{a} + t\,\mathbf{b}, traces a path between two meanings. With a=[1,0]\mathbf{a} = [1,0] and b=[0,1]\mathbf{b} = [0,1], a full 90° apart, the t=0.5t = 0.5 midpoint [0.5,0.5][0.5, 0.5] has length only 0.70710.7071 — the straight line cuts through the inside of the sphere instead of following its surface. Renormalise to [0.7071,0.7071][0.7071, 0.7071] and you are back on the sphere at the true 45° point. Always renormalise after adding, subtracting or averaging embeddings.

Turning scores into results: three thresholding strategies

Cosine gives you a ranked list of numbers. Deciding which to actually show is a separate decision, and getting it wrong is a common cause of a search feature that "feels broken".

Take two real queries against the same corpus:

  • Query A (well-covered topic) scores: 0.81, 0.74, 0.70, 0.23, 0.09
  • Query B (nothing relevant in the corpus) scores: 0.31, 0.29, 0.28, 0.26, 0.11
StrategyRuleQuery A returnsQuery B returnsUse when
Hard thresholdKeep score > 0.53 results (0.81, 0.74, 0.70)0 resultsAn empty result beats a wrong one — support search, legal lookup
Top-kKeep the best 55 results, including the 0.09 junk5 results, all irrelevantThe slot must always be filled — "related articles"
PercentileKeep score > 90th percentile1 result (cutoff 0.782)1 result (cutoff 0.302)Score ranges differ wildly per query; you want relative quality

Work the Query A cutoff so the number is not magic. Sorted ascending: 0.09, 0.23, 0.70, 0.74, 0.81. The 90th percentile with linear interpolation sits at position 0.9×(5−1)=3.60.9 \times (5-1) = 3.6, which is 60% of the way from 0.74 to 0.81: 0.74+0.6(0.07)=0.7820.74 + 0.6(0.07) = 0.782. Only 0.81 clears it. For Query B the same computation gives 0.29+0.6(0.02)=0.3020.29 + 0.6(0.02) = 0.302, and 0.31 clears it — the strategy's weakness. A percentile threshold can never return nothing. It hands back the best of a bad batch.

Choose the strategy from the cost of being wrong. If a bad result merely wastes a click, use top-k. If a bad result gets fed to a language model and turned into a confident false answer, use a hard threshold and accept the empty page.

Thresholds do not transfer between models

A 0.7 cutoff that works with one encoder can be nonsense with another, because models differ in how tightly they pack the space. With all-MiniLM-L6-v2, unrelated English pairs typically score 0.05–0.15, so 0.7 is a demanding bar. Older encoders trained without hard negatives put unrelated pairs at 0.35–0.50, where 0.7 lets weak matches through. Swap the model, keep the threshold, and precision changes silently. Fix it by labelling a few hundred pairs yourself and picking the cutoff where the two distributions separate.

The four bugs that actually happen

1. A zero vector produces NaN, and NaN sorts to the top

Dividing by a zero norm gives nan. That alone would be survivable; the damage comes from what happens next. NumPy's argsort treats nan as larger than everything, so it lands at the end of an ascending sort. Reverse that array to get descending order — the standard idiom — and your nan is now rank 1.

Python
import numpy as npscores = np.array([0.81, np.nan, 0.74, 0.09])order = np.argsort(scores)[::-1]print(order)            # [1 0 2 3]  -> the NaN is your top hitprint(scores[order])    # [nan 0.81 0.74 0.09]# Fix: neutralise non-finite scores before rankingscores = np.nan_to_num(scores, nan=-1.0, posinf=-1.0, neginf=-1.0)print(np.argsort(scores)[::-1])   # [0 2 3 1]

Zero vectors arrive from whitespace-only documents, from averaging an empty list of chunk vectors, and from hand-built sparse features. Filter empty text before encoding and guard the division.

2. Raw dot product on unnormalised vectors ranks by length

This bites when you move from scikit-learn (which normalises for you) to a vector index that computes inner product directly. Suppose the query is q=[1,0]q = [1, 0] and there are two documents:

  • d1=[0.9,0.436]d_1 = [0.9, 0.436], length 1.0, so cos⁡(q,d1)=0.9\cos(q, d_1) = 0.9 and inner product =0.9= 0.9
  • d2=[1.5,2.598]d_2 = [1.5, 2.598], length 3.0, so cos⁡(q,d2)=1.5/3.0=0.5\cos(q, d_2) = 1.5/3.0 = 0.5 and inner product =1.5= 1.5

Rank by cosine and d1d_1 wins (0.9 vs 0.5). Rank by raw inner product and d2d_2 wins (1.5 vs 0.9) — a worse match promoted purely because its vector is three times longer. Nothing throws, every result looks plausible, and the ranking is upside down. Normalise at write time and inner product is cosine, so the bug becomes impossible.

3. Comparing vectors from two different models

Two models can both output 768 dimensions and still be incomparable, because each learned its own axes from scratch. Dimension 42 means one thing in one model and something unrelated in the other. The comparison does not error — it returns a smallish number that reads as "unrelated" and is meaningless. This surfaces during upgrades: someone embeds new documents with a better model and leaves the old vectors in the index. Quality degrades slowly and nobody finds the cause.

4. Treating a high score as a factual match

Cosine similarity measures directional closeness in a learned space. It does not measure truth. "The API returns a 200 on success" and "The API returns a 500 on success" share nearly every token and the entire topic, and score above 0.9 with most encoders. Negation, numbers, dates and named entities are exactly where embeddings are weakest, because one flipped word barely moves a vector built from averaged token representations. If those distinctions matter, you need a keyword or structured check alongside.

Choosing a metric for real work

MetricFormulaRangeCost per pair (d dims)Choose it when
Cosine similaritya⋅b∥a∥∥b∥\frac{a \cdot b}{\lVert a\rVert \lVert b\rVert}−1 to 13d3d ops + 2 square rootsDefault for text; magnitude is noise.
Dot product∑aibi\sum a_i b_iunbounded2d2d opsVectors already normalised, or magnitude deliberately encodes importance.
Euclidean (L2)∑(ai−bi)2\sqrt{\sum (a_i-b_i)^2}0 to ∞3d3d ops + 1 square rootCoordinates are real measurements, or the library offers only L2.
Manhattan (L1)∑∣ai−bi∣\sum \lvert a_i-b_i \rvert0 to ∞2d2d ops, no multiplyVery high dimensions where L2 concentrates.

The decision is smaller than it looks: normalise, and cosine, dot product and Euclidean all give the same ranking, differing only in arithmetic cost. The real decision is whether to normalise, and for text the answer is yes.

What this means when you build something

Normalise once, at write time, and record that you did. Pass normalize_embeddings=True when you encode, store the unit vectors, and put the model name and dimensionality in the same row. Every downstream component can then use a plain dot product and be correct by construction. The alternative is normalising in four places, forgetting one, and shipping a length detector.

Before picking a threshold, look at the distribution. Take 200 pairs from your own data, label them relevant or not by hand, embed them, and plot the two score histograms. You learn three things in ten minutes: where relevant pairs sit, where irrelevant pairs sit, and whether they overlap so badly that no threshold separates them. That last case is the useful discovery — it says the problem is the model or the chunking, not the cutoff.

Do not loop. The moment you write for doc in docs: similarity(query, doc), you have thrown away a hundredfold speedup for nothing. Stack the vectors into one array, normalise, and multiply. When that matrix stops fitting in memory, around a few hundred thousand vectors, that is the signal to move to an approximate nearest-neighbour index rather than to buy a bigger machine.

Treat every score as evidence, not as an answer. A 0.92 means two texts point in similar directions in a space some model learned from some corpus. It does not mean they say the same thing, nor that a 0.91 result is meaningfully worse. Build pipelines that survive the top result occasionally being wrong, because it will be.