RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is Cosine Similarity and why is it used in vector searches?


Query (2, 1, 0) against four chunks104.5830.97653.1620.70723.1620.283159.4870.707dotlengthcosineA (4, 2, 1)B (1, 3, 0)C (1, 0, 3)F = 3 times B
Dot product ranks F first only because it is three times longer; cosine gives F and B the same score.

What you need to know

Text
cosine(a, b) = (a · b) / (|a| * |b|)a · b = a1*b1 + a2*b2 + ... + an*bn        (dot product)|a|   = sqrt(a1² + a2² + ... + an²)         (length, or L2 norm)

A worked example by hand

Real embeddings have hundreds of dimensions, but the arithmetic is the same in three. Query q = (2, 1, 0) and four chunks:

ChunkVectorDot with qLengthCosine
A(4, 2, 1)2·4 + 1·2 + 0·1 = 10√21 = 4.58310 / (2.236 × 4.583) = 0.976
B(1, 3, 0)2 + 3 = 5√10 = 3.1625 / (2.236 × 3.162) = 0.707
C(1, 0, 3)2√10 = 3.1622 / 7.071 = 0.283
F(3, 9, 0) = 3 × B6 + 9 = 15√90 = 9.48715 / (2.236 × 9.487) = 0.707

(|q| = √5 = 2.236.) Two things to see:

  • Cosine ranks A first. A points almost the same way as q.
  • Raw dot product ranks F first (15 against 10), only because F is three times longer than B. Its direction is identical to B's, and cosine correctly gives them the same score. That is the length bias cosine removes.

Normalise, and the metrics agree

If every vector is scaled to length 1, the denominator becomes 1, so cosine equals the dot product. And squared Euclidean distance becomes 2 - 2 × cosine:

Text
A: 2 - 2 × 0.976 = 0.048     (closest)B: 2 - 2 × 0.707 = 0.586C: 2 - 2 × 0.283 = 1.434     (farthest)

Same order as cosine. That is why many systems normalise at write time and then use the fastest operation, the inner product (FAISS IndexFlatIP, pgvector <#> or <=>, Chroma with space: cosine).

Reading real scores

With real models, scores cluster in a model-specific band. With bge-small-en-v1.5, an unrelated sentence scored about 0.38 and a strong match about 0.77. Do not hard-code thresholds like "0.8 means relevant"; calibrate them per model on labelled data.

When not to use cosine

Use whatever the model card says. A few models are trained with dot product on unnormalised vectors, where length carries a signal such as confidence. Euclidean distance on unnormalised text embeddings is rarely right.

A real-life example

A bank's FAQ bot stores chunks in a store whose default metric is L2, without normalising vectors. Long product-terms chunks have larger vectors than short FAQ answers. For "what is the minimum balance for a savings account?", the short FAQ answer is the best match in direction, but the ranking is dominated by vector length, and a long, general "Savings account terms" chunk ranks above it.

The fix is two lines: normalise embeddings at write and query time, and set the collection's space to cosine. The short, precise FAQ answer moves to the top. Nothing else changed, which is why interviewers like this question: a metric mismatch produces no error, only worse answers.

Follow-up questions to expect

  • "Can cosine similarity be negative for text?" — Yes in principle, but most text embedding models produce mostly positive similarities; unrelated text tends to sit at a low positive value.
  • "Why is cosine distance written as 1 - cosine?" — So smaller means closer, like other distances. Chroma, for example, returns this distance, so lower is better.
  • "Is cosine slower than dot product?" — Only if you compute the lengths each time. Normalise once at write time, and cosine is a dot product.