Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Which similarity metric should your RAG system use, and why don't the scores mean what people think?
What you need to know
The three metrics
| Metric | Formula in words | Sensitive to length? |
|---|---|---|
| Cosine similarity | dot product divided by the product of the two lengths | No: compares direction only |
| Dot product | sum of element-wise products | Yes: longer vectors score higher |
| Euclidean (L2) distance | straight-line distance between the points | Yes |
Why they agree on normalised vectors
For two vectors of length 1:
distance(a, b) squared = 2 - 2 x cosine(a, b)dot(a, b) = cosine(a, b)Higher cosine always means smaller distance and higher dot product, so all three produce the same top-k. Dot product skips the division, so it is slightly cheaper, which is why many vector databases recommend it for normalised embeddings. Check rather than assume:
import numpy as npnorms = np.linalg.norm(embeddings, axis=1)print(norms.min(), norms.max()) # both close to 1.0 means already normalisedWhen magnitude does matter
In recommendation systems, a vector's length can encode popularity or confidence. There, dot product deliberately rewards it and is the right choice. For semantic text search, length mostly tracks things like text length, which you do not want to reward.
The threshold trap
People write "if score < 0.8, say I don't know". That is fragile for three reasons:
- Scores are model-specific. Many embedding models place almost all real text pairs between about 0.6 and 0.9 cosine. A 0.8 may be a strong match for one model and a weak one for another.
- Scores shift when you change models. Upgrade the embedding model and the old threshold silently means something else.
- Scores are not probabilities. 0.82 does not mean "82% relevant".
What to do instead:
- Label — 200 query-chunk pairs marked relevant or not.
- Plot — the score distributions for relevant and irrelevant pairs.
- Pick — the threshold that gives the precision or recall you need.
- Re-calibrate — every time the embedding model changes.
Or rely on a cross-encoder reranker. It reads the query and chunk together and is trained to score relevance, so its scores separate relevant from irrelevant more cleanly, though they still need a calibrated cut-off.
A real-life example
Scenario (illustrative numbers). A food-delivery company's help bot refuses to answer when the top chunk scores below 0.80. After switching to a newer embedding model, refusals jump from 6% to 31% overnight, although retrieval quality on the eval set actually improved.
The team plots scores on 300 labelled pairs. With the new model, relevant pairs sit between 0.55 and 0.75, and irrelevant pairs between 0.30 and 0.52. The old 0.80 threshold rejects almost everything. They set the new cut-off at 0.54, which keeps 97% of relevant pairs and rejects 95% of irrelevant ones, and add a calibration step to the embedding-upgrade checklist. Refusals drop to 5%.
Follow-up questions to expect
- "Which metric should I set in my vector database?" — Match what the embedding model was trained for, usually cosine; if vectors are normalised, dot product gives identical ranking faster.
- "Can I compare scores across two different embedding models?" — No. Each model has its own score distribution; compare ranking quality on an eval set instead.
- "Why is a reranker score better for thresholds?" — It is trained on relevance judgements for query-document pairs, so its scale tracks relevance more directly than a geometric similarity does.