Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Which similarity metric should your RAG system use, and why don't the scores mean what people think?


What you need to know

The three metrics

MetricFormula in wordsSensitive to length?
Cosine similaritydot product divided by the product of the two lengthsNo: compares direction only
Dot productsum of element-wise productsYes: longer vectors score higher
Euclidean (L2) distancestraight-line distance between the pointsYes

Why they agree on normalised vectors

For two vectors of length 1:

Text
distance(a, b) squared = 2 - 2 x cosine(a, b)dot(a, b)              = cosine(a, b)

Higher cosine always means smaller distance and higher dot product, so all three produce the same top-k. Dot product skips the division, so it is slightly cheaper, which is why many vector databases recommend it for normalised embeddings. Check rather than assume:

Python
import numpy as npnorms = np.linalg.norm(embeddings, axis=1)print(norms.min(), norms.max())      # both close to 1.0 means already normalised

When magnitude does matter

In recommendation systems, a vector's length can encode popularity or confidence. There, dot product deliberately rewards it and is the right choice. For semantic text search, length mostly tracks things like text length, which you do not want to reward.

The threshold trap

People write "if score < 0.8, say I don't know". That is fragile for three reasons:

  • Scores are model-specific. Many embedding models place almost all real text pairs between about 0.6 and 0.9 cosine. A 0.8 may be a strong match for one model and a weak one for another.
  • Scores shift when you change models. Upgrade the embedding model and the old threshold silently means something else.
  • Scores are not probabilities. 0.82 does not mean "82% relevant".

What to do instead:

  1. Label — 200 query-chunk pairs marked relevant or not.
  2. Plot — the score distributions for relevant and irrelevant pairs.
  3. Pick — the threshold that gives the precision or recall you need.
  4. Re-calibrate — every time the embedding model changes.

Or rely on a cross-encoder reranker. It reads the query and chunk together and is trained to score relevance, so its scores separate relevant from irrelevant more cleanly, though they still need a calibrated cut-off.

A real-life example

Scenario (illustrative numbers). A food-delivery company's help bot refuses to answer when the top chunk scores below 0.80. After switching to a newer embedding model, refusals jump from 6% to 31% overnight, although retrieval quality on the eval set actually improved.

The team plots scores on 300 labelled pairs. With the new model, relevant pairs sit between 0.55 and 0.75, and irrelevant pairs between 0.30 and 0.52. The old 0.80 threshold rejects almost everything. They set the new cut-off at 0.54, which keeps 97% of relevant pairs and rejects 95% of irrelevant ones, and add a calibration step to the embedding-upgrade checklist. Refusals drop to 5%.

Follow-up questions to expect

  • "Which metric should I set in my vector database?" — Match what the embedding model was trained for, usually cosine; if vectors are normalised, dot product gives identical ranking faster.
  • "Can I compare scores across two different embedding models?" — No. Each model has its own score distribution; compare ranking quality on an eval set instead.
  • "Why is a reranker score better for thresholds?" — It is trained on relevance judgements for query-document pairs, so its scale tracks relevance more directly than a geometric similarity does.