Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 4: Similarity Metric Decision


What you need to know

The scenario: the team is creating a new vector index and must choose cosine, dot product or Euclidean distance.

The three metrics

MetricWhat it measuresFormula in words
Cosine similarityAngle between vectors, ignores lengthdot(a, b) divided by (length of a times length of b)
Dot (inner) productAngle and length togethersum of a_i times b_i
Euclidean (L2) distanceStraight-line distancesquare root of the sum of (a_i minus b_i) squared
Python
import numpy as npa, b = np.random.rand(768), np.random.rand(768)a, b = a / np.linalg.norm(a), b / np.linalg.norm(b)      # normalise once, at index timecos = a @ b / (np.linalg.norm(a) * np.linalg.norm(b))print(np.isclose(cos, a @ b))                                  # True: dot == cosineprint(np.isclose(np.sum((a - b) ** 2), 2 - 2 * cos))           # True: L2 squared == 2 - 2cos

How to decide

  1. Read the model card — it states the training objective and whether outputs are normalised.
  2. Normalise if the model expects cosine — then configure the index for inner product.
  3. Keep raw dot product only when length means something — some retrieval models encode popularity or confidence in the vector's length.
  4. Verify — measure recall@10 on a golden set with each candidate metric. If cosine and dot product disagree noticeably, your vectors are not normalised the way you assumed.

Euclidean distance is a natural choice for some image and audio embeddings from metric-learning setups; for text retrieval, cosine or normalised dot product is the norm.

Two operational rules

  • In most vector stores, the metric is fixed when the index or collection is created. Changing it means a rebuild.
  • The query must use the same metric and the same normalisation as the index. A mismatch does not crash; it quietly returns worse results.

A real-life example

Scenario, numbers made up. A team building a legal-search index copies an old config that uses Euclidean distance with unnormalised vectors from a new embedding model. Recall@10 on their 300-query golden set is 71%.

Reading the model card, they see the model is trained with a cosine objective and does not normalise outputs by default. They normalise at index time and switch the collection to inner product. Recall@10 becomes 83%, and search is slightly faster because inner product is cheaper to compute than distance. Rebuilding the index took four hours — a cost they could have avoided by checking before the first build.

Follow-up questions to expect

  • "Is inner product faster than cosine?" — Yes, slightly: cosine needs the lengths, while inner product on pre-normalised vectors is just multiply and add.
  • "When would dot product without normalising be right?" — When the model was trained with it and vector length carries a signal, such as item popularity in some recommendation models.
  • "What happens if the metric is wrong?" — Nothing crashes; rankings get worse. That is why you verify with recall on a golden set.