Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 4: Similarity Metric Decision
What you need to know
The scenario: the team is creating a new vector index and must choose cosine, dot product or Euclidean distance.
The three metrics
| Metric | What it measures | Formula in words |
|---|---|---|
| Cosine similarity | Angle between vectors, ignores length | dot(a, b) divided by (length of a times length of b) |
| Dot (inner) product | Angle and length together | sum of a_i times b_i |
| Euclidean (L2) distance | Straight-line distance | square root of the sum of (a_i minus b_i) squared |
1import numpy as np23a, b = np.random.rand(768), np.random.rand(768)4a, b = a / np.linalg.norm(a), b / np.linalg.norm(b) # normalise once, at index time5cos = a @ b / (np.linalg.norm(a) * np.linalg.norm(b))6print(np.isclose(cos, a @ b)) # True: dot == cosine7print(np.isclose(np.sum((a - b) ** 2), 2 - 2 * cos)) # True: L2 squared == 2 - 2cosHow to decide
- Read the model card — it states the training objective and whether outputs are normalised.
- Normalise if the model expects cosine — then configure the index for inner product.
- Keep raw dot product only when length means something — some retrieval models encode popularity or confidence in the vector's length.
- Verify — measure recall@10 on a golden set with each candidate metric. If cosine and dot product disagree noticeably, your vectors are not normalised the way you assumed.
Euclidean distance is a natural choice for some image and audio embeddings from metric-learning setups; for text retrieval, cosine or normalised dot product is the norm.
Two operational rules
- In most vector stores, the metric is fixed when the index or collection is created. Changing it means a rebuild.
- The query must use the same metric and the same normalisation as the index. A mismatch does not crash; it quietly returns worse results.
A real-life example
Scenario, numbers made up. A team building a legal-search index copies an old config that uses Euclidean distance with unnormalised vectors from a new embedding model. Recall@10 on their 300-query golden set is 71%.
Reading the model card, they see the model is trained with a cosine objective and does not normalise outputs by default. They normalise at index time and switch the collection to inner product. Recall@10 becomes 83%, and search is slightly faster because inner product is cheaper to compute than distance. Rebuilding the index took four hours — a cost they could have avoided by checking before the first build.
Follow-up questions to expect
- "Is inner product faster than cosine?" — Yes, slightly: cosine needs the lengths, while inner product on pre-normalised vectors is just multiply and add.
- "When would dot product without normalising be right?" — When the model was trained with it and vector length carries a signal, such as item popularity in some recommendation models.
- "What happens if the metric is wrong?" — Nothing crashes; rankings get worse. That is why you verify with recall on a golden set.