LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What role do embeddings play in LLMs?


Cosine similarity between three toy embeddings1.000.990.210.991.000.320.210.321.00sneakerstrainerslaptopsneakerstrainerslaptop3-dimension vectors from the lesson; real ones have hundreds of dimensions.
Similarity is an angle, not a shared word — which is how 'flat feet' finds 'arch support' in the catalogue.

What you need to know

The embedding table

The table has one row per token. For a model with a 128,000-token vocabulary and 4,096 dimensions:

Text
128,000 rows x 4,096 columns = about 524 million parameters

Looking up token ID 9123 just returns row 9123. The numbers start random and are learned by backpropagation, so tokens used in similar contexts end up with similar rows.

Four places embeddings appear in an LLM

  • Input embeddings — the lookup above.
  • Position information — older models added a position vector to each token vector. Most modern LLMs use RoPE instead, which rotates the query and key vectors inside attention rather than adding anything to the embedding.
  • Contextual hidden states — after each layer, a token's vector has absorbed information from the tokens around it. This is how "it" comes to carry the meaning of the noun it refers to.
  • Output projection — the final vector is multiplied by a matrix of vocabulary size to give one score (logit) per token. Some models reuse the input table here (weight tying) to save parameters.

Similarity with small numbers

Cosine similarity measures the angle between two vectors: 1 means same direction, 0 means unrelated.

Python
import numpy as npdef cos(a, b):    return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))sneakers = np.array([0.9, 0.1, 0.3])trainers = np.array([0.8, 0.2, 0.35])laptop   = np.array([0.1, 0.9, 0.0])print(round(cos(sneakers, trainers), 2))  # 0.99 -> near-synonymsprint(round(cos(sneakers, laptop), 2))    # 0.21 -> unrelated

Real embeddings have hundreds or thousands of dimensions, but the maths is the same.

Model embeddings versus embedding models

The vectors inside a chat LLM are per token and tuned for predicting the next token. For search you use a separate embedding model trained so that a whole sentence maps to one vector, and matching question–answer pairs land close together.

A real-life example

An e-commerce search assistant used keyword search. A shopper typing "running shoes for flat feet" got zero results, because the catalogue says "stability trainers with arch support".

The team embeds every product description once (2 million products, each a 768-dimension vector stored in a vector index). At query time, they embed the shopper's text and fetch the 50 nearest products by cosine similarity. "Running shoes for flat feet" now lands next to "stability trainers" because both appeared in similar contexts during the embedding model's training. Keyword search is kept alongside for exact matches like model numbers ("ASX-209"), which embeddings handle poorly — this mix is called hybrid search.

Follow-up questions to expect

  • "What is the difference between static and contextual embeddings?" — Static ones (word2vec, GloVe) give a word one vector everywhere; contextual ones (from transformer layers) change with the sentence, so "bank" has different vectors in different contexts.
  • "Why cosine similarity rather than Euclidean distance?" — Cosine ignores vector length and compares direction, which is more stable for meaning. For normalised vectors the two give the same ranking.
  • "What dimension should embeddings have?" — A trade-off: more dimensions capture more nuance but cost more storage and search time. Many embedding models let you truncate vectors to a shorter length with a small quality loss.