Course Content
LLMs Deep Dive
10 sections · 40 lessons
What role do embeddings play in LLMs?
What you need to know
The embedding table
The table has one row per token. For a model with a 128,000-token vocabulary and 4,096 dimensions:
128,000 rows x 4,096 columns = about 524 million parametersLooking up token ID 9123 just returns row 9123. The numbers start random and are learned by backpropagation, so tokens used in similar contexts end up with similar rows.
Four places embeddings appear in an LLM
- Input embeddings — the lookup above.
- Position information — older models added a position vector to each token vector. Most modern LLMs use RoPE instead, which rotates the query and key vectors inside attention rather than adding anything to the embedding.
- Contextual hidden states — after each layer, a token's vector has absorbed information from the tokens around it. This is how "it" comes to carry the meaning of the noun it refers to.
- Output projection — the final vector is multiplied by a matrix of vocabulary size to give one score (logit) per token. Some models reuse the input table here (weight tying) to save parameters.
Similarity with small numbers
Cosine similarity measures the angle between two vectors: 1 means same direction, 0 means unrelated.
1import numpy as np23def cos(a, b):4 return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))56sneakers = np.array([0.9, 0.1, 0.3])7trainers = np.array([0.8, 0.2, 0.35])8laptop = np.array([0.1, 0.9, 0.0])910print(round(cos(sneakers, trainers), 2)) # 0.99 -> near-synonyms11print(round(cos(sneakers, laptop), 2)) # 0.21 -> unrelatedReal embeddings have hundreds or thousands of dimensions, but the maths is the same.
Model embeddings versus embedding models
The vectors inside a chat LLM are per token and tuned for predicting the next token. For search you use a separate embedding model trained so that a whole sentence maps to one vector, and matching question–answer pairs land close together.
A real-life example
An e-commerce search assistant used keyword search. A shopper typing "running shoes for flat feet" got zero results, because the catalogue says "stability trainers with arch support".
The team embeds every product description once (2 million products, each a 768-dimension vector stored in a vector index). At query time, they embed the shopper's text and fetch the 50 nearest products by cosine similarity. "Running shoes for flat feet" now lands next to "stability trainers" because both appeared in similar contexts during the embedding model's training. Keyword search is kept alongside for exact matches like model numbers ("ASX-209"), which embeddings handle poorly — this mix is called hybrid search.
Follow-up questions to expect
- "What is the difference between static and contextual embeddings?" — Static ones (word2vec, GloVe) give a word one vector everywhere; contextual ones (from transformer layers) change with the sentence, so "bank" has different vectors in different contexts.
- "Why cosine similarity rather than Euclidean distance?" — Cosine ignores vector length and compares direction, which is more stable for meaning. For normalised vectors the two give the same ranking.
- "What dimension should embeddings have?" — A trade-off: more dimensions capture more nuance but cost more storage and search time. Many embedding models let you truncate vectors to a shorter length with a small quality loss.