Natural Language Processing Basics

Word Embeddings (Word2Vec, GloVe, FastText)


Two reviews arrive at your search system. One says "the food was excellent". The other says "the food was superb". A human reads them as saying the same thing. Ask a count-based representation how similar they are and it disagrees.

Python
from sklearn.feature_extraction.text import TfidfVectorizerfrom sklearn.metrics.pairwise import cosine_similarityv = TfidfVectorizer()X = v.fit_transform(["excellent", "superb", "terrible"])print(cosine_similarity(X[0], X[1]))  # excellent vs superb  -> [[0.]]print(cosine_similarity(X[0], X[2]))  # excellent vs terrible -> [[0.]]

Zero, and zero. As far as this representation is concerned, excellent is exactly as related to superb as it is to terrible — which is to say, not at all.

The reason is structural, not a tuning problem. In a one-hot or count representation, every word gets its own dimension. excellent is [0,…,1,…,0][0,\dots,1,\dots,0] with the 1 in position 4,821. superb has its 1 in position 11,203. Their dot product is zero because they never both have a non-zero value in the same slot. Every pair of distinct words is orthogonal. The geometry has no room to express that two words mean similar things.

You cannot fix this by weighting. You have to change what a word is.

The neighbourhood counting could never findexcellentsuperb — cosine 0.83outstanding — 0.79terrific — 0.74great — 0.71awful — close too:antonyms share contexts
None of these share a character with "excellent"; they are close because they appear in the same contexts.

The distributional hypothesis

Where would the meaning even come from? You are not going to hand-write a similarity score for every pair in a 50,000-word vocabulary — that is 1.25 billion pairs.

The idea that unlocks it is nearly a century old and stated most memorably by the linguist J.R. Firth: you shall know a word by the company it keeps.

Consider a word you have genuinely never seen:

Text
The chef poured the ganache over the cake.I ordered a dark ganache tart for dessert.This ganache is too bitter; add more cream.Warm the ganache until it is smooth and glossy.

You now know a great deal about ganache. It is a food, sweet-adjacent, chocolate-related, pourable when warm, used in desserts. Nobody defined it. You inferred all of it from the surrounding words — the contexts.

The distributional hypothesis: words that appear in similar contexts tend to have similar meanings. It is not obviously true, and it is not perfectly true, but it is true enough that it turns meaning into something you can compute from raw text with no labels at all.

If excellent and superb both appear near food, service, absolutely, really, and recommend, and neither appears near refund or disappointing, then a system that represents each word by its contexts will place them close together automatically.

What an embedding is

A word embedding is a dense vector of real numbers — typically 50 to 300 of them — that represents a word, learned so that words with similar contexts get similar vectors.

One-hot / countEmbedding
DimensionsVocabulary size (30k–500k)50–300
ValuesMostly zero, integersAll non-zero, real
Where it comes fromAssigned by an indexLearned from a corpus
Similar wordsOrthogonal — distance is meaninglessNearby — distance is meaningful
Unseen word<UNK><UNK>, unless subword-aware
Interpretable dimensionsYes — one word eachNo — individual dimensions mean nothing

Similarity is measured with cosine similarity, the cosine of the angle between two vectors:

cos(a,b)=a⋅b∥a∥ ∥b∥\text{cos}(a, b) = \frac{a \cdot b}{\|a\| \, \|b\|}

It ranges from −1-1 to 11 and ignores vector length, which matters because frequent words tend to develop larger vectors during training. You want to compare direction, not magnitude.

Concretely, with toy 3-dimensional vectors:

Text
excellent = [0.8, 0.6, 0.1]superb    = [0.7, 0.7, 0.2]terrible  = [-0.7, 0.5, -0.4]excellent . superb   = 0.56 + 0.42 + 0.02 = 1.00|excellent| = 1.005,  |superb| = 1.010cos = 1.00 / (1.005 * 1.010) = 0.985excellent . terrible = -0.56 + 0.30 - 0.04 = -0.30|terrible| = 0.949cos = -0.30 / (1.005 * 0.949) = -0.315

0.985 versus −0.315. The geometry now carries the meaning that the count representation could not express.

The famous arithmetic

Because directions in the space encode consistent relationships, vector arithmetic produces results that look uncanny:

Text
king - man + woman  ≈  queenparis - france + italy  ≈  romewalking - walked + swam  ≈  swimming

This is not magic and it is worth understanding why it happens. During training, the only reliable difference between the contexts of king and queen is gendered language — the same difference that separates man from woman. The model has no reason to encode that difference twice, so it settles on roughly one direction in the space that means "gendered male versus female". Subtracting man and adding woman moves you along that direction.

Be sceptical about how impressive this is. The standard evaluation excludes the three input words from the answer candidates, which does a lot of work. Many analogies fail. Treat it as evidence that the space has structure, not as proof of understanding.

Word2Vec: learn embeddings by predicting context

Word2Vec turns the distributional hypothesis into a prediction task. The trick is that the prediction itself is disposable — you only want the weights it learns along the way.

Skip-gram

Given a centre word, predict the words around it. Take the sentence and a window of 2:

Text
"the quick brown fox jumps over the lazy dog"                 ^centretraining pairs from centre = "fox", window = 2:  (fox, quick)  (fox, brown)  (fox, jumps)  (fox, over)

Slide the window across the whole corpus and you generate millions of these pairs, all from unlabelled text.

The model is deliberately tiny: a single hidden layer with no activation function.

Text
one-hot input (V)  ->  W_in (V x D)  ->  hidden (D)  ->  W_out (D x V)  ->  softmax (V)

Because the input is one-hot, multiplying by WinW_{in} just selects one row. That row is the word's embedding. The whole network is a lookup table with a training objective bolted on. When training finishes, you throw away WoutW_{out} and keep WinW_{in}.

Why the naive version is unusable, and how negative sampling fixes it

The softmax over the output layer normalises across the entire vocabulary. With V=100,000V = 100{,}000 and D=300D = 300, every single training pair requires 30 million multiply-adds for the output layer alone. Multiply by a corpus with a billion tokens and the training run never finishes.

Negative sampling replaces the multi-class softmax with a much cheaper binary question: is this (centre, context) pair real, or did I make it up?

Text
Real pair:      (fox, brown)     -> label 1Fake pairs:     (fox, database)  -> label 0                (fox, tuesday)   -> label 0                (fox, the)       -> label 0                (fox, plastic)   -> label 0                (fox, orbit)     -> label 0

The objective becomes logistic regression on 1 positive and kk negatives, typically k=5k = 5 for large corpora and k=15k = 15 for small ones. Cost per pair drops from V×D=30,000,000V \times D = 30{,}000{,}000 to (k+1)×D=1,800(k+1) \times D = 1{,}800 — roughly a 16,000-fold reduction, which is the difference between a research curiosity and a tool.

Negative words are drawn from a modified unigram distribution, P(w)∝f(w)0.75P(w) \propto f(w)^{0.75}. The exponent flattens the distribution so that very frequent words are sampled as negatives somewhat less often than raw frequency would dictate, and rare words somewhat more.

CBOW

Continuous bag-of-words runs the same architecture backwards: average the context vectors and predict the centre word.

Skip-gramCBOW
DirectionCentre → contextContext → centre
Training examples per window2×2 \times window size1
SpeedSlowerSeveral times faster
Rare wordsMuch better — each occurrence generates many updatesWeaker — rare words get averaged away in the context
Frequent wordsGoodSlightly better, smoothed by averaging
Use whenCorpus is small, or rare-word quality mattersCorpus is very large and time is tight

For most work, skip-gram with negative sampling is the default.

GloVe: fit the co-occurrence statistics directly

Word2Vec looks at one window at a time and never sees the corpus as a whole. GloVe takes the opposite route: build the full word-by-word co-occurrence matrix first, then find vectors that explain it.

The insight that motivates GloVe is about ratios. Take the probe words ice and steam, and count how often each appears near various other words:

Context word kkP(k∣ice)P(k \mid \text{ice})P(k∣steam)P(k \mid \text{steam})RatioWhat the ratio says
solid1.9×10−41.9 \times 10^{-4}2.2×10−52.2 \times 10^{-5}8.9Related to ice, not steam
gas6.6×10−56.6 \times 10^{-5}7.8×10−47.8 \times 10^{-4}0.085Related to steam, not ice
water3.0×10−33.0 \times 10^{-3}2.2×10−32.2 \times 10^{-3}1.36Related to both
fashion1.7×10−51.7 \times 10^{-5}1.8×10−51.8 \times 10^{-5}0.96Related to neither

Raw probabilities are dominated by how common each word is. The ratio cancels that out and isolates the discriminating information: values far from 1 in either direction are meaningful, values near 1 are not. GloVe's objective is designed so that the dot product of two word vectors approximates the log of their co-occurrence count, which makes vector differences correspond to these log-ratios.

J=∑i,jf(Xij)(wi⊤w~j+bi+b~j−log⁡Xij)2J = \sum_{i,j} f(X_{ij}) \left( w_i^\top \tilde{w}_j + b_i + \tilde{b}_j - \log X_{ij} \right)^2

The weighting function ff is the practical part: it caps the influence of extremely frequent pairs and zeroes out pairs that never co-occur, so the model is not dominated by the-of and does not attempt to take log⁡0\log 0.

FastText: words made of pieces

Both models so far share a hard limit. If a word was not in the training corpus, there is no row for it. Ask for a vector and you get an error or an <UNK>.

This bites constantly in practice: proper nouns, typos, technical terms, hashtags, and morphologically rich languages where a single stem can generate dozens of surface forms.

FastText represents each word as the sum of vectors for its character n-grams, plus a vector for the word itself.

Text
"where" with n-grams of length 3-6 (angle brackets mark word boundaries):  <wh, whe, her, ere, re>,  <whe, wher, here, ere>,  <wher, where, here>,  <where, where>  plus the whole-word vector <where>vector("where") = sum of all of the above

The consequence is that an unseen word still gets a sensible vector. unbelievableness may never have appeared, but believ, able, ness and un all have. Misspellings land near their correct forms because they share most of their n-grams.

Word2VecGloVeFastText
Training signalLocal windows, predictiveGlobal co-occurrence countsLocal windows, subword units
Handles unseen wordsNoNoYes
Rare word qualityModerateModerateStrong
Morphology-rich languagesWeakWeakStrong
MemoryV×DV \times DV×DV \times DLarger — n-gram hash buckets too
Training speedFastFast once the matrix is builtSlower
Best forClean, large, single-domain corporaGeneral-purpose pretrained vectorsNoisy text, typos, non-English

Training and using them

Python
from gensim.models import Word2Vec, FastTextsentences = [["the", "food", "was", "excellent"],             ["the", "food", "was", "superb"],             ["service", "was", "terrible"]]   # in reality: millionsmodel = Word2Vec(    sentences,    vector_size=100,   # D    window=5,          # context radius    min_count=1,       # ignore rarer words; use 5 or more on a real corpus    sg=1,              # 1 = skip-gram, 0 = CBOW    negative=10,       # negative samples per positive    workers=4,    epochs=10,)print(model.wv["excellent"].shape)            # (100,)print(model.wv.most_similar("excellent", topn=5))print(model.wv.similarity("excellent", "superb"))

Loading pretrained vectors is usually the better move, and takes one line:

Python
import gensim.downloader as apiglove = api.load("glove-wiki-gigaword-100")   # 400k words, 100 dims, ~130 MBprint(glove.most_similar(positive=["king", "woman"], negative=["man"], topn=3))# [('queen', 0.77), ('monarch', 0.68), ('throne', 0.68)]print(glove.doesnt_match(["breakfast", "lunch", "dinner", "cricket"]))# 'cricket'

FastText's defining behaviour, shown directly:

Python
ft = FastText(sentences, vector_size=100, window=5, min_count=1,              min_n=3, max_n=6, sg=1, epochs=10)print("excellentness" in ft.wv.key_to_index)   # False - never seenprint(ft.wv["excellentness"].shape)            # (100,) - vector anywayprint(ft.wv.most_similar("excellentness", topn=3))

Turning word vectors into document vectors

The simplest usable approach is to average the vectors of a document's words. It is crude — averaging destroys order exactly the way bag-of-words does — but it is a strong, fast baseline.

Python
import numpy as npfrom sklearn.linear_model import LogisticRegressiondef doc_vector(tokens, kv):    vecs = [kv[t] for t in tokens if t in kv]    if not vecs:        return np.zeros(kv.vector_size)    return np.mean(vecs, axis=0)X = np.vstack([doc_vector(d, glove) for d in tokenised_docs])clf = LogisticRegression(max_iter=1000).fit(X, y)

Note what you have gained and lost against TF-IDF. You now have 100 dense features instead of hundreds of thousands of sparse ones, and a review saying superb lands near one saying excellent even if neither word appeared in the other's training data. But you have thrown away order and, crucially, you have thrown away emphasis — averaging gives the the same weight as appalling. Weighting each word vector by its TF-IDF score before averaging usually recovers a point or two.

Where this breaks

One vector per word, forever

This is the fundamental limitation and it is worth stating plainly. bank gets exactly one vector, which must simultaneously serve river bank and investment bank. The training process averages the two senses into a single compromise point that represents neither well.

Python
print([w for w, _ in glove.most_similar("bank", topn=6)])# ['banks', 'banking', 'credit', 'investment', 'financial', 'securities']

Every neighbour is financial. In news and Wikipedia text the money sense is far more common, so it has swallowed the river sense: a sentence about a muddy river bank gets a vector that says finance. No amount of extra training data fixes this, because the representation has no slot for context. Contextual models — the kind built on transformers — solve it by producing a different vector for each occurrence of the word, which is the single largest reason they displaced static embeddings.

The vectors absorb whatever is in the corpus

Embeddings learn statistical regularities, and human text is full of regularities we would not endorse. The widely-reproduced result is that man : computer programmer :: woman : homemaker falls out of vectors trained on news text. Nobody intended it; the corpus contained it.

If your embeddings feed a system that screens CVs, ranks applicants, or moderates content, this is not an academic concern. Audit the associations that matter for your application before deploying, and know that debiasing techniques reduce measured bias without reliably removing it.

Training your own on a small corpus

Word2Vec needs a lot of text. On a corpus of 10,000 documents, most words appear a handful of times and their vectors are essentially noise. The result looks fine — most_similar returns something — but the neighbours are arbitrary.

A rough guide: below roughly 10 million tokens, use pretrained vectors. Above 100 million domain-specific tokens, training your own is likely to beat generic pretrained vectors on domain terms. In between, start pretrained and fine-tune.

Trusting the similarity numbers too much

Cosine similarity reflects distributional similarity, not synonymy. Antonyms appear in near-identical contexts — hot and cold both precede weather, water, day — so they end up close together.

Python
print(glove.similarity("good", "bad"))   # ~0.77 - very highprint(glove.similarity("good", "cake"))  # ~0.33

A sentiment system that treats cosine distance as semantic agreement will conclude that good and bad are near-synonyms. They are, in the only sense the model was trained to measure. Know what your metric actually measures before you build on it.

Choosing, in practice

Your situationDo this
General English, no special domainDownload pretrained GloVe (100d or 300d) and move on
Noisy text — social media, user input, typosFastText, pretrained or trained on your data
Large domain corpus (legal, medical, logs)Train your own skip-gram; domain terms will be far better
Non-English, especially agglutinative languagesFastText, without hesitation
Polysemy or nuance is central to the taskStatic embeddings are the wrong tool; use a contextual model
Millisecond latency budget, CPU onlyStatic embeddings — a lookup table is unbeatable on speed

When you wire embeddings into a neural network, one decision comes up immediately: initialise the embedding layer with pretrained vectors, then decide whether to freeze it. Freeze when your labelled dataset is small — a few thousand examples will overfit 30,000 × 300 = 9 million embedding parameters within an epoch. Unfreeze when you have hundreds of thousands of labelled examples and your domain uses words differently from the pretraining corpus. A reliable middle path is to freeze for the first two or three epochs while the rest of the network stabilises, then unfreeze with a learning rate an order of magnitude lower than the rest of the model.

The reason embeddings matter is not that they are clever. It is that they let a model generalise across words it has never seen together. Train on reviews containing excellent and the model handles superb at inference — because the corpus already told it, without anyone labelling anything, that those two words live in the same place.