Course Content
Natural Language Processing Basics
4 sections · 10 lessons
Bag-of-Words & TF-IDF Representations
You have 50,000 product reviews and a logistic regression model that refuses to accept anything but numbers. So you need to turn text into numbers. The first idea almost everyone has is the obvious one: give every word an integer.
vocab = {"terrible": 1, "bad": 2, "okay": 3, "good": 4, "excellent": 5}review = "good" # -> [4]review = "terrible" # -> [1]Feed that to a linear model and watch what it assumes. The model computes w⋅x. If excellent is 5 and terrible is 1, the model believes excellent is five times terrible, and that okay (3) sits exactly halfway between them. Sometimes, by pure luck, that ordering is meaningful. Now try it with a real vocabulary:
vocab = {"aardvark": 1, "abacus": 2, ..., "cat": 4821, "dog": 4822, ...}The model now believes dog is one more than cat, and that cat is roughly 2,400 times abacus. These numbers were assigned alphabetically. They encode nothing except spelling, and the model will happily fit patterns to that noise.
This is the ordinal fallacy: using integers as identifiers while a model reads them as quantities. Word identity is categorical, not numeric, and the representation has to reflect that.
One-hot vectors, and why they lead straight to counting
The standard fix for categorical data is a one-hot vector: a vector as long as your vocabulary, with a 1 in the position for that word and 0 everywhere else.
vocabulary = [amazing, bad, film, the, was]"film" -> [0, 0, 1, 0, 0]"bad" -> [0, 1, 0, 0, 0]No word is now larger than any other. All of them are equidistant. But a document is many words, and a model needs one fixed-size vector per document, not a variable number of them. The simplest way to combine them is to add them up.
"the film was bad" the -> [0, 0, 0, 1, 0] film -> [0, 0, 1, 0, 0] was -> [0, 0, 0, 0, 1] bad -> [0, 1, 0, 0, 0] sum = [0, 1, 1, 1, 1]Sum the one-hot vectors of every word in a document and you get a vector of counts. That is the bag-of-words model, arrived at from first principles.
Bag-of-words, worked end to end
Take three tiny documents:
D1: "the film was good"D2: "the film was bad bad"D3: "the acting was good"Step one is to collect the vocabulary — every distinct word, in a fixed order:
[acting, bad, film, good, the, was]Step two is to count occurrences of each vocabulary word in each document:
| acting | bad | film | good | the | was | |
|---|---|---|---|---|---|---|
| D1 | 0 | 0 | 1 | 1 | 1 | 1 |
| D2 | 0 | 2 | 1 | 0 | 1 | 1 |
| D3 | 1 | 0 | 0 | 1 | 1 | 1 |
That is the document-term matrix. Three documents, six features, all numeric. Any model that takes a feature vector can now consume text.
The name is precise and worth taking literally. It really is a bag — you tipped all the words in and shook it. The order is gone.
"dog bites man" -> {bites: 1, dog: 1, man: 1}"man bites dog" -> {bites: 1, dog: 1, man: 1} identicalBag-of-words discards word order completely. For topic classification that is nearly free — a document about football contains football words regardless of arrangement. For anything where order carries meaning, it is a hard ceiling you cannot train your way past.
The frequency problem
Look at the matrix above again and ask which column is useful. the is 1 in every document. was is 1 in every document. They cost you two of six features and separate nothing.
Scale that to real text and it gets worse. In a corpus of film reviews, the most frequent words are roughly:
| Word | Total count | Documents containing it | Useful for classification? |
|---|---|---|---|
the | 336,000 | 24,900 / 25,000 | No |
movie | 44,000 | 18,200 / 25,000 | Barely — it is a film corpus |
good | 21,000 | 9,800 / 25,000 | Yes |
tedious | 310 | 295 / 25,000 | Yes, strongly |
Raw counts give the a value of 336,000 across the corpus and tedious a value of 310. Any model sensitive to feature magnitude — and most are — will be dominated by words that carry no information. You could delete stopwords by hand, but that only handles the extreme cases and requires you to know in advance which words are uninformative in your corpus. In a corpus of medical papers, patient is a stopword. In a general corpus it is not.
You want a principled, automatic way to say: frequent within this document is good evidence; frequent across all documents is not. That is exactly what TF-IDF computes.
TF-IDF from the two ideas it is made of
The score is a product of two terms.
Term frequency
How often the word appears in this document. The simplest form is the raw count, but dividing by document length stops long documents from dominating:
A common refinement is sublinear scaling, 1+log(count), on the reasoning that a word appearing 20 times is not 20 times more relevant than one appearing once.
Inverse document frequency
How rare the word is across the whole corpus. With N documents and df(t) documents containing t:
Ask why there is a logarithm. Without it, with N=25,000, a word in one document scores 25,000 and a word in 100 documents scores 250 — a hundredfold gap that overwhelms everything else. The log compresses that to 10.1 versus 5.5, a difference that is meaningful without being catastrophic. Rarity should be rewarded on a sliding scale, not explosively.
In practice both scikit-learn and most implementations add smoothing so that a term appearing in every document gets a small positive weight rather than exactly zero, and so an unseen term does not divide by zero:
The product
Read the four cases and the design becomes obvious:
| Situation | tf | idf | tfidf | Interpretation |
|---|---|---|---|---|
Common in this doc, common everywhere (the) | high | ≈0 | low | Not distinctive |
Common in this doc, rare elsewhere (tedious) | high | high | high | Strong signal for this doc |
| Rare in this doc, rare elsewhere | low | high | medium | Weak but noteworthy |
| Absent from this doc | 0 | — | 0 | No contribution |
A calculation with real numbers
Use the three documents from earlier, N=3, with the smoothed formula.
Document frequencies: the 3, was 3, good 2, film 2, bad 1, acting 1.
| Term | df | idf = log((1+3)/(1+df)) + 1 |
|---|---|---|
the | 3 | log(4/4) + 1 = 0 + 1 = 1.000 |
was | 3 | 1.000 |
film | 2 | log(4/3) + 1 = 0.288 + 1 = 1.288 |
good | 2 | 1.288 |
bad | 1 | log(4/2) + 1 = 0.693 + 1 = 1.693 |
acting | 1 | 1.693 |
Now take D2, "the film was bad bad", using raw counts for tf:
| Term | count | idf | tf × idf | After L2 normalisation |
|---|---|---|---|---|
bad | 2 | 1.693 | 3.386 | 0.871 |
film | 1 | 1.288 | 1.288 | 0.331 |
the | 1 | 1.000 | 1.000 | 0.257 |
was | 1 | 1.000 | 1.000 | 0.257 |
The L2 norm is 3.3862+1.2882+12+12=11.47+1.66+1+1=3.889, and each value is divided by it so the document vector has length 1. That normalisation matters: without it, a 2,000-word review would have vastly larger values than a 50-word one purely because it is longer, and a distance-based model would cluster documents by length rather than by content.
Read the result. In a document of five words, bad now has a weight of 0.87 and the only 0.26. That is the ranking you wanted, produced automatically, with no hand-written stopword list.
TF-IDF is not a model and it learns nothing. It is a weighting scheme that encodes one assumption — a word is informative about a document in proportion to how concentrated it is in that document relative to the corpus. That assumption is simple, cheap, and remarkably hard to beat on topical tasks.
Doing it in scikit-learn
1from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer23docs = ["the film was good",4 "the film was bad bad",5 "the acting was good"]67cv = CountVectorizer()8X = cv.fit_transform(docs)9print(cv.get_feature_names_out())10# ['acting' 'bad' 'film' 'good' 'the' 'was']11print(X.toarray())12# [[0 0 1 1 1 1]13# [0 2 1 0 1 1]14# [1 0 0 1 1 1]]1516tv = TfidfVectorizer()17T = tv.fit_transform(docs)18print(T.toarray().round(3))19# [[0. 0. 0.558 0.558 0.434 0.434]20# [0. 0.871 0.331 0. 0.257 0.257]21# [0.663 0. 0. 0.504 0.391 0.391]]The middle row matches the hand calculation. The parameters that actually matter:
| Parameter | What it does | Sensible starting value |
|---|---|---|
min_df | Ignore terms in fewer than this many documents | 2 or 5 — kills typos and one-off tokens |
max_df | Ignore terms in more than this fraction of documents | 0.9 — an automatic, corpus-specific stopword list |
max_features | Keep only the top-N by frequency | 20000–50000 |
ngram_range | Include multi-word features | (1, 2) |
sublinear_tf | Use 1+log(tf) instead of raw tf | True for long documents |
norm | Vector normalisation | 'l2' — leave it alone |
Implementing it by hand
Worth doing once, because it removes all mystery about what the library is doing:
1import math2from collections import Counter34def tfidf(docs):5 tokenised = [d.lower().split() for d in docs]6 vocab = sorted({w for d in tokenised for w in d})7 N = len(tokenised)89 df = Counter()10 for d in tokenised:11 df.update(set(d))1213 idf = {t: math.log((1 + N) / (1 + df[t])) + 1 for t in vocab}1415 matrix = []16 for d in tokenised:17 counts = Counter(d)18 row = [counts[t] * idf[t] for t in vocab]19 norm = math.sqrt(sum(v * v for v in row)) or 1.020 matrix.append([v / norm for v in row])21 return vocab, matrix2223vocab, M = tfidf(["the film was good",24 "the film was bad bad",25 "the acting was good"])26print(vocab)27print([round(v, 3) for v in M[1]])28# [0.0, 0.871, 0.331, 0.0, 0.257, 0.257]N-grams: buying back a little word order
Bag-of-words cannot tell not good from good. An n-gram treats each contiguous run of n tokens as its own feature, which recovers local order.
"the film was not good"unigrams: the, film, was, not, goodbigrams: the film, film was, was not, not goodNow not good is a single feature, and a classifier can learn a negative weight for it while keeping a positive weight for good. This is the cheapest available fix for negation in a bag-of-words pipeline, and it works well.
The cost is vocabulary size. With a unigram vocabulary of V words, the theoretical bigram space is V2:
| ngram_range | Features (25k IMDb reviews, min_df=2) | Typical accuracy gain | Fit time |
|---|---|---|---|
(1, 1) | ~45,000 | baseline | 1× |
(1, 2) | ~440,000 | +2 to +3 points | ~3× |
(1, 3) | ~900,000 | +0 to +0.5 more | ~6× |
Bigrams almost always pay for themselves. Trigrams almost never do — most trigrams appear once, get pruned by min_df, and the survivors add little. (1, 2) with min_df=2 is the default worth reaching for.
Sparsity, and why it is fine
A document-term matrix for 25,000 reviews with a 440,000-feature bigram vocabulary has 11 billion cells. Stored densely as 64-bit floats that would be 88 gigabytes, but an average review touches only about 300 of those features. Over 99.9% of the matrix is zero.
Scikit-learn returns a scipy.sparse matrix that stores only the non-zeros. Two consequences follow:
- Never call
.toarray()on a real corpus. It will try to materialise every zero and exhaust your memory. - Prefer models that accept sparse input directly —
LogisticRegression,LinearSVC,MultinomialNB,SGDClassifier. Tree ensembles and neural networks generally want dense input, which is a large part of why linear models remain the standard partner for TF-IDF.
The mistakes that cost real accuracy
Fitting the vectoriser on all your data
This is the single most common bug, and it inflates your reported score without ever failing loudly.
1# WRONG - the vectoriser has seen the test set2X = TfidfVectorizer().fit_transform(all_docs)3X_train, X_test = train_test_split(X, ...)45# RIGHT - fit on train only, transform test with what was learned6X_train_txt, X_test_txt, y_train, y_test = train_test_split(docs, labels)7vec = TfidfVectorizer(min_df=2, ngram_range=(1, 2))8X_train = vec.fit_transform(X_train_txt)9X_test = vec.transform(X_test_txt)In the wrong version, the idf values are computed using document frequencies that include the test set. Your test documents influenced their own feature weights. The measured accuracy is optimistic and the model degrades when it meets genuinely new data. Use a Pipeline and cross-validation and this becomes impossible to get wrong:
1from sklearn.pipeline import make_pipeline2from sklearn.linear_model import LogisticRegression3from sklearn.model_selection import cross_val_score45pipe = make_pipeline(6 TfidfVectorizer(min_df=2, ngram_range=(1, 2), sublinear_tf=True),7 LogisticRegression(max_iter=1000, C=1.0),8)9print(cross_val_score(pipe, docs, labels, cv=5, scoring="accuracy").mean())Assuming unseen words do something
Any word not in the fitted vocabulary is silently dropped at transform time. If your training corpus is from 2018 and your production traffic is from today, a growing fraction of every incoming document contributes nothing at all. Monitor the proportion of out-of-vocabulary tokens in live traffic; a rising number is your early warning that the vectoriser needs refitting.
Reading feature weights as causes
A linear model over TF-IDF features is pleasantly inspectable:
1import numpy as np2names = vec.get_feature_names_out()3coefs = clf.coef_[0]4top = np.argsort(coefs)5print("most negative:", names[top[:10]])6print("most positive:", names[top[-10:]])That is genuinely useful for debugging — if the or a stray HTML tag shows up in the top ten, your preprocessing is broken. But a high weight means the feature correlates with the label in your training data, not that it causes anything. If every positive review in your scrape came from one site that uses a particular template word, that word will top the list.
When this is still the right tool
It is tempting to treat TF-IDF as a historical curiosity now that pretrained transformers exist. That is a mistake, and the reason is economic rather than sentimental.
| TF-IDF + linear model | Fine-tuned transformer | |
|---|---|---|
| Training time (25k docs) | Seconds, on a laptop | Tens of minutes, on a GPU |
| Inference latency | Under a millisecond, CPU | 10–100 ms, GPU preferred |
| Model size | A few MB | 400 MB+ |
| Labelled data needed | Works from a few thousand | Works from hundreds, better with more |
| Topical classification accuracy | Strong | Slightly stronger |
| Tasks needing word order or nuance | Weak | Much stronger |
| Explains its own decisions | Directly, per feature | Only with extra tooling |
Two practical habits follow. First, build the TF-IDF baseline before anything else, always. It takes ten minutes and it tells you what the task's floor looks like. A transformer that beats it by half a point is not worth the operational cost; a transformer that beats it by fifteen points has told you the task genuinely needs semantics and order.
Second, when your data has clear topical vocabulary — spam filtering, routing support tickets to departments, tagging news by section, deduplicating documents, retrieval over a fixed corpus — TF-IDF is frequently not just adequate but the correct final answer. It is fast, it runs anywhere, it has no dependency on a GPU, and when it makes a mistake you can point at the feature that caused it.