Course Content
Natural Language Processing Basics
4 sections · 10 lessons
Mini-Project: IMDb Sentiment Analysis System
A notebook cell that prints val_accuracy: 0.891 is not a sentiment analysis system. It is a number. The system is the thing that takes a string of text over HTTP and returns a label and a confidence, correctly, on input nobody has seen — including input from next month.
The gap between those two things is where this project lives. You will build a classifier for IMDb movie reviews, and the interesting work is not the model architecture. It is the decisions on either side of it: how long to truncate reviews, what to do with a word the vocabulary has never seen, whether 0.5 is the right threshold, and how to package the model so that the thing running in production computes the same function as the thing you validated.
The specification
Build an end-to-end binary sentiment classifier over film reviews and expose it as an inference service.
| Deliverable | Requirement |
|---|---|
| Data pipeline | Load raw reviews, clean, tokenise, build vocabulary, index |
| Model | Bidirectional multi-layer LSTM with attention pooling over variable-length input |
| Accuracy | ≥ 88% on a held-out validation set |
| Length handling | Correct results on 20-word and 800-word reviews alike |
| Augmentation | At least one technique implemented and honestly measured |
| Baseline comparison | A fine-tuned transformer, scored on the same split |
| Service | HTTP endpoint returning label and probability, loading a saved artefact |
The 88% target is chosen deliberately. A bag-of-words model with logistic regression reaches about 87% on this dataset in under a minute. If your recurrent model lands at 86%, it is not a slightly worse model — it is a much more expensive model that lost to a linear one, and something is wrong.
The dataset, and what is hiding in it
The IMDb review corpus holds 50,000 reviews, split 25,000 for training and 25,000 for testing, with balanced positive and negative labels. Reviews were selected to be strongly polarised: ratings of 1–4 count as negative, 7–10 as positive, and the ambiguous middle was excluded.
1import os23def load_imdb(split_dir):4 texts, labels = [], []5 for name, label in (("pos", 1), ("neg", 0)):6 folder = os.path.join(split_dir, name)7 for fn in sorted(os.listdir(folder)):8 with open(os.path.join(folder, fn), encoding="utf-8") as f:9 texts.append(f.read())10 labels.append(label)11 return texts, labels1213train_texts, train_labels = load_imdb("./aclImdb/train")14test_texts, test_labels = load_imdb("./aclImdb/test")Four properties of this data will affect your results, and three of them are easy to miss.
The text contains HTML. Reviews are full of <br /> tags. Leave them in and br becomes one of the most frequent tokens in your entire vocabulary, occupying an embedding row and contributing nothing.
The polarisation makes it easier than real feedback. Excluding 5- and 6-star reviews removes the hardest cases. A model at 90% here may sit at 75% on live customer feedback, which contains the lukewarm reviews this corpus threw away. Do not report the IMDb number as if it predicts production behaviour.
Lengths vary enormously. The mean is around 230 words; the longest exceed 2,400. This single fact drives your most consequential hyperparameter.
There are near-duplicates. Some reviewers post variations of the same text. If a near-duplicate straddles your train and validation split, your validation score is measuring memorisation.
Look at the length distribution before choosing anything
1import numpy as np23lengths = np.array([len(t.split()) for t in train_texts])4for p in (50, 75, 90, 95, 99):5 print(f"p{p}: {np.percentile(lengths, p):.0f} words")6print(f"mean {lengths.mean():.0f} max {lengths.max()}")p50: 174 wordsp75: 284 wordsp90: 458 wordsp95: 598 wordsp99: 913 wordsmean 234 max 2470Now the trade-off is visible:
max_len | Reviews kept whole | Relative compute | Comment |
|---|---|---|---|
| 100 | ~12% | 0.4× | Fast, but most reviews lose their ending — where verdicts live |
| 250 | ~70% | 1.0× | The sensible default |
| 500 | ~92% | 2.0× | Worth it if the extra accuracy shows up on validation |
| 1000 | ~99% | 4.0× | Four times the cost for the last 8% of reviews |
One detail matters more than the number itself: truncate from the front, not the back. Reviews build to a verdict. Keeping the first 250 words of a 600-word review discards the conclusion; keeping the last 250 keeps it. Try both and measure — on this dataset, tail truncation usually wins by a point or more.
Cleaning that helps and cleaning that hurts
1import re23def clean(text):4 text = re.sub(r"<br\s*/?>", " ", text) # the HTML that is actually there5 text = re.sub(r"<[^>]+>", " ", text) # anything else tag-shaped6 text = text.lower()7 text = re.sub(r"http\S+|www\.\S+", " url ", text)8 text = re.sub(r"([!?])\1+", r" \1\1 ", text) # "!!!!!" -> " !! " (keep intensity)9 text = re.sub(r"[^a-z0-9!?'\s]", " ", text) # keep ! ? and apostrophes10 text = re.sub(r"\s+", " ", text).strip()11 return text1213def tokenise(text):14 return clean(text).split()Two choices in that function are deliberate departures from the standard recipe, and both are worth defending.
Exclamation and question marks are kept. They carry sentiment intensity. Collapsing runs of them stops !!!, !!!! and !!!!! from becoming three separate vocabulary entries that each appear a handful of times.
Stopwords are not removed. The standard English stopword list contains not, no and never. Applying it turns "this was not good" into good — you have not cleaned the review, you have inverted it. If you want a smaller vocabulary, use a minimum frequency threshold instead, which achieves the same thing without deleting the negations.
Deduplicate before splitting
1seen, keep = set(), []2for i, t in enumerate(train_texts):3 h = hash(clean(t)[:400]) # first 400 chars of cleaned text4 if h not in seen:5 seen.add(h)6 keep.append(i)7print(f"removed {len(train_texts) - len(keep)} near-duplicates")Vocabulary and batching
1from collections import Counter2import torch3from torch.utils.data import Dataset, DataLoader4from torch.nn.utils.rnn import pad_sequence56PAD, UNK = 0, 178def build_vocab(tokenised_docs, max_size=25000, min_count=2):9 counts = Counter(w for doc in tokenised_docs for w in doc)10 vocab = {"<pad>": PAD, "<unk>": UNK}11 for word, c in counts.most_common():12 if c < min_count or len(vocab) >= max_size:13 break14 vocab[word] = len(vocab)15 return vocab1617class ReviewDataset(Dataset):18 def __init__(self, docs, labels, vocab, max_len=250):19 self.items = []20 for tokens, y in zip(docs, labels):21 ids = [vocab.get(t, UNK) for t in tokens][-max_len:] # tail22 self.items.append((torch.tensor(ids), len(ids), y))23 self.labels = labels2425 def __len__(self):26 return len(self.items)2728 def __getitem__(self, i):29 return self.items[i]3031def collate(batch):32 seqs, lengths, labels = zip(*batch)33 return (pad_sequence(seqs, batch_first=True, padding_value=PAD),34 torch.tensor(lengths),35 torch.tensor(labels, dtype=torch.float))Two rules govern this step, and breaking either inflates your reported score.
Build the vocabulary on the training split only. Not on train plus validation, and certainly not on the test set. A vocabulary derived from data the model is supposed to have never seen is a leak: the identity of rare words in your evaluation set has influenced what the model can represent.
Set min_count to at least 2. A word appearing once in 25,000 reviews receives exactly one gradient update during the whole training run. Its embedding stays essentially at its random initialisation and contributes noise at inference. On this dataset, dropping singletons removes over 40% of the vocabulary and costs nothing.
The model
1import torch.nn as nn23class SentimentModel(nn.Module):4 def __init__(self, vocab_size, embed_dim=100, hidden_dim=128,5 num_layers=2, dropout=0.4, pretrained=None):6 super().__init__()7 self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=PAD)8 if pretrained is not None:9 self.embedding.weight.data.copy_(torch.tensor(pretrained))10 self.embedding.weight.requires_grad = False # unfrozen at epoch 21112 self.lstm = nn.LSTM(embed_dim, hidden_dim, num_layers=num_layers,13 bidirectional=True, batch_first=True,14 dropout=dropout if num_layers > 1 else 0.0)15 self.attn = nn.Linear(hidden_dim * 2, 1)16 self.dropout = nn.Dropout(dropout)17 self.fc = nn.Linear(hidden_dim * 2, 1)1819 # Bias the forget gate towards remembering at initialisation20 for name, p in self.lstm.named_parameters():21 if "bias" in name:22 n = p.size(0)23 p.data[n // 4: n // 2].fill_(1.0)2425 def forward(self, x, lengths):26 mask = x != PAD27 emb = self.dropout(self.embedding(x))28 packed = nn.utils.rnn.pack_padded_sequence(29 emb, lengths.cpu(), batch_first=True, enforce_sorted=False)30 packed_out, _ = self.lstm(packed)31 out, _ = nn.utils.rnn.pad_packed_sequence(32 packed_out, batch_first=True, total_length=x.size(1))3334 scores = self.attn(out).squeeze(-1).masked_fill(~mask, float("-inf"))35 alpha = torch.softmax(scores, dim=1)36 context = torch.bmm(alpha.unsqueeze(1), out).squeeze(1)37 return self.fc(self.dropout(context)).squeeze(-1), alphaFour things in this class are load-bearing, and each corresponds to a common way this project goes wrong.
| Line | Why it is there | What happens without it |
|---|---|---|
padding_idx=PAD | Pad embedding stays zero, gets no gradient | The model learns a vector for "nothing" |
pack_padded_sequence | LSTM stops at each sequence's true end | A 20-word review is pushed through 230 steps of padding; its signal decays away |
masked_fill(~mask, -inf) | Pad positions get zero attention weight | Padding steals probability mass from real tokens |
| Forget-gate bias = 1 | Cell starts out inclined to remember | Slower convergence, weaker long-range behaviour |
Attention pooling rather than the final hidden state is the right call for this data. Reviews are long and the decisive clause can be anywhere; attention lets the model weight it wherever it sits. It also gives you a debugging tool, since you can print which tokens the model actually used.
Pretrained embeddings
1import numpy as np2import gensim.downloader as api34def build_embedding_matrix(vocab, dim=100):5 kv = api.load(f"glove-wiki-gigaword-{dim}")6 rng = np.random.default_rng(0)7 matrix = rng.normal(0, 0.1, (len(vocab), dim)).astype(np.float32)8 matrix[PAD] = 0.09 hits = 010 for word, idx in vocab.items():11 if word in kv:12 matrix[idx] = kv[word]13 hits += 114 print(f"initialised {hits}/{len(vocab)} rows from GloVe")15 return matrixOn 25,000 reviews, initialising from pretrained vectors is typically worth 2–3 accuracy points and gets you there in fewer epochs. Freeze the embedding layer for the first two epochs while the layers above it stabilise, then unfreeze it with a lower learning rate.
Training
1criterion = nn.BCEWithLogitsLoss()2optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)3scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(4 optimizer, mode="max", factor=0.5, patience=1)56def run(loader, train):7 model.train() if train else model.eval()8 total, correct, n = 0.0, 0, 09 with torch.set_grad_enabled(train):10 for x, lengths, y in loader:11 x, y = x.to(device), y.to(device)12 logits, _ = model(x, lengths)13 loss = criterion(logits, y)14 if train:15 optimizer.zero_grad()16 loss.backward()17 nn.utils.clip_grad_norm_(model.parameters(), 5.0)18 optimizer.step()19 total += loss.item() * y.size(0)20 correct += ((torch.sigmoid(logits) > 0.5) == y.bool()).sum().item()21 n += y.size(0)22 return total / n, correct / n2324best, patience, bad_epochs = 0.0, 3, 025for epoch in range(15):26 if epoch == 2:27 model.embedding.weight.requires_grad = True # unfreeze28 tr_loss, tr_acc = run(train_loader, True)29 va_loss, va_acc = run(val_loader, False)30 scheduler.step(va_acc)31 print(f"{epoch:2d} train {tr_acc:.4f} val {va_acc:.4f} loss {va_loss:.4f}")32 if va_acc > best:33 best, bad_epochs = va_acc, 034 torch.save(model.state_dict(), "best.pt")35 else:36 bad_epochs += 137 if bad_epochs >= patience:38 print("early stopping")39 breakWhat a healthy run looks like on this data, and what each pathology means:
| Pattern | Diagnosis | Action |
|---|---|---|
| Train 0.93 / val 0.89, both rising | Healthy | Keep going |
| Train 0.99 / val 0.86, val falling | Overfitting | More dropout, freeze embeddings longer, stop earlier |
| Both stuck near 0.50 | Not learning at all | Check labels are in the batch correctly; check learning rate; check you are not applying sigmoid twice |
Loss becomes NaN | Exploding gradients | Clipping is present but the learning rate is too high |
| Val accuracy 0.87 and flat from epoch 1 | The model is not using sequence structure | Verify packing; verify attention mask; compare against bag-of-words |
That last row is the one worth checking early. If your BiLSTM matches a linear bag-of-words model exactly, the recurrence is probably not contributing — usually because padding is being processed as content.
The transformer comparison
Run a fine-tuned pretrained transformer on the identical split so the comparison is meaningful.
1from transformers import AutoTokenizer, AutoModelForSequenceClassification2from transformers import get_linear_schedule_with_warmup3from torch.optim import AdamW45name = "distilbert-base-uncased"6tok = AutoTokenizer.from_pretrained(name)7bert = AutoModelForSequenceClassification.from_pretrained(name, num_labels=2).to(device)89enc = tok(raw_texts, truncation=True, max_length=256, padding="max_length",10 return_tensors="pt")1112opt = AdamW(bert.parameters(), lr=2e-5, weight_decay=0.01)13steps = len(bert_loader) * 314sched = get_linear_schedule_with_warmup(opt, int(0.1 * steps), steps)Note the learning rate: 2e-5, not 1e-3. Fine-tuning at 1e-3 destroys the pretrained weights within a few hundred steps and lands you near chance. This is the single most common mistake in the comparison step, and the symptom — a large pretrained model performing worse than a small LSTM — looks like an architecture result rather than the configuration error it is.
Also note: feed the transformer raw text, not your cleaned tokens. Its own subword tokeniser was trained against a specific vocabulary, and pre-cleaning shifts every index away from what the model saw during pretraining.
| Model | Accuracy | Params | Train time (GPU) | CPU latency |
|---|---|---|---|---|
| TF-IDF + logistic regression | ~87% | ~0.3 M | < 1 min | < 1 ms |
| BiLSTM + attention, random embeddings | ~86% | ~3 M | ~10 min | ~5 ms |
| BiLSTM + attention, GloVe | ~89% | ~3 M | ~10 min | ~5 ms |
| DistilBERT fine-tuned | ~92% | 66 M | ~20 min | ~60 ms |
| BERT-base fine-tuned | ~93% | 110 M | ~40 min | ~110 ms |
Three to four points of accuracy for twelve to twenty times the latency and twenty to forty times the parameters. Whether that is a good trade depends entirely on what the system is for — which is exactly the judgement this project exists to make you exercise.
Augmentation, measured honestly
1import random2from nltk.corpus import wordnet34def synonym_replace(tokens, n=2, protected=frozenset({"not", "no", "never", "n't"})):5 tokens = list(tokens)6 candidates = [i for i, t in enumerate(tokens)7 if len(t) > 3 and t not in protected and wordnet.synsets(t)]8 random.shuffle(candidates)9 for i in candidates[:n]:10 syns = {l.name().replace("_", " ")11 for s in wordnet.synsets(tokens[i]) for l in s.lemmas()}12 syns.discard(tokens[i])13 if syns:14 tokens[i] = random.choice(sorted(syns))15 return tokensThe protected set is essential. Replacing not with a WordNet "synonym" can flip a review's meaning while its label stays fixed, and you have manufactured mislabelled training data.
| Technique | Cost | Typical gain on 25k examples | Risk |
|---|---|---|---|
| Synonym replacement | Cheap | +0.2 to +0.5 points | Can alter sentiment |
| Random deletion (10%) | Free | +0 to +0.3 points | May delete the negation |
| Back-translation | Expensive | +0.5 to +1.5 points | Needs translation models |
| More real labelled data | Expensive | Reliably the largest | None |
Be sceptical here. On 25,000 balanced examples augmentation buys very little, because the dataset is already large enough for a 3M-parameter model. Its value rises sharply when you have a few thousand examples or a badly imbalanced class. Measure it on your split rather than assuming it helps — and always augment after splitting, or an augmented copy of a training review will appear in validation.
Packaging and serving
This is where projects that scored well quietly break, and the cause is almost always the same: the model was saved and the vocabulary was not.
The weights are meaningless without the exact token-to-index mapping they were trained with. Index 4,821 means brilliant only if the same vocabulary is loaded. Rebuild the vocabulary at deploy time from slightly different data and every index shifts. There is no error, no warning, and no crash — the model simply reads a different sentence from the one you sent it and predicts near-randomly.
1import json23torch.save({4 "state_dict": model.state_dict(),5 "vocab": vocab,6 "config": {"embed_dim": 100, "hidden_dim": 128,7 "num_layers": 2, "max_len": 250, "threshold": 0.5},8}, "sentiment_model.pt")Weights, vocabulary, architecture configuration and threshold travel together as one artefact. Anything less is not deployable.
1from fastapi import FastAPI2from pydantic import BaseModel34app = FastAPI()5ckpt = torch.load("sentiment_model.pt", map_location="cpu")6vocab, cfg = ckpt["vocab"], ckpt["config"]78model = SentimentModel(len(vocab), cfg["embed_dim"], cfg["hidden_dim"],9 cfg["num_layers"])10model.load_state_dict(ckpt["state_dict"])11model.eval()1213class Review(BaseModel):14 text: str1516@app.post("/predict")17def predict(review: Review):18 tokens = tokenise(review.text) # the same function as training19 ids = [vocab.get(t, UNK) for t in tokens][-cfg["max_len"]:] or [UNK]20 x = torch.tensor([ids])21 with torch.no_grad():22 logits, alpha = model(x, torch.tensor([len(ids)]))23 prob = torch.sigmoid(logits).item()24 unknown_rate = sum(i == UNK for i in ids) / len(ids)25 return {26 "label": "positive" if prob > cfg["threshold"] else "negative",27 "confidence": round(prob if prob > 0.5 else 1 - prob, 4),28 "probability_positive": round(prob, 4),29 "unknown_token_rate": round(unknown_rate, 3),30 }Four details in that handler are doing real work. model.eval() disables dropout — forget it and every request gets a slightly different answer. torch.no_grad() avoids building a graph you will never backpropagate through, roughly halving latency and memory. or [UNK] handles empty input, which would otherwise pass a zero-length tensor to the LSTM and raise. And returning the probability rather than only a label lets whatever calls this route borderline cases to a human.
The unknown_token_rate is the cheapest monitoring you will ever add. If it is 3% in validation and 22% in production, your traffic no longer resembles your training data — and you will know that weeks before anyone notices the accuracy has slipped.
Reading the errors, and knowing when it is finished
Once the number clears 88%, the highest-value work is not tuning. It is printing your fifty most confident mistakes and sorting them by hand.
1model.eval()2errors = []3with torch.no_grad():4 for x, lengths, y in val_loader:5 logits, alpha = model(x.to(device), lengths)6 probs = torch.sigmoid(logits).cpu()7 for i in range(len(y)):8 if (probs[i] > 0.5) != bool(y[i]):9 top = alpha[i].topk(5).indices.tolist()10 errors.append({11 "prob": float(probs[i]),12 "true": int(y[i]),13 "attended": [inv_vocab[int(x[i][j])] for j in top],14 })15errors.sort(key=lambda e: abs(e["prob"] - e["true"]), reverse=True)The attention weights turn a vague "it got this wrong" into a specific diagnosis. If the model calls "not good at all" positive and its attention sits on good with nothing on not, you know precisely what is broken, and adding hidden units will not fix it.
Expect the residual errors to sort into a small number of families: sarcasm, mixed reviews where the label is one-sided, plot summaries that quote negative events in a positive review, and reviews where the gold label is simply wrong. Only the first two are model problems; the last is a data problem, and it is more common than people expect.
Finally, keep a handful of hand-written cases and run them on every retrain — a negation, a but-contrast, a very short review, a very long one, and something with unusual vocabulary. Aggregate accuracy moves slowly and hides regressions. A model that starts calling "not good" positive has broken in a way that a 0.3% dip in validation accuracy will never tell you about, and those five examples take two seconds to check.