Course Content
Natural Language Processing Basics
4 sections · 10 lessons
Sentiment Analysis with RNNs
Your first sentiment classifier writes itself in about eight lines. Two word lists, count them, take the difference.
1POSITIVE = {"good", "great", "excellent", "love", "best", "amazing"}2NEGATIVE = {"bad", "terrible", "awful", "hate", "worst", "boring"}34def sentiment(text):5 words = text.lower().split()6 score = sum(w in POSITIVE for w in words) - sum(w in NEGATIVE for w in words)7 return "positive" if score > 0 else "negative"Run it on real reviews and watch it fall over.
| Review | Lexicon says | Truth |
|---|---|---|
| "this film was not good at all" | positive | negative |
| "I expected it to be terrible, but it was wonderful" | 0 → negative | positive |
| "the acting was fine, the plot was fine, everything was fine" | 0 → negative | lukewarm negative |
| "a masterclass in how to waste two hours" | 0 → negative | negative — by luck, not reasoning |
| "the special effects were bad but I loved every minute" | negative | positive |
Two of five right, and both by accident. The failures are not random — they cluster around a single deficiency. The lexicon sees words. It never sees how words modify each other. not before good, but after a concession, expected setting up a contrast: all of these change what the surrounding words mean, and all of them are invisible to something that counts.
Sentiment analysis is worth studying precisely because it looks trivial and is not. It is the standard benchmark for whether a model can handle composition rather than vocabulary.
The task, stated precisely
Sentiment analysis is text classification where the label is an opinion. Three variants come up, and they are genuinely different problems:
| Variant | Input | Output | Difficulty |
|---|---|---|---|
| Document-level | A whole review | One label | Easiest — plenty of redundant signal |
| Sentence-level | One sentence | One label per sentence | Harder — less context to work with |
| Aspect-based | A review plus an aspect | One label per aspect | Hardest — one text, several opposing verdicts |
Aspect-based is the one businesses actually want. "Great food, appalling service, and the bill took twenty minutes to arrive" is not positive or negative — it is positive on food and negative on service and speed. A document-level model forced to emit one label for this loses the information the restaurant needs.
Six ways sentiment defeats naive models
Negation
Negation inverts polarity, and its scope is not fixed:
"not good" -> negation reaches 1 word"not particularly well made" -> reaches 3 words"I would not say this was good" -> reaches 5 words, across a clause"not without merit" -> double negation, back to positiveA bag-of-words model can catch not good as a bigram. It cannot catch not ... good with five words between them. A recurrent model can, because its state carries the negation forward.
Contrastive conjunctions
The word but is a polarity switch that tells you which half of a sentence to weight:
"the plot was weak but the performances were extraordinary" -> positive"the performances were extraordinary but the plot was weak" -> negativeIdentical words, identical counts, opposite labels. Order alone decides it. This pair is the cleanest possible demonstration of why sequence models exist for this task.
Context-dependent vocabulary
Sentiment words are not universally positive or negative:
| Word | Positive context | Negative context |
|---|---|---|
unpredictable | an unpredictable plot | unpredictable steering |
quiet | a quiet engine | quiet dialogue in a cinema |
cheap | cheap flights | cheap materials |
long | long battery life | long queues |
simple | simple to use | simple plot |
This is why sentiment models transfer badly across domains. A model trained on film reviews learns that predictable is negative; deployed on train timetable feedback it will get that exactly backwards.
Intensifiers and hedges
"good" baseline"very good" stronger"absolutely superb" stronger still"fairly good" weaker"good, I suppose" weaker, verging on negativeThese modifiers carry no sentiment on their own and change the magnitude of whatever follows. A counting model has no mechanism for magnitude at all.
Sarcasm and irony
"Brilliant. Another two hours I will never get back.""Oh good, it broke again.""Ten out of ten for the packaging. Shame about the product."The surface words are positive; the meaning is not. Sarcasm depends on tone, shared expectations and world knowledge that is not in the text. Every model gets a meaningful share of these wrong, and it is honest to accept that as a floor rather than a bug to be fixed. Sarcasm is a large slice of the residual error on any consumer-review dataset.
Comparatives and conditionals
"better than their last one" -> positive about this, negative about that"if only the ending had worked" -> negative, stated as a wish"I wanted to like it" -> negative, stated as regretNone of these contain an explicit negative word. All are negative.
Almost every hard case in sentiment analysis is a case where meaning comes from arrangement rather than vocabulary. That is exactly the gap a recurrent model is built to close — and exactly why sentiment is the standard demonstration that word order matters.
Building the classifier
Vocabulary and encoding
The model needs integer indices, so you need a mapping from word to index built from the training data only.
1from collections import Counter23PAD, UNK = 0, 145def build_vocab(tokenised_docs, max_size=25000, min_count=2):6 counts = Counter(w for doc in tokenised_docs for w in doc)7 vocab = {"<pad>": PAD, "<unk>": UNK}8 for word, c in counts.most_common(max_size - 2):9 if c < min_count:10 break11 vocab[word] = len(vocab)12 return vocab1314def encode(tokens, vocab, max_len=300):15 ids = [vocab.get(t, UNK) for t in tokens][:max_len]16 return ids, len(ids)Build this on the training split alone. Building it on all your data before splitting is a leak: the vocabulary — and later the embedding table — is shaped by text the model is supposed to have never seen.
min_count=2 matters more than it looks. Words appearing once are typos, names and one-off constructions. Each gets an embedding row that receives exactly one gradient update in the whole run, so it stays near its random initialisation and contributes pure noise at inference.
Batching variable lengths
1import torch2from torch.utils.data import Dataset, DataLoader3from torch.nn.utils.rnn import pad_sequence45class ReviewDataset(Dataset):6 def __init__(self, docs, labels, vocab, max_len=300):7 self.data = [encode(d, vocab, max_len) for d in docs]8 self.labels = labels910 def __len__(self):11 return len(self.labels)1213 def __getitem__(self, i):14 ids, length = self.data[i]15 return torch.tensor(ids), length, torch.tensor(self.labels[i])1617def collate(batch):18 seqs, lengths, labels = zip(*batch)19 padded = pad_sequence(seqs, batch_first=True, padding_value=PAD)20 return padded, torch.tensor(lengths), torch.stack(labels)2122loader = DataLoader(train_ds, batch_size=64, shuffle=True, collate_fn=collate)Padding is a batching convenience, not data. The model must be told to ignore it, which is what the lengths tensor is for.
The model
1import torch.nn as nn23class SentimentBiLSTM(nn.Module):4 def __init__(self, vocab_size, embed_dim=100, hidden_dim=128,5 num_layers=2, dropout=0.4, pretrained=None):6 super().__init__()7 self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=PAD)8 if pretrained is not None:9 self.embedding.weight.data.copy_(torch.tensor(pretrained))10 self.embedding.weight.requires_grad = False # unfreeze later1112 self.lstm = nn.LSTM(13 embed_dim, hidden_dim, num_layers=num_layers,14 bidirectional=True, batch_first=True,15 dropout=dropout if num_layers > 1 else 0.0,16 )17 self.attn = nn.Linear(hidden_dim * 2, 1)18 self.dropout = nn.Dropout(dropout)19 self.fc = nn.Linear(hidden_dim * 2, 1) # one logit for binary2021 def forward(self, x, lengths):22 mask = x != PAD23 emb = self.dropout(self.embedding(x))24 packed = nn.utils.rnn.pack_padded_sequence(25 emb, lengths.cpu(), batch_first=True, enforce_sorted=False26 )27 packed_out, _ = self.lstm(packed)28 out, _ = nn.utils.rnn.pad_packed_sequence(29 packed_out, batch_first=True, total_length=x.size(1)30 )31 scores = self.attn(out).squeeze(-1).masked_fill(~mask, float("-inf"))32 alpha = torch.softmax(scores, dim=1)33 context = torch.bmm(alpha.unsqueeze(1), out).squeeze(1)34 return self.fc(self.dropout(context)).squeeze(-1), alphaAttention pooling rather than the last hidden state is a deliberate choice here. Reviews are long, and the decisive clause can be anywhere. Pooling by attention lets the model weight the clause that carries the verdict instead of whatever happened to come last.
Training
1model = SentimentBiLSTM(len(vocab)).to(device)2criterion = nn.BCEWithLogitsLoss()3optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)45def run_epoch(loader, train=True):6 model.train() if train else model.eval()7 total_loss, correct, n = 0.0, 0, 08 with torch.set_grad_enabled(train):9 for x, lengths, y in loader:10 x, y = x.to(device), y.to(device).float()11 logits, _ = model(x, lengths)12 loss = criterion(logits, y)13 if train:14 optimizer.zero_grad()15 loss.backward()16 nn.utils.clip_grad_norm_(model.parameters(), 5.0)17 optimizer.step()18 total_loss += loss.item() * y.size(0)19 correct += ((torch.sigmoid(logits) > 0.5) == y.bool()).sum().item()20 n += y.size(0)21 return total_loss / n, correct / n2223best = 0.024for epoch in range(10):25 tr_loss, tr_acc = run_epoch(train_loader, True)26 va_loss, va_acc = run_epoch(val_loader, False)27 print(f"epoch {epoch}: train {tr_acc:.3f} | val {va_acc:.3f}")28 if va_acc > best:29 best = va_acc30 torch.save(model.state_dict(), "best.pt")BCEWithLogitsLoss takes raw logits and applies the sigmoid internally in a numerically stable way. Applying sigmoid yourself and then using BCELoss is a common source of NaN losses.
Two habits that save time. Save on the best validation score, not the last epoch — recurrent models overfit and epoch 10 is often worse than epoch 4. And if the embedding layer is frozen, unfreeze it after two or three epochs with a reduced learning rate, once the layers above it have stopped producing wild gradients.
Accuracy is the wrong number
Suppose you deploy a model to flag negative reviews for the support team. In production, 5% of reviews are negative. Your model reports 95% accuracy. Here is the confusion matrix from 10,000 reviews:
| Predicted negative | Predicted positive | |
|---|---|---|
| Actually negative (500) | 0 | 500 |
| Actually positive (9,500) | 0 | 9,500 |
The model predicts "positive" for every input. It is 95% accurate and completely worthless — it finds none of the reviews it exists to find. Accuracy on imbalanced data measures the class balance, not the model.
Use precision, recall and F1 on the class you care about:
A real model on that data might give:
| Predicted negative | Predicted positive | |
|---|---|---|
| Actually negative (500) | 390 (TP) | 110 (FN) |
| Actually positive (9,500) | 260 (FP) | 9,240 (TN) |
Precision = 390 / 650 = 0.60. Recall = 390 / 500 = 0.78. F1 = 2(0.60 × 0.78) / 1.38 = 0.68. Accuracy = 9,630 / 10,000 = 96.3% — barely above the useless model, while the F1 tells you something real happened.
The threshold is a product decision
The 0.5 cut-off is a default, not a law. Moving it trades precision against recall:
| Threshold | Precision | Recall | F1 | Suits |
|---|---|---|---|---|
| 0.30 | 0.44 | 0.91 | 0.59 | Catch nearly every complaint; humans filter |
| 0.50 | 0.60 | 0.78 | 0.68 | Balanced default |
| 0.70 | 0.79 | 0.61 | 0.69 | Auto-escalation where false alarms are costly |
| 0.90 | 0.93 | 0.31 | 0.46 | Fully automated refunds |
Pick the threshold on the validation set against the cost of each error type in your application, not by leaving it where it was initialised.
Making negation explicit
If your model struggles with negation, one cheap preprocessing trick helps a great deal: mark every token between a negation word and the next punctuation mark.
1import re23NEGATIONS = {"not", "no", "never", "n't", "cannot", "without", "hardly"}4CLAUSE_END = {".", ",", ";", "!", "?", "but", "however", "although"}56def mark_negation(tokens):7 out, negating = [], False8 for t in tokens:9 if t in CLAUSE_END:10 negating = False11 out.append(t)12 continue13 out.append("NEG_" + t if negating else t)14 if t in NEGATIONS:15 negating = True16 return out1718print(mark_negation("this was not good but the ending was great".split()))19# ['this', 'was', 'not', 'NEG_good', 'but', 'the', 'ending', 'was', 'great']NEG_good is now a distinct vocabulary entry that can learn its own, negative embedding, entirely separate from good. The clause boundary stops the marking from running away — without it, "not good, but the ending was great" would mark great as negated too, which is the opposite of what you want.
This is most valuable for bag-of-words and linear models, where negation is otherwise invisible. A well-trained BiLSTM learns much of this from data, but the explicit marker still tends to help on smaller datasets where the model has less evidence to learn from.
More than two classes
Star ratings are ordinal — 1 to 5, ordered — and this changes the right loss.
| Framing | Output layer | Loss | Treats 1★ vs 5★ error as |
|---|---|---|---|
| Multi-class | 5 logits | Cross-entropy | Same as 4★ vs 5★ — wrong |
| Regression | 1 value | MSE | 16× worse than 4★ vs 5★ |
| Ordinal (cumulative) | 4 binary outputs | BCE on "is rating > k?" | Correctly ordered, respects discreteness |
Plain cross-entropy over five classes is the default and it is subtly wrong: it treats the five labels as unrelated categories, so predicting 1★ for a 5★ review costs exactly what predicting 4★ costs. Regression fixes the ordering but produces outputs like 3.7 that you then have to round, and it assumes the gaps between stars are equal. The cumulative-link formulation — four binary classifiers answering "is this above 1?", "above 2?" and so on — usually performs best and is barely more code.
Reading your model's mistakes
The highest-value hour you can spend on a sentiment model is not tuning it. It is printing fifty misclassified examples and sorting them into categories by hand.
1model.eval()2errors = []3with torch.no_grad():4 for x, lengths, y in val_loader:5 logits, alpha = model(x.to(device), lengths)6 probs = torch.sigmoid(logits).cpu()7 for i in range(len(y)):8 if (probs[i] > 0.5) != bool(y[i]):9 errors.append({10 "text": decode(x[i], inv_vocab),11 "true": int(y[i]),12 "prob": float(probs[i]),13 "top_tokens": top_attended(x[i], alpha[i], inv_vocab, k=5),14 })1516errors.sort(key=lambda e: abs(e["prob"] - e["true"]), reverse=True)Sorting by confidence puts the model's most confident mistakes first, and those are the informative ones. Then categorise:
| Error category | What it points at | What to do |
|---|---|---|
| Negation missed | Model is not composing | Negation marking; more layers; check not is not a stopword |
| Sarcasm | Genuinely hard | Accept a floor; do not chase it with capacity |
| Mixed sentiment forced to one label | Task framing | Move to aspect-based, or add a "mixed" class |
| Long review, verdict early | Truncation or pooling | Raise max_len; attention pooling |
| Domain vocabulary | Train/production mismatch | Fine-tune on in-domain data |
| The label is simply wrong | Data quality | Fix the labels — this is more common than people expect |
Attention weights make this diagnosis far faster. If a model calls "not good at all" positive and its attention mass sits on good with almost nothing on not, you know precisely what is broken, and no amount of adding hidden units will fix it.
Taking it to production
Three things separate a model that scores well in a notebook from one that works in a live system.
Ship preprocessing and vocabulary with the weights. The model is a function of the exact token-to-index mapping it was trained with. Save the vocabulary, the tokeniser configuration, and the maximum length in the same artefact as the weights. A mismatch here produces no error — indices simply point at the wrong embeddings, and accuracy silently collapses.
Return the probability, not just the label. Downstream systems need to distinguish 0.51 from 0.99. Routing a borderline case to a human is only possible if the confidence survives the API boundary.
Watch for drift. Language moves. New products, new slang, new complaint patterns. Track the proportion of unknown tokens in live traffic and the distribution of predicted probabilities. A rising unknown-token rate or predictions clustering near 0.5 both mean the model is seeing text unlike its training data — and both show up long before anyone notices the accuracy has dropped.
Finally, keep a fixed set of hand-written adversarial cases and run them on every retrain: a negation, a but-contrast, a mixed review, a sarcastic one. Aggregate metrics move slowly and hide regressions. A model that starts calling "not good" positive has broken in a way that a 0.3% accuracy dip will never tell you about.