Natural Language Processing Basics

Sentiment Analysis with RNNs


Your first sentiment classifier writes itself in about eight lines. Two word lists, count them, take the difference.

Python
POSITIVE = {"good", "great", "excellent", "love", "best", "amazing"}NEGATIVE = {"bad", "terrible", "awful", "hate", "worst", "boring"}def sentiment(text):    words = text.lower().split()    score = sum(w in POSITIVE for w in words) - sum(w in NEGATIVE for w in words)    return "positive" if score > 0 else "negative"

Run it on real reviews and watch it fall over.

ReviewLexicon saysTruth
"this film was not good at all"positivenegative
"I expected it to be terrible, but it was wonderful"0 → negativepositive
"the acting was fine, the plot was fine, everything was fine"0 → negativelukewarm negative
"a masterclass in how to waste two hours"0 → negativenegative — by luck, not reasoning
"the special effects were bad but I loved every minute"negativepositive

Two of five right, and both by accident. The failures are not random — they cluster around a single deficiency. The lexicon sees words. It never sees how words modify each other. not before good, but after a concession, expected setting up a contrast: all of these change what the surrounding words mean, and all of them are invisible to something that counts.

Sentiment analysis is worth studying precisely because it looks trivial and is not. It is the standard benchmark for whether a model can handle composition rather than vocabulary.

Six sentences a word-count classifier gets backwardsWord order carries the sign• "not at all good" — negation flips it• "cheap food, but awful service"• "barely acceptable" versus "truly great"The word is not the meaning• "unpredictable plot" or "brakes"• "Brilliant. Waited ninety minutes."• "better than their last attempt"
Every one of these is fixed by reading the sentence in order, which is precisely what a recurrent model does.

The task, stated precisely

Sentiment analysis is text classification where the label is an opinion. Three variants come up, and they are genuinely different problems:

VariantInputOutputDifficulty
Document-levelA whole reviewOne labelEasiest — plenty of redundant signal
Sentence-levelOne sentenceOne label per sentenceHarder — less context to work with
Aspect-basedA review plus an aspectOne label per aspectHardest — one text, several opposing verdicts

Aspect-based is the one businesses actually want. "Great food, appalling service, and the bill took twenty minutes to arrive" is not positive or negative — it is positive on food and negative on service and speed. A document-level model forced to emit one label for this loses the information the restaurant needs.

Six ways sentiment defeats naive models

Negation

Negation inverts polarity, and its scope is not fixed:

Text
"not good"                        -> negation reaches 1 word"not particularly well made"      -> reaches 3 words"I would not say this was good"   -> reaches 5 words, across a clause"not without merit"               -> double negation, back to positive

A bag-of-words model can catch not good as a bigram. It cannot catch not ... good with five words between them. A recurrent model can, because its state carries the negation forward.

Contrastive conjunctions

The word but is a polarity switch that tells you which half of a sentence to weight:

Text
"the plot was weak but the performances were extraordinary"  -> positive"the performances were extraordinary but the plot was weak"  -> negative

Identical words, identical counts, opposite labels. Order alone decides it. This pair is the cleanest possible demonstration of why sequence models exist for this task.

Context-dependent vocabulary

Sentiment words are not universally positive or negative:

WordPositive contextNegative context
unpredictablean unpredictable plotunpredictable steering
quieta quiet enginequiet dialogue in a cinema
cheapcheap flightscheap materials
longlong battery lifelong queues
simplesimple to usesimple plot

This is why sentiment models transfer badly across domains. A model trained on film reviews learns that predictable is negative; deployed on train timetable feedback it will get that exactly backwards.

Intensifiers and hedges

Text
"good"              baseline"very good"         stronger"absolutely superb" stronger still"fairly good"       weaker"good, I suppose"   weaker, verging on negative

These modifiers carry no sentiment on their own and change the magnitude of whatever follows. A counting model has no mechanism for magnitude at all.

Sarcasm and irony

Text
"Brilliant. Another two hours I will never get back.""Oh good, it broke again.""Ten out of ten for the packaging. Shame about the product."

The surface words are positive; the meaning is not. Sarcasm depends on tone, shared expectations and world knowledge that is not in the text. Every model gets a meaningful share of these wrong, and it is honest to accept that as a floor rather than a bug to be fixed. Sarcasm is a large slice of the residual error on any consumer-review dataset.

Comparatives and conditionals

Text
"better than their last one"      -> positive about this, negative about that"if only the ending had worked"   -> negative, stated as a wish"I wanted to like it"             -> negative, stated as regret

None of these contain an explicit negative word. All are negative.

Almost every hard case in sentiment analysis is a case where meaning comes from arrangement rather than vocabulary. That is exactly the gap a recurrent model is built to close — and exactly why sentiment is the standard demonstration that word order matters.

Building the classifier

Vocabulary and encoding

The model needs integer indices, so you need a mapping from word to index built from the training data only.

Python
from collections import CounterPAD, UNK = 0, 1def build_vocab(tokenised_docs, max_size=25000, min_count=2):    counts = Counter(w for doc in tokenised_docs for w in doc)    vocab = {"<pad>": PAD, "<unk>": UNK}    for word, c in counts.most_common(max_size - 2):        if c < min_count:            break        vocab[word] = len(vocab)    return vocabdef encode(tokens, vocab, max_len=300):    ids = [vocab.get(t, UNK) for t in tokens][:max_len]    return ids, len(ids)

Build this on the training split alone. Building it on all your data before splitting is a leak: the vocabulary — and later the embedding table — is shaped by text the model is supposed to have never seen.

min_count=2 matters more than it looks. Words appearing once are typos, names and one-off constructions. Each gets an embedding row that receives exactly one gradient update in the whole run, so it stays near its random initialisation and contributes pure noise at inference.

Batching variable lengths

Python
import torchfrom torch.utils.data import Dataset, DataLoaderfrom torch.nn.utils.rnn import pad_sequenceclass ReviewDataset(Dataset):    def __init__(self, docs, labels, vocab, max_len=300):        self.data = [encode(d, vocab, max_len) for d in docs]        self.labels = labels    def __len__(self):        return len(self.labels)    def __getitem__(self, i):        ids, length = self.data[i]        return torch.tensor(ids), length, torch.tensor(self.labels[i])def collate(batch):    seqs, lengths, labels = zip(*batch)    padded = pad_sequence(seqs, batch_first=True, padding_value=PAD)    return padded, torch.tensor(lengths), torch.stack(labels)loader = DataLoader(train_ds, batch_size=64, shuffle=True, collate_fn=collate)

Padding is a batching convenience, not data. The model must be told to ignore it, which is what the lengths tensor is for.

The model

Python
import torch.nn as nnclass SentimentBiLSTM(nn.Module):    def __init__(self, vocab_size, embed_dim=100, hidden_dim=128,                 num_layers=2, dropout=0.4, pretrained=None):        super().__init__()        self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=PAD)        if pretrained is not None:            self.embedding.weight.data.copy_(torch.tensor(pretrained))            self.embedding.weight.requires_grad = False   # unfreeze later        self.lstm = nn.LSTM(            embed_dim, hidden_dim, num_layers=num_layers,            bidirectional=True, batch_first=True,            dropout=dropout if num_layers > 1 else 0.0,        )        self.attn = nn.Linear(hidden_dim * 2, 1)        self.dropout = nn.Dropout(dropout)        self.fc = nn.Linear(hidden_dim * 2, 1)   # one logit for binary    def forward(self, x, lengths):        mask = x != PAD        emb = self.dropout(self.embedding(x))        packed = nn.utils.rnn.pack_padded_sequence(            emb, lengths.cpu(), batch_first=True, enforce_sorted=False        )        packed_out, _ = self.lstm(packed)        out, _ = nn.utils.rnn.pad_packed_sequence(            packed_out, batch_first=True, total_length=x.size(1)        )        scores = self.attn(out).squeeze(-1).masked_fill(~mask, float("-inf"))        alpha = torch.softmax(scores, dim=1)        context = torch.bmm(alpha.unsqueeze(1), out).squeeze(1)        return self.fc(self.dropout(context)).squeeze(-1), alpha

Attention pooling rather than the last hidden state is a deliberate choice here. Reviews are long, and the decisive clause can be anywhere. Pooling by attention lets the model weight the clause that carries the verdict instead of whatever happened to come last.

Training

Python
model = SentimentBiLSTM(len(vocab)).to(device)criterion = nn.BCEWithLogitsLoss()optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)def run_epoch(loader, train=True):    model.train() if train else model.eval()    total_loss, correct, n = 0.0, 0, 0    with torch.set_grad_enabled(train):        for x, lengths, y in loader:            x, y = x.to(device), y.to(device).float()            logits, _ = model(x, lengths)            loss = criterion(logits, y)            if train:                optimizer.zero_grad()                loss.backward()                nn.utils.clip_grad_norm_(model.parameters(), 5.0)                optimizer.step()            total_loss += loss.item() * y.size(0)            correct += ((torch.sigmoid(logits) > 0.5) == y.bool()).sum().item()            n += y.size(0)    return total_loss / n, correct / nbest = 0.0for epoch in range(10):    tr_loss, tr_acc = run_epoch(train_loader, True)    va_loss, va_acc = run_epoch(val_loader, False)    print(f"epoch {epoch}: train {tr_acc:.3f} | val {va_acc:.3f}")    if va_acc > best:        best = va_acc        torch.save(model.state_dict(), "best.pt")

BCEWithLogitsLoss takes raw logits and applies the sigmoid internally in a numerically stable way. Applying sigmoid yourself and then using BCELoss is a common source of NaN losses.

Two habits that save time. Save on the best validation score, not the last epoch — recurrent models overfit and epoch 10 is often worse than epoch 4. And if the embedding layer is frozen, unfreeze it after two or three epochs with a reduced learning rate, once the layers above it have stopped producing wild gradients.

Accuracy is the wrong number

Suppose you deploy a model to flag negative reviews for the support team. In production, 5% of reviews are negative. Your model reports 95% accuracy. Here is the confusion matrix from 10,000 reviews:

Predicted negativePredicted positive
Actually negative (500)0500
Actually positive (9,500)09,500

The model predicts "positive" for every input. It is 95% accurate and completely worthless — it finds none of the reviews it exists to find. Accuracy on imbalanced data measures the class balance, not the model.

Use precision, recall and F1 on the class you care about:

precision=TPTP+FP,recall=TPTP+FN,F1=2⋅P⋅RP+R\text{precision} = \frac{TP}{TP + FP}, \qquad \text{recall} = \frac{TP}{TP + FN}, \qquad F_1 = \frac{2 \cdot P \cdot R}{P + R}

A real model on that data might give:

Predicted negativePredicted positive
Actually negative (500)390 (TP)110 (FN)
Actually positive (9,500)260 (FP)9,240 (TN)

Precision = 390 / 650 = 0.60. Recall = 390 / 500 = 0.78. F1 = 2(0.60 × 0.78) / 1.38 = 0.68. Accuracy = 9,630 / 10,000 = 96.3% — barely above the useless model, while the F1 tells you something real happened.

The threshold is a product decision

The 0.5 cut-off is a default, not a law. Moving it trades precision against recall:

ThresholdPrecisionRecallF1Suits
0.300.440.910.59Catch nearly every complaint; humans filter
0.500.600.780.68Balanced default
0.700.790.610.69Auto-escalation where false alarms are costly
0.900.930.310.46Fully automated refunds

Pick the threshold on the validation set against the cost of each error type in your application, not by leaving it where it was initialised.

Making negation explicit

If your model struggles with negation, one cheap preprocessing trick helps a great deal: mark every token between a negation word and the next punctuation mark.

Python
import reNEGATIONS = {"not", "no", "never", "n't", "cannot", "without", "hardly"}CLAUSE_END = {".", ",", ";", "!", "?", "but", "however", "although"}def mark_negation(tokens):    out, negating = [], False    for t in tokens:        if t in CLAUSE_END:            negating = False            out.append(t)            continue        out.append("NEG_" + t if negating else t)        if t in NEGATIONS:            negating = True    return outprint(mark_negation("this was not good but the ending was great".split()))# ['this', 'was', 'not', 'NEG_good', 'but', 'the', 'ending', 'was', 'great']

NEG_good is now a distinct vocabulary entry that can learn its own, negative embedding, entirely separate from good. The clause boundary stops the marking from running away — without it, "not good, but the ending was great" would mark great as negated too, which is the opposite of what you want.

This is most valuable for bag-of-words and linear models, where negation is otherwise invisible. A well-trained BiLSTM learns much of this from data, but the explicit marker still tends to help on smaller datasets where the model has less evidence to learn from.

More than two classes

Star ratings are ordinal — 1 to 5, ordered — and this changes the right loss.

FramingOutput layerLossTreats 1★ vs 5★ error as
Multi-class5 logitsCross-entropySame as 4★ vs 5★ — wrong
Regression1 valueMSE16× worse than 4★ vs 5★
Ordinal (cumulative)4 binary outputsBCE on "is rating > k?"Correctly ordered, respects discreteness

Plain cross-entropy over five classes is the default and it is subtly wrong: it treats the five labels as unrelated categories, so predicting 1★ for a 5★ review costs exactly what predicting 4★ costs. Regression fixes the ordering but produces outputs like 3.7 that you then have to round, and it assumes the gaps between stars are equal. The cumulative-link formulation — four binary classifiers answering "is this above 1?", "above 2?" and so on — usually performs best and is barely more code.

Reading your model's mistakes

The highest-value hour you can spend on a sentiment model is not tuning it. It is printing fifty misclassified examples and sorting them into categories by hand.

Python
model.eval()errors = []with torch.no_grad():    for x, lengths, y in val_loader:        logits, alpha = model(x.to(device), lengths)        probs = torch.sigmoid(logits).cpu()        for i in range(len(y)):            if (probs[i] > 0.5) != bool(y[i]):                errors.append({                    "text": decode(x[i], inv_vocab),                    "true": int(y[i]),                    "prob": float(probs[i]),                    "top_tokens": top_attended(x[i], alpha[i], inv_vocab, k=5),                })errors.sort(key=lambda e: abs(e["prob"] - e["true"]), reverse=True)

Sorting by confidence puts the model's most confident mistakes first, and those are the informative ones. Then categorise:

Error categoryWhat it points atWhat to do
Negation missedModel is not composingNegation marking; more layers; check not is not a stopword
SarcasmGenuinely hardAccept a floor; do not chase it with capacity
Mixed sentiment forced to one labelTask framingMove to aspect-based, or add a "mixed" class
Long review, verdict earlyTruncation or poolingRaise max_len; attention pooling
Domain vocabularyTrain/production mismatchFine-tune on in-domain data
The label is simply wrongData qualityFix the labels — this is more common than people expect

Attention weights make this diagnosis far faster. If a model calls "not good at all" positive and its attention mass sits on good with almost nothing on not, you know precisely what is broken, and no amount of adding hidden units will fix it.

Taking it to production

Three things separate a model that scores well in a notebook from one that works in a live system.

Ship preprocessing and vocabulary with the weights. The model is a function of the exact token-to-index mapping it was trained with. Save the vocabulary, the tokeniser configuration, and the maximum length in the same artefact as the weights. A mismatch here produces no error — indices simply point at the wrong embeddings, and accuracy silently collapses.

Return the probability, not just the label. Downstream systems need to distinguish 0.51 from 0.99. Routing a borderline case to a human is only possible if the confidence survives the API boundary.

Watch for drift. Language moves. New products, new slang, new complaint patterns. Track the proportion of unknown tokens in live traffic and the distribution of predicted probabilities. A rising unknown-token rate or predictions clustering near 0.5 both mean the model is seeing text unlike its training data — and both show up long before anyone notices the accuracy has dropped.

Finally, keep a fixed set of hand-written adversarial cases and run them on every retrain: a negation, a but-contrast, a mixed review, a sarcastic one. Aggregate metrics move slowly and hide regressions. A model that starts calling "not good" positive has broken in a way that a 0.3% accuracy dip will never tell you about.