Natural Language Processing Basics

Mini-Project: IMDb Sentiment Analysis System


A notebook cell that prints val_accuracy: 0.891 is not a sentiment analysis system. It is a number. The system is the thing that takes a string of text over HTTP and returns a label and a confidence, correctly, on input nobody has seen — including input from next month.

The gap between those two things is where this project lives. You will build a classifier for IMDb movie reviews, and the interesting work is not the model architecture. It is the decisions on either side of it: how long to truncate reviews, what to do with a word the vocabulary has never seen, whether 0.5 is the right threshold, and how to package the model so that the thing running in production computes the same function as the thing you validated.

The system, not the notebook numberRaw reviewtext over HTTPCleaning thatkeeps negationsVocabulary andpadded batchTrained model,best checkpointLabel plus acalibrated scoretopbottomDeduplicate before splitting, or the same review scores you in both training and test.
val_accuracy 0.891 measures one layer of this stack; everything above and below it can still be wrong.

The specification

Build an end-to-end binary sentiment classifier over film reviews and expose it as an inference service.

DeliverableRequirement
Data pipelineLoad raw reviews, clean, tokenise, build vocabulary, index
ModelBidirectional multi-layer LSTM with attention pooling over variable-length input
Accuracy≥ 88% on a held-out validation set
Length handlingCorrect results on 20-word and 800-word reviews alike
AugmentationAt least one technique implemented and honestly measured
Baseline comparisonA fine-tuned transformer, scored on the same split
ServiceHTTP endpoint returning label and probability, loading a saved artefact

The 88% target is chosen deliberately. A bag-of-words model with logistic regression reaches about 87% on this dataset in under a minute. If your recurrent model lands at 86%, it is not a slightly worse model — it is a much more expensive model that lost to a linear one, and something is wrong.

The dataset, and what is hiding in it

The IMDb review corpus holds 50,000 reviews, split 25,000 for training and 25,000 for testing, with balanced positive and negative labels. Reviews were selected to be strongly polarised: ratings of 1–4 count as negative, 7–10 as positive, and the ambiguous middle was excluded.

Python
import osdef load_imdb(split_dir):    texts, labels = [], []    for name, label in (("pos", 1), ("neg", 0)):        folder = os.path.join(split_dir, name)        for fn in sorted(os.listdir(folder)):            with open(os.path.join(folder, fn), encoding="utf-8") as f:                texts.append(f.read())            labels.append(label)    return texts, labelstrain_texts, train_labels = load_imdb("./aclImdb/train")test_texts,  test_labels  = load_imdb("./aclImdb/test")

Four properties of this data will affect your results, and three of them are easy to miss.

The text contains HTML. Reviews are full of <br /> tags. Leave them in and br becomes one of the most frequent tokens in your entire vocabulary, occupying an embedding row and contributing nothing.

The polarisation makes it easier than real feedback. Excluding 5- and 6-star reviews removes the hardest cases. A model at 90% here may sit at 75% on live customer feedback, which contains the lukewarm reviews this corpus threw away. Do not report the IMDb number as if it predicts production behaviour.

Lengths vary enormously. The mean is around 230 words; the longest exceed 2,400. This single fact drives your most consequential hyperparameter.

There are near-duplicates. Some reviewers post variations of the same text. If a near-duplicate straddles your train and validation split, your validation score is measuring memorisation.

Look at the length distribution before choosing anything

Python
import numpy as nplengths = np.array([len(t.split()) for t in train_texts])for p in (50, 75, 90, 95, 99):    print(f"p{p}: {np.percentile(lengths, p):.0f} words")print(f"mean {lengths.mean():.0f}  max {lengths.max()}")
Text
p50: 174 wordsp75: 284 wordsp90: 458 wordsp95: 598 wordsp99: 913 wordsmean 234  max 2470

Now the trade-off is visible:

max_lenReviews kept wholeRelative computeComment
100~12%0.4×Fast, but most reviews lose their ending — where verdicts live
250~70%1.0×The sensible default
500~92%2.0×Worth it if the extra accuracy shows up on validation
1000~99%4.0×Four times the cost for the last 8% of reviews

One detail matters more than the number itself: truncate from the front, not the back. Reviews build to a verdict. Keeping the first 250 words of a 600-word review discards the conclusion; keeping the last 250 keeps it. Try both and measure — on this dataset, tail truncation usually wins by a point or more.

Cleaning that helps and cleaning that hurts

Python
import redef clean(text):    text = re.sub(r"<br\s*/?>", " ", text)         # the HTML that is actually there    text = re.sub(r"<[^>]+>", " ", text)          # anything else tag-shaped    text = text.lower()    text = re.sub(r"http\S+|www\.\S+", " url ", text)    text = re.sub(r"([!?])\1+", r" \1\1 ", text)   # "!!!!!" -> " !! "  (keep intensity)    text = re.sub(r"[^a-z0-9!?'\s]", " ", text)    # keep ! ? and apostrophes    text = re.sub(r"\s+", " ", text).strip()    return textdef tokenise(text):    return clean(text).split()

Two choices in that function are deliberate departures from the standard recipe, and both are worth defending.

Exclamation and question marks are kept. They carry sentiment intensity. Collapsing runs of them stops !!!, !!!! and !!!!! from becoming three separate vocabulary entries that each appear a handful of times.

Stopwords are not removed. The standard English stopword list contains not, no and never. Applying it turns "this was not good" into good — you have not cleaned the review, you have inverted it. If you want a smaller vocabulary, use a minimum frequency threshold instead, which achieves the same thing without deleting the negations.

Deduplicate before splitting

Python
seen, keep = set(), []for i, t in enumerate(train_texts):    h = hash(clean(t)[:400])          # first 400 chars of cleaned text    if h not in seen:        seen.add(h)        keep.append(i)print(f"removed {len(train_texts) - len(keep)} near-duplicates")

Vocabulary and batching

Python
from collections import Counterimport torchfrom torch.utils.data import Dataset, DataLoaderfrom torch.nn.utils.rnn import pad_sequencePAD, UNK = 0, 1def build_vocab(tokenised_docs, max_size=25000, min_count=2):    counts = Counter(w for doc in tokenised_docs for w in doc)    vocab = {"<pad>": PAD, "<unk>": UNK}    for word, c in counts.most_common():        if c < min_count or len(vocab) >= max_size:            break        vocab[word] = len(vocab)    return vocabclass ReviewDataset(Dataset):    def __init__(self, docs, labels, vocab, max_len=250):        self.items = []        for tokens, y in zip(docs, labels):            ids = [vocab.get(t, UNK) for t in tokens][-max_len:]   # tail            self.items.append((torch.tensor(ids), len(ids), y))        self.labels = labels    def __len__(self):        return len(self.items)    def __getitem__(self, i):        return self.items[i]def collate(batch):    seqs, lengths, labels = zip(*batch)    return (pad_sequence(seqs, batch_first=True, padding_value=PAD),            torch.tensor(lengths),            torch.tensor(labels, dtype=torch.float))

Two rules govern this step, and breaking either inflates your reported score.

Build the vocabulary on the training split only. Not on train plus validation, and certainly not on the test set. A vocabulary derived from data the model is supposed to have never seen is a leak: the identity of rare words in your evaluation set has influenced what the model can represent.

Set min_count to at least 2. A word appearing once in 25,000 reviews receives exactly one gradient update during the whole training run. Its embedding stays essentially at its random initialisation and contributes noise at inference. On this dataset, dropping singletons removes over 40% of the vocabulary and costs nothing.

The model

Python
import torch.nn as nnclass SentimentModel(nn.Module):    def __init__(self, vocab_size, embed_dim=100, hidden_dim=128,                 num_layers=2, dropout=0.4, pretrained=None):        super().__init__()        self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=PAD)        if pretrained is not None:            self.embedding.weight.data.copy_(torch.tensor(pretrained))            self.embedding.weight.requires_grad = False   # unfrozen at epoch 2        self.lstm = nn.LSTM(embed_dim, hidden_dim, num_layers=num_layers,                            bidirectional=True, batch_first=True,                            dropout=dropout if num_layers > 1 else 0.0)        self.attn = nn.Linear(hidden_dim * 2, 1)        self.dropout = nn.Dropout(dropout)        self.fc = nn.Linear(hidden_dim * 2, 1)        # Bias the forget gate towards remembering at initialisation        for name, p in self.lstm.named_parameters():            if "bias" in name:                n = p.size(0)                p.data[n // 4: n // 2].fill_(1.0)    def forward(self, x, lengths):        mask = x != PAD        emb = self.dropout(self.embedding(x))        packed = nn.utils.rnn.pack_padded_sequence(            emb, lengths.cpu(), batch_first=True, enforce_sorted=False)        packed_out, _ = self.lstm(packed)        out, _ = nn.utils.rnn.pad_packed_sequence(            packed_out, batch_first=True, total_length=x.size(1))        scores = self.attn(out).squeeze(-1).masked_fill(~mask, float("-inf"))        alpha = torch.softmax(scores, dim=1)        context = torch.bmm(alpha.unsqueeze(1), out).squeeze(1)        return self.fc(self.dropout(context)).squeeze(-1), alpha

Four things in this class are load-bearing, and each corresponds to a common way this project goes wrong.

LineWhy it is thereWhat happens without it
padding_idx=PADPad embedding stays zero, gets no gradientThe model learns a vector for "nothing"
pack_padded_sequenceLSTM stops at each sequence's true endA 20-word review is pushed through 230 steps of padding; its signal decays away
masked_fill(~mask, -inf)Pad positions get zero attention weightPadding steals probability mass from real tokens
Forget-gate bias = 1Cell starts out inclined to rememberSlower convergence, weaker long-range behaviour

Attention pooling rather than the final hidden state is the right call for this data. Reviews are long and the decisive clause can be anywhere; attention lets the model weight it wherever it sits. It also gives you a debugging tool, since you can print which tokens the model actually used.

Pretrained embeddings

Python
import numpy as npimport gensim.downloader as apidef build_embedding_matrix(vocab, dim=100):    kv = api.load(f"glove-wiki-gigaword-{dim}")    rng = np.random.default_rng(0)    matrix = rng.normal(0, 0.1, (len(vocab), dim)).astype(np.float32)    matrix[PAD] = 0.0    hits = 0    for word, idx in vocab.items():        if word in kv:            matrix[idx] = kv[word]            hits += 1    print(f"initialised {hits}/{len(vocab)} rows from GloVe")    return matrix

On 25,000 reviews, initialising from pretrained vectors is typically worth 2–3 accuracy points and gets you there in fewer epochs. Freeze the embedding layer for the first two epochs while the layers above it stabilise, then unfreeze it with a lower learning rate.

Training

Python
criterion = nn.BCEWithLogitsLoss()optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(    optimizer, mode="max", factor=0.5, patience=1)def run(loader, train):    model.train() if train else model.eval()    total, correct, n = 0.0, 0, 0    with torch.set_grad_enabled(train):        for x, lengths, y in loader:            x, y = x.to(device), y.to(device)            logits, _ = model(x, lengths)            loss = criterion(logits, y)            if train:                optimizer.zero_grad()                loss.backward()                nn.utils.clip_grad_norm_(model.parameters(), 5.0)                optimizer.step()            total += loss.item() * y.size(0)            correct += ((torch.sigmoid(logits) > 0.5) == y.bool()).sum().item()            n += y.size(0)    return total / n, correct / nbest, patience, bad_epochs = 0.0, 3, 0for epoch in range(15):    if epoch == 2:        model.embedding.weight.requires_grad = True     # unfreeze    tr_loss, tr_acc = run(train_loader, True)    va_loss, va_acc = run(val_loader, False)    scheduler.step(va_acc)    print(f"{epoch:2d}  train {tr_acc:.4f}  val {va_acc:.4f}  loss {va_loss:.4f}")    if va_acc > best:        best, bad_epochs = va_acc, 0        torch.save(model.state_dict(), "best.pt")    else:        bad_epochs += 1        if bad_epochs >= patience:            print("early stopping")            break

What a healthy run looks like on this data, and what each pathology means:

PatternDiagnosisAction
Train 0.93 / val 0.89, both risingHealthyKeep going
Train 0.99 / val 0.86, val fallingOverfittingMore dropout, freeze embeddings longer, stop earlier
Both stuck near 0.50Not learning at allCheck labels are in the batch correctly; check learning rate; check you are not applying sigmoid twice
Loss becomes NaNExploding gradientsClipping is present but the learning rate is too high
Val accuracy 0.87 and flat from epoch 1The model is not using sequence structureVerify packing; verify attention mask; compare against bag-of-words

That last row is the one worth checking early. If your BiLSTM matches a linear bag-of-words model exactly, the recurrence is probably not contributing — usually because padding is being processed as content.

The transformer comparison

Run a fine-tuned pretrained transformer on the identical split so the comparison is meaningful.

Python
from transformers import AutoTokenizer, AutoModelForSequenceClassificationfrom transformers import get_linear_schedule_with_warmupfrom torch.optim import AdamWname = "distilbert-base-uncased"tok = AutoTokenizer.from_pretrained(name)bert = AutoModelForSequenceClassification.from_pretrained(name, num_labels=2).to(device)enc = tok(raw_texts, truncation=True, max_length=256, padding="max_length",          return_tensors="pt")opt = AdamW(bert.parameters(), lr=2e-5, weight_decay=0.01)steps = len(bert_loader) * 3sched = get_linear_schedule_with_warmup(opt, int(0.1 * steps), steps)

Note the learning rate: 2e-5, not 1e-3. Fine-tuning at 1e-3 destroys the pretrained weights within a few hundred steps and lands you near chance. This is the single most common mistake in the comparison step, and the symptom — a large pretrained model performing worse than a small LSTM — looks like an architecture result rather than the configuration error it is.

Also note: feed the transformer raw text, not your cleaned tokens. Its own subword tokeniser was trained against a specific vocabulary, and pre-cleaning shifts every index away from what the model saw during pretraining.

ModelAccuracyParamsTrain time (GPU)CPU latency
TF-IDF + logistic regression~87%~0.3 M< 1 min< 1 ms
BiLSTM + attention, random embeddings~86%~3 M~10 min~5 ms
BiLSTM + attention, GloVe~89%~3 M~10 min~5 ms
DistilBERT fine-tuned~92%66 M~20 min~60 ms
BERT-base fine-tuned~93%110 M~40 min~110 ms

Three to four points of accuracy for twelve to twenty times the latency and twenty to forty times the parameters. Whether that is a good trade depends entirely on what the system is for — which is exactly the judgement this project exists to make you exercise.

Augmentation, measured honestly

Python
import randomfrom nltk.corpus import wordnetdef synonym_replace(tokens, n=2, protected=frozenset({"not", "no", "never", "n't"})):    tokens = list(tokens)    candidates = [i for i, t in enumerate(tokens)                  if len(t) > 3 and t not in protected and wordnet.synsets(t)]    random.shuffle(candidates)    for i in candidates[:n]:        syns = {l.name().replace("_", " ")                for s in wordnet.synsets(tokens[i]) for l in s.lemmas()}        syns.discard(tokens[i])        if syns:            tokens[i] = random.choice(sorted(syns))    return tokens

The protected set is essential. Replacing not with a WordNet "synonym" can flip a review's meaning while its label stays fixed, and you have manufactured mislabelled training data.

TechniqueCostTypical gain on 25k examplesRisk
Synonym replacementCheap+0.2 to +0.5 pointsCan alter sentiment
Random deletion (10%)Free+0 to +0.3 pointsMay delete the negation
Back-translationExpensive+0.5 to +1.5 pointsNeeds translation models
More real labelled dataExpensiveReliably the largestNone

Be sceptical here. On 25,000 balanced examples augmentation buys very little, because the dataset is already large enough for a 3M-parameter model. Its value rises sharply when you have a few thousand examples or a badly imbalanced class. Measure it on your split rather than assuming it helps — and always augment after splitting, or an augmented copy of a training review will appear in validation.

Packaging and serving

This is where projects that scored well quietly break, and the cause is almost always the same: the model was saved and the vocabulary was not.

The weights are meaningless without the exact token-to-index mapping they were trained with. Index 4,821 means brilliant only if the same vocabulary is loaded. Rebuild the vocabulary at deploy time from slightly different data and every index shifts. There is no error, no warning, and no crash — the model simply reads a different sentence from the one you sent it and predicts near-randomly.

Python
import jsontorch.save({    "state_dict": model.state_dict(),    "vocab": vocab,    "config": {"embed_dim": 100, "hidden_dim": 128,               "num_layers": 2, "max_len": 250, "threshold": 0.5},}, "sentiment_model.pt")

Weights, vocabulary, architecture configuration and threshold travel together as one artefact. Anything less is not deployable.

Python
from fastapi import FastAPIfrom pydantic import BaseModelapp = FastAPI()ckpt = torch.load("sentiment_model.pt", map_location="cpu")vocab, cfg = ckpt["vocab"], ckpt["config"]model = SentimentModel(len(vocab), cfg["embed_dim"], cfg["hidden_dim"],                       cfg["num_layers"])model.load_state_dict(ckpt["state_dict"])model.eval()class Review(BaseModel):    text: str@app.post("/predict")def predict(review: Review):    tokens = tokenise(review.text)                     # the same function as training    ids = [vocab.get(t, UNK) for t in tokens][-cfg["max_len"]:] or [UNK]    x = torch.tensor([ids])    with torch.no_grad():        logits, alpha = model(x, torch.tensor([len(ids)]))        prob = torch.sigmoid(logits).item()    unknown_rate = sum(i == UNK for i in ids) / len(ids)    return {        "label": "positive" if prob > cfg["threshold"] else "negative",        "confidence": round(prob if prob > 0.5 else 1 - prob, 4),        "probability_positive": round(prob, 4),        "unknown_token_rate": round(unknown_rate, 3),    }

Four details in that handler are doing real work. model.eval() disables dropout — forget it and every request gets a slightly different answer. torch.no_grad() avoids building a graph you will never backpropagate through, roughly halving latency and memory. or [UNK] handles empty input, which would otherwise pass a zero-length tensor to the LSTM and raise. And returning the probability rather than only a label lets whatever calls this route borderline cases to a human.

The unknown_token_rate is the cheapest monitoring you will ever add. If it is 3% in validation and 22% in production, your traffic no longer resembles your training data — and you will know that weeks before anyone notices the accuracy has slipped.

Reading the errors, and knowing when it is finished

Once the number clears 88%, the highest-value work is not tuning. It is printing your fifty most confident mistakes and sorting them by hand.

Python
model.eval()errors = []with torch.no_grad():    for x, lengths, y in val_loader:        logits, alpha = model(x.to(device), lengths)        probs = torch.sigmoid(logits).cpu()        for i in range(len(y)):            if (probs[i] > 0.5) != bool(y[i]):                top = alpha[i].topk(5).indices.tolist()                errors.append({                    "prob": float(probs[i]),                    "true": int(y[i]),                    "attended": [inv_vocab[int(x[i][j])] for j in top],                })errors.sort(key=lambda e: abs(e["prob"] - e["true"]), reverse=True)

The attention weights turn a vague "it got this wrong" into a specific diagnosis. If the model calls "not good at all" positive and its attention sits on good with nothing on not, you know precisely what is broken, and adding hidden units will not fix it.

Expect the residual errors to sort into a small number of families: sarcasm, mixed reviews where the label is one-sided, plot summaries that quote negative events in a positive review, and reviews where the gold label is simply wrong. Only the first two are model problems; the last is a data problem, and it is more common than people expect.

Finally, keep a handful of hand-written cases and run them on every retrain — a negation, a but-contrast, a very short review, a very long one, and something with unusual vocabulary. Aggregate accuracy moves slowly and hides regressions. A model that starts calling "not good" positive has broken in a way that a 0.3% dip in validation accuracy will never tell you about, and those five examples take two seconds to check.