Course Content
Natural Language Processing Basics
4 sections · 10 lessons
Text Classification & Named Entity Recognition
Two requests land on your desk in the same week.
Request one: route incoming support emails to the right team — billing, technical, shipping, or returns.
Request two: from the same emails, pull out every order number, product name, and customer name so they can be linked to records in the database.
These sound like the same kind of problem. They are not, and the difference is not about difficulty. It is about the shape of the output.
Input: "My order 84-2213 for the Aeron chair has not arrived. — Priya Nair"Task 1 output: "shipping" one label, whole documentTask 2 output: My/O order/O 84-2213/B-ORDER for/O the/O Aeron/B-PRODUCT chair/I-PRODUCT has/O not/O arrived/O Priya/B-PERSON Nair/I-PERSON one label per tokenTask 1 collapses a sequence into a single decision. Task 2 keeps the sequence and decides at every position. That distinction drives every other choice — the output layer, the loss, the masking, and above all the evaluation metric, which is where most people get into trouble.
| Text classification | Named entity recognition | |
|---|---|---|
| Output shape | (batch,classes) | (batch,seq,tags) |
| Sequence handling | Pool it into one vector | Keep every position |
| Loss | Cross-entropy over one prediction | Cross-entropy over every non-pad position |
| Labels per example | 1 | As many as there are tokens |
| Predictions are independent? | N/A | No — adjacent tags constrain each other |
| Natural metric | Accuracy, macro-F1 | Entity-level precision/recall/F1 |
Text classification beyond two classes
Binary classification uses one output and a sigmoid. With K classes you emit K logits and apply a softmax, which turns them into probabilities summing to 1:
The loss is cross-entropy, which for a single example with true class c reduces to −logP(y=c∣x). In PyTorch, nn.CrossEntropyLoss expects raw logits and applies the log-softmax itself. Applying a softmax in your model and then passing the result to CrossEntropyLoss is a real and common bug: the loss is computed on already-normalised values, gradients are wrong, and the model trains to something mediocre without ever erroring.
Multi-class versus multi-label
A support email might belong to two teams at once. That is a different problem and needs a different output layer.
| Multi-class | Multi-label | |
|---|---|---|
| Question | Which one of these? | Which of these, possibly several? |
| Activation | Softmax over K | Sigmoid on each of K independently |
| Loss | CrossEntropyLoss | BCEWithLogitsLoss |
| Probabilities sum to 1 | Yes | No |
| Prediction | argmax | Every class above a threshold |
Using softmax for a multi-label task actively fights you: the normalisation forces the classes to compete, so raising confidence in "billing" mechanically lowers confidence in "shipping" even when both are true.
Class imbalance
Real routing data is never balanced. If 70% of emails are technical and 3% are returns, an unweighted model learns to ignore the small class entirely — predicting "technical" always is a strong local optimum.
1import torch2import numpy as np3from sklearn.utils.class_weight import compute_class_weight45classes = np.unique(y_train)6weights = compute_class_weight("balanced", classes=classes, y=y_train)7criterion = nn.CrossEntropyLoss(8 weight=torch.tensor(weights, dtype=torch.float).to(device)9)The "balanced" weights are n/(K⋅nk), so a class with a tenth of the examples gets ten times the loss weight per example. This does not create information — it just stops the model from finding it profitable to ignore rare classes.
Macro-F1 versus micro-F1
How you average across classes changes what you are measuring. Take four classes on 1,000 emails:
| Class | Support | Precision | Recall | F1 |
|---|---|---|---|---|
| technical | 700 | 0.94 | 0.97 | 0.955 |
| billing | 200 | 0.85 | 0.80 | 0.824 |
| shipping | 70 | 0.62 | 0.51 | 0.560 |
| returns | 30 | 0.40 | 0.20 | 0.267 |
Micro-F1 pools all predictions before computing the score, so every example counts equally. Here it comes out at roughly 0.88 — dominated by the 700 technical emails.
Macro-F1 averages the per-class F1 scores, so every class counts equally: (0.955+0.824+0.560+0.267)/4=0.652.
0.88 versus 0.65 from the same predictions. Neither is wrong; they answer different questions. Micro tells you how often the system is right. Macro tells you whether it works for every category. If the returns team never receives a correctly-routed email, micro-F1 will not tell you and macro-F1 will shout about it. Report macro-F1 whenever the small classes matter.
Pooling: how to collapse a sequence
A recurrent encoder produces one vector per position and you need one vector per document. The choice matters more than people expect.
| Strategy | Mechanism | Fails when |
|---|---|---|
| Last hidden state | Take hn | The decisive content is early in a long document |
| Mean pooling | Average over positions | One decisive word is diluted by hundreds of neutral ones |
| Max pooling | Element-wise max | Cheap and surprisingly strong; ignores how much of the text agreed |
| Attention pooling | Learned weighted average | Adds parameters; weights are easy to over-interpret |
1class AttentionPool(nn.Module):2 def __init__(self, dim, attn_dim=128):3 super().__init__()4 self.proj = nn.Linear(dim, attn_dim)5 self.v = nn.Linear(attn_dim, 1, bias=False)67 def forward(self, states, mask):8 # states: (B, T, dim) mask: (B, T) True where real9 scores = self.v(torch.tanh(self.proj(states))).squeeze(-1)10 scores = scores.masked_fill(~mask, float("-inf"))11 alpha = torch.softmax(scores, dim=1)12 return torch.bmm(alpha.unsqueeze(1), states).squeeze(1), alphaThe masked_fill is not optional. Padding positions produce scores like everything else, and escore at a pad position is a positive number that steals probability mass from real tokens. Feeding −∞ zeroes them exactly.
1class DocumentClassifier(nn.Module):2 def __init__(self, vocab_size, num_classes, embed_dim=100,3 hidden_dim=128, num_layers=2, dropout=0.4, pad_idx=0):4 super().__init__()5 self.pad_idx = pad_idx6 self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=pad_idx)7 self.lstm = nn.LSTM(embed_dim, hidden_dim, num_layers=num_layers,8 bidirectional=True, batch_first=True,9 dropout=dropout if num_layers > 1 else 0.0)10 self.pool = AttentionPool(hidden_dim * 2)11 self.dropout = nn.Dropout(dropout)12 self.fc = nn.Linear(hidden_dim * 2, num_classes)1314 def forward(self, x, lengths):15 mask = x != self.pad_idx16 packed = nn.utils.rnn.pack_padded_sequence(17 self.dropout(self.embedding(x)), lengths.cpu(),18 batch_first=True, enforce_sorted=False)19 packed_out, _ = self.lstm(packed)20 out, _ = nn.utils.rnn.pad_packed_sequence(21 packed_out, batch_first=True, total_length=x.size(1))22 pooled, alpha = self.pool(out, mask)23 return self.fc(self.dropout(pooled)), alphaNamed entity recognition
NER finds spans of text that refer to real-world things and labels them by type. Standard types are person, organisation, location, date, money and percentage; domain systems add their own — gene names, drug names, part numbers.
Two properties make it harder than it looks.
Entities are spans, not words. Bank of England is one entity across three tokens. Getting two of the three right is not partial credit — it is a wrong entity.
The same string has different types in different contexts. Washington is a person, a state, a city, or a football club depending entirely on the surrounding words. This is why NER models need to read in both directions.
BIO tagging
Sequence models predict one label per token, so spans have to be encoded as per-token labels. The BIO scheme does it with three prefixes:
- B- — Beginning of an entity of this type
- I- — Inside (a continuation of) an entity of this type
- O — Outside any entity
The Bank of England raised rates in March .O B-ORG I-ORG I-ORG O O O B-DATE OWhy not simply tag every token with its type and skip the prefixes? Because you could not tell adjacent entities apart:
Without B/I: Priya Nair Amit Shah PERSON PERSON PERSON PERSON -> one entity? two? four?With B/I: B-PER I-PER B-PER I-PER -> unambiguously two peopleThe B- prefix marks exactly where one entity ends and the next begins. With T entity types the tag set has 2T+1 labels.
The tagger
1class BiLSTMTagger(nn.Module):2 def __init__(self, vocab_size, num_tags, embed_dim=100,3 hidden_dim=128, num_layers=2, dropout=0.4, pad_idx=0):4 super().__init__()5 self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=pad_idx)6 self.lstm = nn.LSTM(embed_dim, hidden_dim, num_layers=num_layers,7 bidirectional=True, batch_first=True,8 dropout=dropout if num_layers > 1 else 0.0)9 self.dropout = nn.Dropout(dropout)10 self.fc = nn.Linear(hidden_dim * 2, num_tags)1112 def forward(self, x):13 emb = self.dropout(self.embedding(x))14 out, _ = self.lstm(emb)15 return self.fc(self.dropout(out)) # (B, T, num_tags)Bidirectionality is doing heavy lifting here. Tagging Washington requires the words that follow it, which a forward-only model has not read yet.
Masking padding out of the loss
This is the step people skip, and it produces a specific, misleading symptom.
1IGNORE = -10023# Build targets so padded positions are IGNORE4tags = tags.masked_fill(x == PAD, IGNORE)56criterion = nn.CrossEntropyLoss(ignore_index=IGNORE)7logits = model(x) # (B, T, num_tags)8loss = criterion(logits.reshape(-1, num_tags), tags.reshape(-1))Without ignore_index, padded positions contribute to the loss. In a typical batch, most positions are padding, so the model's fastest route to a low loss is to predict the pad tag everywhere. It will report 96% token accuracy while finding essentially no entities.
The same trap applies to metrics. In a corpus of ordinary English, roughly 85–90% of real tokens are O. A model that predicts O for everything scores about 88% token accuracy and extracts nothing. Token accuracy is not a usable metric for NER.
Why predictions need to constrain each other
A per-token classifier makes each decision independently, so it can produce tag sequences that are structurally impossible:
Bank of EnglandB-ORG O I-ORG <- I-ORG following O is invalidPriya NairI-PER I-PER <- an entity cannot start with I-Bank of EnglandB-ORG I-PER I-ORG <- type switches mid-entityNothing in the architecture forbids these. The model has learned per-position probabilities and each one is locally plausible.
A conditional random field layer fixes this by scoring the whole tag sequence rather than each tag alone. It adds a learned transition matrix A, where Aij is the score of moving from tag i to tag j:
During training the CRF learns strongly negative transitions for illegal moves — O → I-ORG, B-PER → I-ORG — from the data alone. At inference, the Viterbi algorithm finds the highest-scoring valid sequence in O(T⋅K2) time.
1# pip install pytorch-crf (imported as torchcrf; the PyPI package named "torchcrf" is a different library)2from torchcrf import CRF34class BiLSTMCRF(nn.Module):5 def __init__(self, vocab_size, num_tags, **kw):6 super().__init__()7 self.tagger = BiLSTMTagger(vocab_size, num_tags, **kw)8 self.crf = CRF(num_tags, batch_first=True)910 def loss(self, x, tags, mask):11 emissions = self.tagger(x)12 return -self.crf(emissions, tags, mask=mask, reduction="mean")1314 def decode(self, x, mask):15 return self.crf.decode(self.tagger(x), mask=mask) # list of tag listsA CRF typically adds 1–3 entity-level F1 points over a plain BiLSTM tagger. The gain comes almost entirely from eliminating malformed spans, and it is largest when entities are long or several types are adjacent.
Turning tags back into entities
1def extract_entities(tokens, tags):2 """Return (text, type, start, end) for each entity."""3 entities, current = [], None4 for i, (tok, tag) in enumerate(zip(tokens, tags)):5 if tag.startswith("B-"):6 if current:7 entities.append(current)8 current = {"tokens": [tok], "type": tag[2:], "start": i, "end": i}9 elif tag.startswith("I-") and current and current["type"] == tag[2:]:10 current["tokens"].append(tok)11 current["end"] = i12 else: # O, or an invalid I- with no open entity13 if current:14 entities.append(current)15 current = None16 if current:17 entities.append(current)18 return [(" ".join(e["tokens"]), e["type"], e["start"], e["end"])19 for e in entities]2021toks = "The Bank of England raised rates in March".split()22tags = ["O", "B-ORG", "I-ORG", "I-ORG", "O", "O", "O", "B-DATE"]23print(extract_entities(toks, tags))24# [('Bank of England', 'ORG', 1, 3), ('March', 'DATE', 7, 7)]Note the else branch. A stray I- with no open entity, or one whose type does not match, is dropped rather than crashing. Without a CRF, your model will produce these, and the extraction code has to survive them.
Two ways to score NER, and why they disagree
This is where NER evaluation goes wrong most often. Take one sentence:
tokens: The Bank of England raisedgold: O B-ORG I-ORG I-ORG Opred: O B-ORG I-ORG O OToken-level: 4 of 5 tokens correct → 80% accuracy. Sounds good.
Entity-level: gold contains one entity, Bank of England. The prediction contains one entity, Bank of. They are not the same span, so this is one false positive and one false negative. Precision 0, recall 0, F1 = 0.
80% and 0% for the same prediction. The entity-level score is the honest one, because a downstream system looking up "Bank of" in a database of organisations finds nothing. Partial spans are not partially useful.
1# pip install seqeval2from seqeval.metrics import classification_report, f1_score34y_true = [["O", "B-ORG", "I-ORG", "I-ORG", "O"]]5y_pred = [["O", "B-ORG", "I-ORG", "O", "O"]]67print(f1_score(y_true, y_pred)) # 0.08print(classification_report(y_true, y_pred))seqeval implements entity-level scoring correctly, including the boundary rules. Do not compute NER metrics with sklearn.metrics on flattened tag lists — that gives you the token-level number, which flatters your model by a wide margin and is dominated by the O class.
| Metric | Counts | Reports for the example | Use it |
|---|---|---|---|
| Token accuracy | Correct tags / all tags | 80% | Never, for NER |
Token F1 excluding O | Correct non-O tags | 80% | Debugging only |
| Entity F1 (exact) | Exactly matching spans and types | 0% | The one to report |
What breaks in practice
| Symptom | Cause | Fix |
|---|---|---|
| 96% token accuracy, no entities found | Padding or O dominating the loss | ignore_index; report entity F1 |
Malformed tag sequences (O then I-ORG) | Independent per-token predictions | Add a CRF, or repair spans post hoc |
| Classifier ignores small classes | Imbalance | Class weights; report macro-F1 |
| Validation score far above production | Near-duplicate documents split across train and test | Deduplicate, and split by source or date |
| Tags misaligned with tokens | Retokenised text without remapping labels | Keep tokens and labels in lockstep; for subword tokenisers, label the first sub-token and set the rest to -100 |
| Rare entity types score near zero | Too few examples of that type | Merge types, or gather targeted data — capacity will not help |
The tokenisation-alignment row deserves emphasis because it is the quietest failure. Your NER data is annotated over one tokenisation. If you feed the raw text to a different tokeniser — particularly a subword tokeniser that splits England into Eng + ##land — the label at index 3 no longer describes the token at index 3, and every label after the first split word is off by one or more. The model trains on shuffled labels and reports something around 40 F1 for no visible reason.
What this means when you build one
Two habits carry most of the value.
Choose the metric before you choose the model. For classification, decide whether you care about overall correctness (micro) or every class working (macro), because they can differ by more than 20 points and only one of them matches what your users need. For NER, use entity-level F1 from seqeval and treat token accuracy as a debugging aid only. The metric determines what you optimise, and optimising the wrong one produces a model that looks finished and is not.
Build the baseline that the neural model has to beat. For classification, that is TF-IDF plus logistic regression — it trains in seconds and often lands within a few points of a BiLSTM on topical tasks. For NER, it is a gazetteer: a lookup table of known entity strings. Both are unfashionable and both are strong where the task is genuinely about vocabulary rather than context. If your BiLSTM only matches the gazetteer, the task did not need a sequence model, and you have just learned something valuable about the problem rather than about the architecture.