Natural Language Processing Basics

Text Classification & Named Entity Recognition


Two requests land on your desk in the same week.

Request one: route incoming support emails to the right team — billing, technical, shipping, or returns.

Request two: from the same emails, pull out every order number, product name, and customer name so they can be linked to records in the database.

These sound like the same kind of problem. They are not, and the difference is not about difficulty. It is about the shape of the output.

Text
Input:  "My order 84-2213 for the Aeron chair has not arrived. — Priya Nair"Task 1 output:  "shipping"                       one label, whole documentTask 2 output:  My/O order/O 84-2213/B-ORDER for/O the/O                Aeron/B-PRODUCT chair/I-PRODUCT has/O not/O                arrived/O Priya/B-PERSON Nair/I-PERSON                                                 one label per token

Task 1 collapses a sequence into a single decision. Task 2 keeps the sequence and decides at every position. That distinction drives every other choice — the output layer, the loss, the masking, and above all the evaluation metric, which is where most people get into trouble.

Text classificationNamed entity recognition
Output shape(batch,classes)(\text{batch}, \text{classes})(batch,seq,tags)(\text{batch}, \text{seq}, \text{tags})
Sequence handlingPool it into one vectorKeep every position
LossCross-entropy over one predictionCross-entropy over every non-pad position
Labels per example1As many as there are tokens
Predictions are independent?N/ANo — adjacent tags constrain each other
Natural metricAccuracy, macro-F1Entity-level precision/recall/F1
BIO tags: the span is in the prefixesAngelaMerkelvisitedNewYorkB-PERI-PEROB-LOCI-LOCtokentagB starts a span, I continues it, O is outside — an I with no B before it is an impossible sequence.
Tagging each token alone lets the model emit that impossible pair, which is why the tags must constrain each other.

Text classification beyond two classes

Binary classification uses one output and a sigmoid. With KK classes you emit KK logits and apply a softmax, which turns them into probabilities summing to 1:

P(y=k∣x)=ezk∑j=1KezjP(y = k \mid x) = \frac{e^{z_k}}{\sum_{j=1}^{K} e^{z_j}}

The loss is cross-entropy, which for a single example with true class cc reduces to −log⁡P(y=c∣x)-\log P(y = c \mid x). In PyTorch, nn.CrossEntropyLoss expects raw logits and applies the log-softmax itself. Applying a softmax in your model and then passing the result to CrossEntropyLoss is a real and common bug: the loss is computed on already-normalised values, gradients are wrong, and the model trains to something mediocre without ever erroring.

Multi-class versus multi-label

A support email might belong to two teams at once. That is a different problem and needs a different output layer.

Multi-classMulti-label
QuestionWhich one of these?Which of these, possibly several?
ActivationSoftmax over KKSigmoid on each of KK independently
LossCrossEntropyLossBCEWithLogitsLoss
Probabilities sum to 1YesNo
PredictionargmaxEvery class above a threshold

Using softmax for a multi-label task actively fights you: the normalisation forces the classes to compete, so raising confidence in "billing" mechanically lowers confidence in "shipping" even when both are true.

Class imbalance

Real routing data is never balanced. If 70% of emails are technical and 3% are returns, an unweighted model learns to ignore the small class entirely — predicting "technical" always is a strong local optimum.

Python
import torchimport numpy as npfrom sklearn.utils.class_weight import compute_class_weightclasses = np.unique(y_train)weights = compute_class_weight("balanced", classes=classes, y=y_train)criterion = nn.CrossEntropyLoss(    weight=torch.tensor(weights, dtype=torch.float).to(device))

The "balanced" weights are n/(K⋅nk)n / (K \cdot n_k), so a class with a tenth of the examples gets ten times the loss weight per example. This does not create information — it just stops the model from finding it profitable to ignore rare classes.

Macro-F1 versus micro-F1

How you average across classes changes what you are measuring. Take four classes on 1,000 emails:

ClassSupportPrecisionRecallF1
technical7000.940.970.955
billing2000.850.800.824
shipping700.620.510.560
returns300.400.200.267

Micro-F1 pools all predictions before computing the score, so every example counts equally. Here it comes out at roughly 0.88 — dominated by the 700 technical emails.

Macro-F1 averages the per-class F1 scores, so every class counts equally: (0.955+0.824+0.560+0.267)/4=0.652(0.955 + 0.824 + 0.560 + 0.267)/4 = 0.652.

0.88 versus 0.65 from the same predictions. Neither is wrong; they answer different questions. Micro tells you how often the system is right. Macro tells you whether it works for every category. If the returns team never receives a correctly-routed email, micro-F1 will not tell you and macro-F1 will shout about it. Report macro-F1 whenever the small classes matter.

Pooling: how to collapse a sequence

A recurrent encoder produces one vector per position and you need one vector per document. The choice matters more than people expect.

StrategyMechanismFails when
Last hidden stateTake hnh_nThe decisive content is early in a long document
Mean poolingAverage over positionsOne decisive word is diluted by hundreds of neutral ones
Max poolingElement-wise maxCheap and surprisingly strong; ignores how much of the text agreed
Attention poolingLearned weighted averageAdds parameters; weights are easy to over-interpret
Python
class AttentionPool(nn.Module):    def __init__(self, dim, attn_dim=128):        super().__init__()        self.proj = nn.Linear(dim, attn_dim)        self.v = nn.Linear(attn_dim, 1, bias=False)    def forward(self, states, mask):        # states: (B, T, dim)  mask: (B, T) True where real        scores = self.v(torch.tanh(self.proj(states))).squeeze(-1)        scores = scores.masked_fill(~mask, float("-inf"))        alpha = torch.softmax(scores, dim=1)        return torch.bmm(alpha.unsqueeze(1), states).squeeze(1), alpha

The masked_fill is not optional. Padding positions produce scores like everything else, and escoree^{\text{score}} at a pad position is a positive number that steals probability mass from real tokens. Feeding −∞-\infty zeroes them exactly.

Python
class DocumentClassifier(nn.Module):    def __init__(self, vocab_size, num_classes, embed_dim=100,                 hidden_dim=128, num_layers=2, dropout=0.4, pad_idx=0):        super().__init__()        self.pad_idx = pad_idx        self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=pad_idx)        self.lstm = nn.LSTM(embed_dim, hidden_dim, num_layers=num_layers,                            bidirectional=True, batch_first=True,                            dropout=dropout if num_layers > 1 else 0.0)        self.pool = AttentionPool(hidden_dim * 2)        self.dropout = nn.Dropout(dropout)        self.fc = nn.Linear(hidden_dim * 2, num_classes)    def forward(self, x, lengths):        mask = x != self.pad_idx        packed = nn.utils.rnn.pack_padded_sequence(            self.dropout(self.embedding(x)), lengths.cpu(),            batch_first=True, enforce_sorted=False)        packed_out, _ = self.lstm(packed)        out, _ = nn.utils.rnn.pad_packed_sequence(            packed_out, batch_first=True, total_length=x.size(1))        pooled, alpha = self.pool(out, mask)        return self.fc(self.dropout(pooled)), alpha

Named entity recognition

NER finds spans of text that refer to real-world things and labels them by type. Standard types are person, organisation, location, date, money and percentage; domain systems add their own — gene names, drug names, part numbers.

Two properties make it harder than it looks.

Entities are spans, not words. Bank of England is one entity across three tokens. Getting two of the three right is not partial credit — it is a wrong entity.

The same string has different types in different contexts. Washington is a person, a state, a city, or a football club depending entirely on the surrounding words. This is why NER models need to read in both directions.

BIO tagging

Sequence models predict one label per token, so spans have to be encoded as per-token labels. The BIO scheme does it with three prefixes:

  • B- — Beginning of an entity of this type
  • I- — Inside (a continuation of) an entity of this type
  • O — Outside any entity
Text
The    Bank    of      England  raised  rates  in   March  .O      B-ORG   I-ORG   I-ORG    O       O      O    B-DATE O

Why not simply tag every token with its type and skip the prefixes? Because you could not tell adjacent entities apart:

Text
Without B/I:   Priya  Nair   Amit   Shah               PERSON PERSON PERSON PERSON     -> one entity? two? four?With B/I:      B-PER  I-PER  B-PER  I-PER      -> unambiguously two people

The B- prefix marks exactly where one entity ends and the next begins. With TT entity types the tag set has 2T+12T + 1 labels.

The tagger

Python
class BiLSTMTagger(nn.Module):    def __init__(self, vocab_size, num_tags, embed_dim=100,                 hidden_dim=128, num_layers=2, dropout=0.4, pad_idx=0):        super().__init__()        self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=pad_idx)        self.lstm = nn.LSTM(embed_dim, hidden_dim, num_layers=num_layers,                            bidirectional=True, batch_first=True,                            dropout=dropout if num_layers > 1 else 0.0)        self.dropout = nn.Dropout(dropout)        self.fc = nn.Linear(hidden_dim * 2, num_tags)    def forward(self, x):        emb = self.dropout(self.embedding(x))        out, _ = self.lstm(emb)        return self.fc(self.dropout(out))      # (B, T, num_tags)

Bidirectionality is doing heavy lifting here. Tagging Washington requires the words that follow it, which a forward-only model has not read yet.

Masking padding out of the loss

This is the step people skip, and it produces a specific, misleading symptom.

Python
IGNORE = -100# Build targets so padded positions are IGNOREtags = tags.masked_fill(x == PAD, IGNORE)criterion = nn.CrossEntropyLoss(ignore_index=IGNORE)logits = model(x)                                    # (B, T, num_tags)loss = criterion(logits.reshape(-1, num_tags), tags.reshape(-1))

Without ignore_index, padded positions contribute to the loss. In a typical batch, most positions are padding, so the model's fastest route to a low loss is to predict the pad tag everywhere. It will report 96% token accuracy while finding essentially no entities.

The same trap applies to metrics. In a corpus of ordinary English, roughly 85–90% of real tokens are O. A model that predicts O for everything scores about 88% token accuracy and extracts nothing. Token accuracy is not a usable metric for NER.

Why predictions need to constrain each other

A per-token classifier makes each decision independently, so it can produce tag sequences that are structurally impossible:

Text
Bank    of     EnglandB-ORG   O      I-ORG      <- I-ORG following O is invalidPriya   NairI-PER   I-PER              <- an entity cannot start with I-Bank    of     EnglandB-ORG   I-PER  I-ORG      <- type switches mid-entity

Nothing in the architecture forbids these. The model has learned per-position probabilities and each one is locally plausible.

A conditional random field layer fixes this by scoring the whole tag sequence rather than each tag alone. It adds a learned transition matrix AA, where AijA_{ij} is the score of moving from tag ii to tag jj:

score(x,y)=∑t=1T(Pt,yt+Ayt−1,yt)\text{score}(x, y) = \sum_{t=1}^{T} \left( P_{t, y_t} + A_{y_{t-1}, y_t} \right)

During training the CRF learns strongly negative transitions for illegal moves — O → I-ORG, B-PER → I-ORG — from the data alone. At inference, the Viterbi algorithm finds the highest-scoring valid sequence in O(T⋅K2)O(T \cdot K^2) time.

Python
# pip install pytorch-crf   (imported as torchcrf; the PyPI package named "torchcrf" is a different library)from torchcrf import CRFclass BiLSTMCRF(nn.Module):    def __init__(self, vocab_size, num_tags, **kw):        super().__init__()        self.tagger = BiLSTMTagger(vocab_size, num_tags, **kw)        self.crf = CRF(num_tags, batch_first=True)    def loss(self, x, tags, mask):        emissions = self.tagger(x)        return -self.crf(emissions, tags, mask=mask, reduction="mean")    def decode(self, x, mask):        return self.crf.decode(self.tagger(x), mask=mask)   # list of tag lists

A CRF typically adds 1–3 entity-level F1 points over a plain BiLSTM tagger. The gain comes almost entirely from eliminating malformed spans, and it is largest when entities are long or several types are adjacent.

Turning tags back into entities

Python
def extract_entities(tokens, tags):    """Return (text, type, start, end) for each entity."""    entities, current = [], None    for i, (tok, tag) in enumerate(zip(tokens, tags)):        if tag.startswith("B-"):            if current:                entities.append(current)            current = {"tokens": [tok], "type": tag[2:], "start": i, "end": i}        elif tag.startswith("I-") and current and current["type"] == tag[2:]:            current["tokens"].append(tok)            current["end"] = i        else:                       # O, or an invalid I- with no open entity            if current:                entities.append(current)            current = None    if current:        entities.append(current)    return [(" ".join(e["tokens"]), e["type"], e["start"], e["end"])            for e in entities]toks = "The Bank of England raised rates in March".split()tags = ["O", "B-ORG", "I-ORG", "I-ORG", "O", "O", "O", "B-DATE"]print(extract_entities(toks, tags))# [('Bank of England', 'ORG', 1, 3), ('March', 'DATE', 7, 7)]

Note the else branch. A stray I- with no open entity, or one whose type does not match, is dropped rather than crashing. Without a CRF, your model will produce these, and the extraction code has to survive them.

Two ways to score NER, and why they disagree

This is where NER evaluation goes wrong most often. Take one sentence:

Text
tokens: The    Bank    of     England  raisedgold:   O      B-ORG   I-ORG  I-ORG    Opred:   O      B-ORG   I-ORG  O        O

Token-level: 4 of 5 tokens correct → 80% accuracy. Sounds good.

Entity-level: gold contains one entity, Bank of England. The prediction contains one entity, Bank of. They are not the same span, so this is one false positive and one false negative. Precision 0, recall 0, F1 = 0.

80% and 0% for the same prediction. The entity-level score is the honest one, because a downstream system looking up "Bank of" in a database of organisations finds nothing. Partial spans are not partially useful.

Python
# pip install seqevalfrom seqeval.metrics import classification_report, f1_scorey_true = [["O", "B-ORG", "I-ORG", "I-ORG", "O"]]y_pred = [["O", "B-ORG", "I-ORG", "O",     "O"]]print(f1_score(y_true, y_pred))          # 0.0print(classification_report(y_true, y_pred))

seqeval implements entity-level scoring correctly, including the boundary rules. Do not compute NER metrics with sklearn.metrics on flattened tag lists — that gives you the token-level number, which flatters your model by a wide margin and is dominated by the O class.

MetricCountsReports for the exampleUse it
Token accuracyCorrect tags / all tags80%Never, for NER
Token F1 excluding OCorrect non-O tags80%Debugging only
Entity F1 (exact)Exactly matching spans and types0%The one to report

What breaks in practice

SymptomCauseFix
96% token accuracy, no entities foundPadding or O dominating the lossignore_index; report entity F1
Malformed tag sequences (O then I-ORG)Independent per-token predictionsAdd a CRF, or repair spans post hoc
Classifier ignores small classesImbalanceClass weights; report macro-F1
Validation score far above productionNear-duplicate documents split across train and testDeduplicate, and split by source or date
Tags misaligned with tokensRetokenised text without remapping labelsKeep tokens and labels in lockstep; for subword tokenisers, label the first sub-token and set the rest to -100
Rare entity types score near zeroToo few examples of that typeMerge types, or gather targeted data — capacity will not help

The tokenisation-alignment row deserves emphasis because it is the quietest failure. Your NER data is annotated over one tokenisation. If you feed the raw text to a different tokeniser — particularly a subword tokeniser that splits England into Eng + ##land — the label at index 3 no longer describes the token at index 3, and every label after the first split word is off by one or more. The model trains on shuffled labels and reports something around 40 F1 for no visible reason.

What this means when you build one

Two habits carry most of the value.

Choose the metric before you choose the model. For classification, decide whether you care about overall correctness (micro) or every class working (macro), because they can differ by more than 20 points and only one of them matches what your users need. For NER, use entity-level F1 from seqeval and treat token accuracy as a debugging aid only. The metric determines what you optimise, and optimising the wrong one produces a model that looks finished and is not.

Build the baseline that the neural model has to beat. For classification, that is TF-IDF plus logistic regression — it trains in seconds and often lands within a few points of a BiLSTM on topical tasks. For NER, it is a gazetteer: a lookup table of known entity strings. Both are unfashionable and both are strong where the task is genuinely about vocabulary rather than context. If your BiLSTM only matches the gazetteer, the task did not need a sequence model, and you have just learned something valuable about the problem rather than about the architecture.