Natural Language Processing Basics

Introduction to Transformers for NLP


A recurrent model is reading this sentence, one token at a time:

Text
"The report that the committee, after months of deliberation and three separate rounds of public consultation, finally published last Tuesday was widely condemned."

To resolve was condemned the model needs report — 21 tokens back, counting punctuation. For that information to arrive, it must pass through 21 consecutive state updates, each one squeezing it through the same gates alongside everything else the model has read since. The signal survives, degraded, if it survives at all.

Count the steps another way. In a recurrent network, the path length between two positions ii and jj is ∣i−j∣|i - j|. Information between distant words traverses a long chain, and every hop is a chance to be diluted.

There is a second problem, and commercially it is the bigger one. Step 24 cannot begin until step 23 has finished. A 500-token document requires 500 dependent matrix operations. A modern GPU can perform thousands of operations simultaneously and is left almost entirely idle, waiting. You cannot buy your way out of this — adding hardware does not shorten a chain of dependencies.

Recurrent layerSelf-attention layer
Path between any two positionsO(n)O(n)O(1)O(1)
Sequential operationsO(n)O(n)O(1)O(1)
Compute per layerO(n⋅d2)O(n \cdot d^2)O(n2⋅d)O(n^2 \cdot d)
Parallel across positionsNoYes

Read the last row of the compute line honestly: attention is more expensive in raw arithmetic once nn exceeds dd. It wins anyway, because that arithmetic all happens at once on hardware built for exactly that, while recurrence is a queue.

What one attention layer does to one tokenProject eachtoken to q, k, vScore: q of "it"against every kDivide by thesquare root of dSoftmax toweightsthat sum to 1Output is theweighted sum of vEvery token is scored against every other in one matrix product, so the path between any two is length 1.
The recurrent model reaches token 40 after 39 rewrites; attention reaches it in a single step, in parallel.

Self-attention, before any notation

Here is the idea without a single symbol.

Every word looks at every other word in the sentence, decides how relevant each one is to itself, and builds a new representation of itself as a weighted blend of them all.

Text
"The animal did not cross the street because it was too tired."When the model builds a representation for "it":    animal   0.51    <- the antecedent    tired    0.19    it       0.11    street   0.06    was      0.05    ...everything else shares the remaining 0.08

The representation of it becomes mostly animal, with some tired mixed in. And crucially, animal is seven tokens away but reached in one step. There is no chain. Every position is one hop from every other position.

Change one word and the weights change with it:

Text
"The animal did not cross the street because it was too wide."    street   0.48    <- now the antecedent    wide     0.22    it       0.10

Nothing was hand-coded. The scoring function was learned, and it produces different attention for different contexts.

Queries, keys and values

How does a position decide what is relevant? Each token produces three vectors from its input embedding, using three learned matrices:

VectorComputed asIts job
Query qqxWQx W^QWhat this position is looking for
Key kkxWKx W^KWhat this position offers to others
Value vvxWVx W^VThe content this position contributes if attended to

The database analogy is apt. You issue a query, match it against every key, and retrieve a blend of the values weighted by how well each key matched. The difference from a real database is that retrieval is soft — you get a weighted mixture, not a single row.

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\text{Attention}(Q, K, V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right) V

Read it as four steps:

  1. QK⊤QK^\top — dot every query against every key, giving an n×nn \times n matrix of raw relevance scores.
  2. ÷dk\div \sqrt{d_k} — scale down. This is not cosmetic; see below.
  3. softmax\text{softmax} — normalise each row so weights are positive and sum to 1.
  4. ×V\times V — take the weighted average of the value vectors.

Why divide by the square root of the dimension

Suppose qq and kk have independent components with mean 0 and variance 1. Their dot product is a sum of dkd_k such products, so it has variance dkd_k and standard deviation dk\sqrt{d_k}. With dk=64d_k = 64, typical scores land around ±8\pm 8; with dk=512d_k = 512, around ±23\pm 23.

Feed scores of that size into a softmax and see what happens:

Text
scores [1.0, 2.0, 3.0]        -> softmax [0.090, 0.245, 0.665]scores [8.0, 16.0, 24.0]      -> softmax [0.0000001, 0.0003, 0.9997]

The second distribution is effectively one-hot. Attention has become hard selection rather than a blend, and — the part that actually kills training — the softmax gradient at a saturated point is nearly zero. The layer stops learning.

Dividing by dk\sqrt{d_k} returns the scores to unit variance regardless of dimension, keeping the softmax in a range where it produces a meaningful distribution and a usable gradient.

A worked example

Three tokens, dk=2d_k = 2, arithmetic small enough to check by hand.

Text
Q = [[1, 0],     K = [[1, 0],     V = [[1, 0],     [0, 1],          [0, 1],          [0, 2],     [1, 1]]          [1, 1]]          [1, 1]]Scores = Q K^T:  row 1: [1*1+0*0, 1*0+0*1, 1*1+0*1] = [1, 0, 1]  row 2: [0, 1, 1]  row 3: [1, 1, 2]Scaled by sqrt(2) = 1.414:  row 1: [0.707, 0.000, 0.707]  row 3: [0.707, 0.707, 1.414]Softmax of row 1:  exp: [2.028, 1.000, 2.028]   sum = 5.056  weights: [0.401, 0.198, 0.401]Output for position 1 = 0.401*[1,0] + 0.198*[0,2] + 0.401*[1,1]                      = [0.401, 0] + [0, 0.396] + [0.401, 0.401]                      = [0.802, 0.797]

Position 1 attends to positions 1 and 3 equally (both score 1) and to position 2 less (score 0), and its output is the corresponding blend of value vectors. That is the entire mechanism.

Multiple heads

One attention layer produces one set of weights per position — one notion of relevance. But relevance is not one thing. Resolving a pronoun, tracking the subject of a verb, and noticing a negation are different relationships, and a single weighting cannot express all three at once.

Multi-head attention runs hh attention operations in parallel, each with its own WQ,WK,WVW^Q, W^K, W^V, each on a slice of the dimensions, then concatenates the results and projects them back.

MultiHead(X)=[head1;… ;headh]WO\text{MultiHead}(X) = \left[\text{head}_1 ; \dots ; \text{head}_h\right] W^O

With dmodel=512d_{model} = 512 and h=8h = 8, each head works in 64 dimensions, so the total compute matches single-head attention at full width — you get eight relationship types for the price of one.

Probing trained models finds heads that specialise in recognisable ways: some attend to the immediately preceding token, some link verbs to their subjects, some connect pronouns to antecedents, and a good many appear to do nothing useful and can be pruned with little loss.

Attention has no idea what order things are in

Here is a property that surprises people. Shuffle the input tokens and self-attention produces the same outputs, merely shuffled to match. Every position attends to every other with no notion of near or far, left or right.

Text
"dog bites man"  and  "man bites dog"-> identical sets of attention outputs, just reordered

A recurrent network gets order for free from the order in which it reads. Attention has to be told.

The fix is to add a position-dependent vector to each input embedding before the first layer. The original formulation uses sinusoids of different frequencies:

PE(pos,2i)=sin⁡ ⁣(pos100002i/d),PE(pos,2i+1)=cos⁡ ⁣(pos100002i/d)PE_{(pos, 2i)} = \sin\!\left(\frac{pos}{10000^{2i/d}}\right), \qquad PE_{(pos, 2i+1)} = \cos\!\left(\frac{pos}{10000^{2i/d}}\right)

Python
import numpy as npdef positional_encoding(max_len, d_model):    pos = np.arange(max_len)[:, None]    i = np.arange(d_model)[None, :]    angle = pos / np.power(10000, (2 * (i // 2)) / d_model)    pe = np.zeros((max_len, d_model))    pe[:, 0::2] = np.sin(angle[:, 0::2])    pe[:, 1::2] = np.cos(angle[:, 1::2])    return pe

Low-frequency dimensions change slowly and encode coarse position; high-frequency dimensions distinguish neighbours. Because sin⁡(a+b)\sin(a + b) expands into terms involving sin⁡a\sin a, cos⁡a\cos a, sin⁡b\sin b and cos⁡b\cos b, a fixed offset corresponds to a linear transformation of the encoding — which gives the model a route to learning relative position, not just absolute.

BERT and GPT-2 instead learn a position embedding table directly. It is simpler, works at least as well within the trained length, and cannot extrapolate beyond it. Most current large language models use a third option, rotary position embeddings (RoPE), which rotate the query and key vectors by a position-dependent angle so that attention scores depend on the relative distance between tokens.

The block

A transformer layer is attention plus a feed-forward network, each wrapped in a residual connection and layer normalisation.

Text
x ──┬──> MultiHeadAttention ──> (+) ──> LayerNorm ──┬──> FeedForward ──> (+) ──> LayerNorm ──>    └───────────────────────────┘                    └──────────────────────┘              residual                                        residual
Python
import torch.nn as nnclass TransformerBlock(nn.Module):    def __init__(self, d_model, n_heads, d_ff, dropout=0.1):        super().__init__()        self.attn = nn.MultiheadAttention(d_model, n_heads,                                          dropout=dropout, batch_first=True)        self.norm1 = nn.LayerNorm(d_model)        self.norm2 = nn.LayerNorm(d_model)        self.ff = nn.Sequential(            nn.Linear(d_model, d_ff),            nn.GELU(),            nn.Dropout(dropout),            nn.Linear(d_ff, d_model),        )        self.dropout = nn.Dropout(dropout)    def forward(self, x, key_padding_mask=None):        a, _ = self.attn(x, x, x, key_padding_mask=key_padding_mask)        x = self.norm1(x + self.dropout(a))        x = self.norm2(x + self.dropout(self.ff(x)))        return x

Each piece earns its place. Residual connections give the gradient a direct path to every layer, which is what makes 12-, 24- and 96-layer stacks trainable. Layer normalisation stabilises the scale of activations, and unlike batch normalisation it does not depend on the batch, so it behaves identically at batch size 1. The feed-forward network — usually four times as wide as dmodeld_{model} — is applied to each position independently and holds a large share of the model's parameters; attention mixes information between positions, the feed-forward layer processes it.

This is the original post-norm arrangement, with LayerNorm after each residual addition. Most current models use pre-norm instead — normalise the input to each sub-layer, add the residual afterwards, and put one final LayerNorm after the last block — because deep pre-norm stacks train more stably. In PyTorch's built-in nn.TransformerEncoderLayer that is norm_first=True.

key_padding_mask is the practical detail that gets forgotten. Without it, padding positions receive attention weight and contribute to every real token's representation. There is no error — only a model that quietly performs worse on short sequences in a batch of long ones.

Encoders, decoders, and why the distinction matters

The same block appears in two configurations, and they are suited to different tasks.

Encoder (BERT-style)Decoder (GPT-style)
Attention rangeAll positions, both directionsOnly positions to the left (causally masked)
Pretraining taskPredict masked-out tokensPredict the next token
Natural useClassification, tagging, retrievalGeneration
Can generate textNot naturallyYes

Masked language modelling is what makes an encoder bidirectional. Hide 15% of the tokens and train the model to recover them from both sides:

Text
input:  "the [MASK] sat on the mat"target: "cat"

Every representation is built from full left and right context. A decoder cannot do this — it must not see the future, or next-token prediction becomes trivial and the model learns nothing usable at generation time.

Using a pretrained model

The reason transformers dominate is not only the architecture. It is that pretraining on enormous unlabelled corpora is a one-time cost someone else has already paid, and adapting the result to your task takes a few thousand labelled examples.

Python
from transformers import AutoTokenizer, AutoModelForSequenceClassificationname = "bert-base-uncased"tokenizer = AutoTokenizer.from_pretrained(name)model = AutoModelForSequenceClassification.from_pretrained(name, num_labels=2)enc = tokenizer(    ["the film was not good at all", "an absolute triumph"],    padding=True, truncation=True, max_length=256, return_tensors="pt",)print(enc["input_ids"].shape)      # (2, seq_len)print(tokenizer.convert_ids_to_tokens(enc["input_ids"][0]))# ['[CLS]', 'the', 'film', 'was', 'not', 'good', 'at', 'all', '[SEP]', ...]

Two special tokens do specific jobs. [CLS] sits at the front and its final-layer representation is used as the sequence summary for classification. [SEP] marks the end, and separates the two halves in sentence-pair tasks.

Fine-tuning

Python
from torch.optim import AdamWfrom transformers import get_linear_schedule_with_warmupoptimizer = AdamW(model.parameters(), lr=2e-5, weight_decay=0.01)total = len(train_loader) * EPOCHSscheduler = get_linear_schedule_with_warmup(    optimizer, num_warmup_steps=int(0.1 * total), num_training_steps=total)model.train()for epoch in range(EPOCHS):    for batch in train_loader:        batch = {k: v.to(device) for k, v in batch.items()}        out = model(**batch)                 # labels in batch -> loss computed        out.loss.backward()        torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)        optimizer.step()        scheduler.step()        optimizer.zero_grad()

The hyperparameters here are not arbitrary and are the most common thing people get wrong.

SettingValueWhat happens otherwise
Learning rate2e-5 (range 1e-5 to 5e-5)At 1e-3 the pretrained weights are destroyed in the first few hundred steps; accuracy lands near chance
Epochs2–4Beyond that it memorises the training set
Warmup~10% of stepsLarge early updates on a fresh classifier head disrupt the encoder
Batch size16–32Larger needs a proportionally higher learning rate
Max length128–512Compute scales with n2n^2; 512 costs roughly 4× what 256 does

The learning rate is worth stating twice. Training a network from scratch at 2e-5 would take forever; fine-tuning at 1e-3 destroys everything that made the pretrained model valuable. If your fine-tuned model performs near chance, check the learning rate before anything else.

The comparison you actually have to make

BiLSTM + attentionFine-tuned BERT-base
Parameters~2–5 M110 M
Disk~20 MB~440 MB
Training (25k examples)10–20 min, GPU optional20–40 min, GPU required in practice
Inference, CPU, one document~5 ms~100 ms
Typical binary sentiment accuracy88–90%93–95%
Accuracy with 1,000 labels~78%~90%
Input length limitEffectively unbounded512 tokens, hard
Handles unseen wordsNo — <unk>Yes — subword tokenisation

The row that decides most real projects is not accuracy — it is accuracy with 1,000 labels. Pretraining means the model already knows English before it sees your data, so it needs far fewer labelled examples to reach a useful score. When labelling is expensive, that gap is the whole argument.

The 512-token limit is the sharpest practical constraint, because attention costs O(n2)O(n^2) and that quadratic is what caps the window. A 2,000-word document does not fit. Your options are to truncate (losing content), to chunk and pool over chunks (adding complexity), or to use a long-context variant. A BiLSTM has no such limit and remains genuinely competitive on long documents.

Self-attention's contribution is a single structural change: every position reaches every other in one step, and all positions compute at once. Everything else — multiple heads, positional encodings, residuals — is machinery that makes that one idea trainable at scale.

Choosing, and getting it running

Reach for a pretrained transformer when you have fewer than roughly 50,000 labelled examples, when nuance matters, when your text contains rare words or typos, or when a few points of accuracy are worth 20× the inference cost. Reach for a recurrent model when latency is measured in single-digit milliseconds, when you have no GPU, when documents run to thousands of tokens, or when you have millions of labelled in-domain examples and pretraining's advantage has evaporated.

Three habits will save you most of the debugging time.

Use the tokeniser that shipped with the checkpoint, and give it raw text. The model learned its embeddings against that exact vocabulary. Lowercasing before a case-sensitive checkpoint, or stripping punctuation the tokeniser expected, shifts every index away from what the model saw during pretraining. There is no error message — just accuracy that is inexplicably a few points low.

Start with the smallest model that could work. A distilled six-layer model is typically 60% faster and within one point of the full-size version on straightforward classification. Establish that the task needs more capacity before paying for it.

Sanity-check with a zero-shot pipeline before you fine-tune anything. Three lines tell you whether an off-the-shelf model already solves your problem:

Python
from transformers import pipelineclf = pipeline("sentiment-analysis")print(clf("the film was not good at all"))# [{'label': 'NEGATIVE', 'score': 0.9997}]

Sometimes it does, and the fine-tuning run you were about to launch was never necessary.