Course Content
Natural Language Processing Basics
4 sections · 10 lessons
Introduction to Transformers for NLP
A recurrent model is reading this sentence, one token at a time:
"The report that the committee, after months of deliberation and three separate rounds of public consultation, finally published last Tuesday was widely condemned."To resolve was condemned the model needs report — 21 tokens back, counting punctuation. For that information to arrive, it must pass through 21 consecutive state updates, each one squeezing it through the same gates alongside everything else the model has read since. The signal survives, degraded, if it survives at all.
Count the steps another way. In a recurrent network, the path length between two positions i and j is ∣i−j∣. Information between distant words traverses a long chain, and every hop is a chance to be diluted.
There is a second problem, and commercially it is the bigger one. Step 24 cannot begin until step 23 has finished. A 500-token document requires 500 dependent matrix operations. A modern GPU can perform thousands of operations simultaneously and is left almost entirely idle, waiting. You cannot buy your way out of this — adding hardware does not shorten a chain of dependencies.
| Recurrent layer | Self-attention layer | |
|---|---|---|
| Path between any two positions | O(n) | O(1) |
| Sequential operations | O(n) | O(1) |
| Compute per layer | O(n⋅d2) | O(n2⋅d) |
| Parallel across positions | No | Yes |
Read the last row of the compute line honestly: attention is more expensive in raw arithmetic once n exceeds d. It wins anyway, because that arithmetic all happens at once on hardware built for exactly that, while recurrence is a queue.
Self-attention, before any notation
Here is the idea without a single symbol.
Every word looks at every other word in the sentence, decides how relevant each one is to itself, and builds a new representation of itself as a weighted blend of them all.
"The animal did not cross the street because it was too tired."When the model builds a representation for "it": animal 0.51 <- the antecedent tired 0.19 it 0.11 street 0.06 was 0.05 ...everything else shares the remaining 0.08The representation of it becomes mostly animal, with some tired mixed in. And crucially, animal is seven tokens away but reached in one step. There is no chain. Every position is one hop from every other position.
Change one word and the weights change with it:
"The animal did not cross the street because it was too wide." street 0.48 <- now the antecedent wide 0.22 it 0.10Nothing was hand-coded. The scoring function was learned, and it produces different attention for different contexts.
Queries, keys and values
How does a position decide what is relevant? Each token produces three vectors from its input embedding, using three learned matrices:
| Vector | Computed as | Its job |
|---|---|---|
| Query q | xWQ | What this position is looking for |
| Key k | xWK | What this position offers to others |
| Value v | xWV | The content this position contributes if attended to |
The database analogy is apt. You issue a query, match it against every key, and retrieve a blend of the values weighted by how well each key matched. The difference from a real database is that retrieval is soft — you get a weighted mixture, not a single row.
Read it as four steps:
- QK⊤ — dot every query against every key, giving an n×n matrix of raw relevance scores.
- ÷dk — scale down. This is not cosmetic; see below.
- softmax — normalise each row so weights are positive and sum to 1.
- ×V — take the weighted average of the value vectors.
Why divide by the square root of the dimension
Suppose q and k have independent components with mean 0 and variance 1. Their dot product is a sum of dk such products, so it has variance dk and standard deviation dk. With dk=64, typical scores land around ±8; with dk=512, around ±23.
Feed scores of that size into a softmax and see what happens:
scores [1.0, 2.0, 3.0] -> softmax [0.090, 0.245, 0.665]scores [8.0, 16.0, 24.0] -> softmax [0.0000001, 0.0003, 0.9997]The second distribution is effectively one-hot. Attention has become hard selection rather than a blend, and — the part that actually kills training — the softmax gradient at a saturated point is nearly zero. The layer stops learning.
Dividing by dk returns the scores to unit variance regardless of dimension, keeping the softmax in a range where it produces a meaningful distribution and a usable gradient.
A worked example
Three tokens, dk=2, arithmetic small enough to check by hand.
Q = [[1, 0], K = [[1, 0], V = [[1, 0], [0, 1], [0, 1], [0, 2], [1, 1]] [1, 1]] [1, 1]]Scores = Q K^T: row 1: [1*1+0*0, 1*0+0*1, 1*1+0*1] = [1, 0, 1] row 2: [0, 1, 1] row 3: [1, 1, 2]Scaled by sqrt(2) = 1.414: row 1: [0.707, 0.000, 0.707] row 3: [0.707, 0.707, 1.414]Softmax of row 1: exp: [2.028, 1.000, 2.028] sum = 5.056 weights: [0.401, 0.198, 0.401]Output for position 1 = 0.401*[1,0] + 0.198*[0,2] + 0.401*[1,1] = [0.401, 0] + [0, 0.396] + [0.401, 0.401] = [0.802, 0.797]Position 1 attends to positions 1 and 3 equally (both score 1) and to position 2 less (score 0), and its output is the corresponding blend of value vectors. That is the entire mechanism.
Multiple heads
One attention layer produces one set of weights per position — one notion of relevance. But relevance is not one thing. Resolving a pronoun, tracking the subject of a verb, and noticing a negation are different relationships, and a single weighting cannot express all three at once.
Multi-head attention runs h attention operations in parallel, each with its own WQ,WK,WV, each on a slice of the dimensions, then concatenates the results and projects them back.
With dmodel=512 and h=8, each head works in 64 dimensions, so the total compute matches single-head attention at full width — you get eight relationship types for the price of one.
Probing trained models finds heads that specialise in recognisable ways: some attend to the immediately preceding token, some link verbs to their subjects, some connect pronouns to antecedents, and a good many appear to do nothing useful and can be pruned with little loss.
Attention has no idea what order things are in
Here is a property that surprises people. Shuffle the input tokens and self-attention produces the same outputs, merely shuffled to match. Every position attends to every other with no notion of near or far, left or right.
"dog bites man" and "man bites dog"-> identical sets of attention outputs, just reorderedA recurrent network gets order for free from the order in which it reads. Attention has to be told.
The fix is to add a position-dependent vector to each input embedding before the first layer. The original formulation uses sinusoids of different frequencies:
1import numpy as np23def positional_encoding(max_len, d_model):4 pos = np.arange(max_len)[:, None]5 i = np.arange(d_model)[None, :]6 angle = pos / np.power(10000, (2 * (i // 2)) / d_model)7 pe = np.zeros((max_len, d_model))8 pe[:, 0::2] = np.sin(angle[:, 0::2])9 pe[:, 1::2] = np.cos(angle[:, 1::2])10 return peLow-frequency dimensions change slowly and encode coarse position; high-frequency dimensions distinguish neighbours. Because sin(a+b) expands into terms involving sina, cosa, sinb and cosb, a fixed offset corresponds to a linear transformation of the encoding — which gives the model a route to learning relative position, not just absolute.
BERT and GPT-2 instead learn a position embedding table directly. It is simpler, works at least as well within the trained length, and cannot extrapolate beyond it. Most current large language models use a third option, rotary position embeddings (RoPE), which rotate the query and key vectors by a position-dependent angle so that attention scores depend on the relative distance between tokens.
The block
A transformer layer is attention plus a feed-forward network, each wrapped in a residual connection and layer normalisation.
x ──┬──> MultiHeadAttention ──> (+) ──> LayerNorm ──┬──> FeedForward ──> (+) ──> LayerNorm ──> └───────────────────────────┘ └──────────────────────┘ residual residual1import torch.nn as nn23class TransformerBlock(nn.Module):4 def __init__(self, d_model, n_heads, d_ff, dropout=0.1):5 super().__init__()6 self.attn = nn.MultiheadAttention(d_model, n_heads,7 dropout=dropout, batch_first=True)8 self.norm1 = nn.LayerNorm(d_model)9 self.norm2 = nn.LayerNorm(d_model)10 self.ff = nn.Sequential(11 nn.Linear(d_model, d_ff),12 nn.GELU(),13 nn.Dropout(dropout),14 nn.Linear(d_ff, d_model),15 )16 self.dropout = nn.Dropout(dropout)1718 def forward(self, x, key_padding_mask=None):19 a, _ = self.attn(x, x, x, key_padding_mask=key_padding_mask)20 x = self.norm1(x + self.dropout(a))21 x = self.norm2(x + self.dropout(self.ff(x)))22 return xEach piece earns its place. Residual connections give the gradient a direct path to every layer, which is what makes 12-, 24- and 96-layer stacks trainable. Layer normalisation stabilises the scale of activations, and unlike batch normalisation it does not depend on the batch, so it behaves identically at batch size 1. The feed-forward network — usually four times as wide as dmodel — is applied to each position independently and holds a large share of the model's parameters; attention mixes information between positions, the feed-forward layer processes it.
This is the original post-norm arrangement, with LayerNorm after each residual addition. Most current models use pre-norm instead — normalise the input to each sub-layer, add the residual afterwards, and put one final LayerNorm after the last block — because deep pre-norm stacks train more stably. In PyTorch's built-in nn.TransformerEncoderLayer that is norm_first=True.
key_padding_mask is the practical detail that gets forgotten. Without it, padding positions receive attention weight and contribute to every real token's representation. There is no error — only a model that quietly performs worse on short sequences in a batch of long ones.
Encoders, decoders, and why the distinction matters
The same block appears in two configurations, and they are suited to different tasks.
| Encoder (BERT-style) | Decoder (GPT-style) | |
|---|---|---|
| Attention range | All positions, both directions | Only positions to the left (causally masked) |
| Pretraining task | Predict masked-out tokens | Predict the next token |
| Natural use | Classification, tagging, retrieval | Generation |
| Can generate text | Not naturally | Yes |
Masked language modelling is what makes an encoder bidirectional. Hide 15% of the tokens and train the model to recover them from both sides:
input: "the [MASK] sat on the mat"target: "cat"Every representation is built from full left and right context. A decoder cannot do this — it must not see the future, or next-token prediction becomes trivial and the model learns nothing usable at generation time.
Using a pretrained model
The reason transformers dominate is not only the architecture. It is that pretraining on enormous unlabelled corpora is a one-time cost someone else has already paid, and adapting the result to your task takes a few thousand labelled examples.
1from transformers import AutoTokenizer, AutoModelForSequenceClassification23name = "bert-base-uncased"4tokenizer = AutoTokenizer.from_pretrained(name)5model = AutoModelForSequenceClassification.from_pretrained(name, num_labels=2)67enc = tokenizer(8 ["the film was not good at all", "an absolute triumph"],9 padding=True, truncation=True, max_length=256, return_tensors="pt",10)11print(enc["input_ids"].shape) # (2, seq_len)12print(tokenizer.convert_ids_to_tokens(enc["input_ids"][0]))13# ['[CLS]', 'the', 'film', 'was', 'not', 'good', 'at', 'all', '[SEP]', ...]Two special tokens do specific jobs. [CLS] sits at the front and its final-layer representation is used as the sequence summary for classification. [SEP] marks the end, and separates the two halves in sentence-pair tasks.
Fine-tuning
1from torch.optim import AdamW2from transformers import get_linear_schedule_with_warmup34optimizer = AdamW(model.parameters(), lr=2e-5, weight_decay=0.01)5total = len(train_loader) * EPOCHS6scheduler = get_linear_schedule_with_warmup(7 optimizer, num_warmup_steps=int(0.1 * total), num_training_steps=total8)910model.train()11for epoch in range(EPOCHS):12 for batch in train_loader:13 batch = {k: v.to(device) for k, v in batch.items()}14 out = model(**batch) # labels in batch -> loss computed15 out.loss.backward()16 torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)17 optimizer.step()18 scheduler.step()19 optimizer.zero_grad()The hyperparameters here are not arbitrary and are the most common thing people get wrong.
| Setting | Value | What happens otherwise |
|---|---|---|
| Learning rate | 2e-5 (range 1e-5 to 5e-5) | At 1e-3 the pretrained weights are destroyed in the first few hundred steps; accuracy lands near chance |
| Epochs | 2–4 | Beyond that it memorises the training set |
| Warmup | ~10% of steps | Large early updates on a fresh classifier head disrupt the encoder |
| Batch size | 16–32 | Larger needs a proportionally higher learning rate |
| Max length | 128–512 | Compute scales with n2; 512 costs roughly 4× what 256 does |
The learning rate is worth stating twice. Training a network from scratch at 2e-5 would take forever; fine-tuning at 1e-3 destroys everything that made the pretrained model valuable. If your fine-tuned model performs near chance, check the learning rate before anything else.
The comparison you actually have to make
| BiLSTM + attention | Fine-tuned BERT-base | |
|---|---|---|
| Parameters | ~2–5 M | 110 M |
| Disk | ~20 MB | ~440 MB |
| Training (25k examples) | 10–20 min, GPU optional | 20–40 min, GPU required in practice |
| Inference, CPU, one document | ~5 ms | ~100 ms |
| Typical binary sentiment accuracy | 88–90% | 93–95% |
| Accuracy with 1,000 labels | ~78% | ~90% |
| Input length limit | Effectively unbounded | 512 tokens, hard |
| Handles unseen words | No — <unk> | Yes — subword tokenisation |
The row that decides most real projects is not accuracy — it is accuracy with 1,000 labels. Pretraining means the model already knows English before it sees your data, so it needs far fewer labelled examples to reach a useful score. When labelling is expensive, that gap is the whole argument.
The 512-token limit is the sharpest practical constraint, because attention costs O(n2) and that quadratic is what caps the window. A 2,000-word document does not fit. Your options are to truncate (losing content), to chunk and pool over chunks (adding complexity), or to use a long-context variant. A BiLSTM has no such limit and remains genuinely competitive on long documents.
Self-attention's contribution is a single structural change: every position reaches every other in one step, and all positions compute at once. Everything else — multiple heads, positional encodings, residuals — is machinery that makes that one idea trainable at scale.
Choosing, and getting it running
Reach for a pretrained transformer when you have fewer than roughly 50,000 labelled examples, when nuance matters, when your text contains rare words or typos, or when a few points of accuracy are worth 20× the inference cost. Reach for a recurrent model when latency is measured in single-digit milliseconds, when you have no GPU, when documents run to thousands of tokens, or when you have millions of labelled in-domain examples and pretraining's advantage has evaporated.
Three habits will save you most of the debugging time.
Use the tokeniser that shipped with the checkpoint, and give it raw text. The model learned its embeddings against that exact vocabulary. Lowercasing before a case-sensitive checkpoint, or stripping punctuation the tokeniser expected, shifts every index away from what the model saw during pretraining. There is no error message — just accuracy that is inexplicably a few points low.
Start with the smallest model that could work. A distilled six-layer model is typically 60% faster and within one point of the full-size version on straightforward classification. Establish that the task needs more capacity before paying for it.
Sanity-check with a zero-shot pipeline before you fine-tune anything. Three lines tell you whether an off-the-shelf model already solves your problem:
1from transformers import pipeline2clf = pipeline("sentiment-analysis")3print(clf("the film was not good at all"))4# [{'label': 'NEGATIVE', 'score': 0.9997}]Sometimes it does, and the fine-tuning run you were about to launch was never necessary.