How an LLM actually works: tokens, attention, sampling, context

JR

Jai Rao

August 22, 202621 min read

A mechanism-level walk through a language model: subword tokenisation and the bugs it causes, attention, temperature and top-p, the three training stages, and what a long context costs.


Ask a language model a question and a paragraph comes back. What actually happened is a loop that ran a few hundred times. The model looked at everything in front of it, produced a score for every entry in its vocabulary, something picked one entry, that entry got appended to the input, and the whole thing ran again. There is no plan held in reserve and no draft being revised. The fluency is what you get from doing one small step competently, several hundred times in a row.

Almost every surprising behaviour of these systems falls out of the mechanics of that loop: the letter-counting failures, the arithmetic that collapses on long numbers, the way quality sags in the middle of a very long document, the fact that asking for intermediate steps genuinely produces better answers. This is a walk through the loop, from the moment your text is cut into pieces to the moment a token is chosen, with the consequences named as we go.

What the tokeniser decides before the model sees anything

A model never sees characters. Text first passes through a tokeniser that maps it to a sequence of integers drawn from a fixed vocabulary, usually somewhere in the range of 30,000 to 200,000 entries. Those entries are not words. They are fragments,, and that fragmentation reaches all the way to the output.

Words look like the obvious unit, and they fail badly. English has a long tail: you would need an enormous vocabulary and would still meet a typo, a surname, a product code, or a Python identifier that has no entry at all. Characters are the opposite trade: nothing is ever unrepresentable, but sequences get four or five times longer, and every position you add costs compute in every layer. Subword tokenisation sits in the middle. Frequent words get a single entry, rare words get assembled out of pieces, and nothing is unrepresentable.

Byte pair encoding is the usual way to build that vocabulary. Start with individual characters, count every adjacent pair across the corpus, merge the most frequent pair into a new symbol, and repeat, recording the merges in order. Encoding new text means replaying those merges. Here is the whole idea in twenty lines, on a corpus of six words.

Text
from collections import Counter# "_" marks the start of a word, so "low" inside "slow" is a different# symbol from "low" at the start of a word.words = [list("_" + w) for w in "low lower lowest slow slower slowest".split()]def merge(seq, a, b):    out, i = [], 0    while i < len(seq):        if i + 1 < len(seq) and seq[i] == a and seq[i + 1] == b:            out.append(a + b); i += 2        else:            out.append(seq[i]); i += 1    return outfor _ in range(5):    pairs = Counter(p for w in words for p in zip(w, w[1:]))    (a, b), n = pairs.most_common(1)[0]    print(f"merge {a!r}+{b!r} (seen {n}x) -> {a + b!r}")    words = [merge(w, a, b) for w in words]print(words[0], words[3], words[5])

The merges come out as 'l'+'o', then 'lo'+'w', then 'low'+'e', and the final line prints ['_', 'low'] ['_s', 'low'] ['_s', 'lowe', 's', 't']. Frequency bought low its own symbol; slowest is still four fragments. That asymmetry is the whole story of tokenisation at scale. Production tokenisers run this over raw bytes rather than characters, starting from a base vocabulary of 256 so that any byte sequence in any script can be encoded, and they learn hundreds of thousands of merges from a corpus that is mostly English web text.

One detail that bites people constantly: whitespace is usually attached to the token that follows it. " the", "the" and "The" are three distinct vocabulary entries with three unrelated integer IDs. A stray leading space changes the token sequence, which changes the prediction, which is why an otherwise identical prompt can behave differently after an innocuous edit to your string formatting.

The bugs you inherit from the tokeniser

Take the famous one. Asked how many times the letter r appears in "strawberry", models have historically got it wrong. The reason is not a failure of intelligence, it is a failure of access. That word arrives as a handful of opaque integers, plausibly something like str + aw + berry. The letters are not there to be counted any more than the individual grooves are there when you look at a record sleeve. A model can answer correctly by having read enough discussion of spelling to memorise the answer, and current models often do, but that is recall rather than inspection, and it fails the moment you ask about a word nobody has written about.

The general rule: any task that requires looking at characters is working against the representation. Reversing a string, counting occurrences, finding the seventh letter, constructing an acrostic, judging whether two words rhyme, applying a cipher. None of these are hard problems. They are problems phrased in a unit the model does not have.

Long numbers and unlucky splits

Arithmetic breaks for a related reason with an extra twist. Digits get grouped by the tokeniser, commonly into chunks of up to three, and the grouping depends on where the boundaries happen to land. 1000 might be one token while 1001 is two, and the same digit can sit in a differently sized chunk in two numbers you are trying to add. Column arithmetic works because the ones column lines up under the ones column. If the operands are chopped into mismatched groups, that alignment is simply not visible.

Then there is the compute ceiling. Every token gets exactly one forward pass through the network: a fixed number of layers, a fixed width, no loops. A twelve-digit multiplication needs more sequential steps than that, and there is nowhere to keep a carry. This is why writing the working out helps, and why handing the problem to a calculator helps far more.

The third consequence is economic. Byte-level BPE trained mostly on English gives common English words their own single tokens. A character in Devanagari, Thai, or Chinese is three bytes in UTF-8, and if the merge table never learned those byte sequences well, a single character can cost two or three tokens. The same sentence, same meaning, can cost several times more tokens in Hindi than in English. That means a higher bill, a context window that fills faster, and less room to work in for the same nominal limit — an inequity that lives entirely in a lookup table built before training started.

Every token becomes a vector, and every vector looks backwards

Once you have integers, the first layer is a lookup. The embedding matrix has one row per vocabulary entry and one column per model dimension, and the token ID is a row index. That is genuinely all it is: token 8,242 means "fetch row 8,242". Those rows are learned during training, and they end up arranged so that tokens used in similar contexts land near each other, which is why a model handles a synonym it has never seen in your exact sentence.

A pile of independent row vectors is not yet language. Two things happen next. Position information is added, either as a learned per-position vector or, in most current models, by rotating the query and key vectors by an angle that depends on position. And then attention lets each position look at the positions before it.

Query, key and value without the linear algebra

Each position produces three different projections of its own vector. The query is what this position is looking for. The key is what this position offers to anyone looking. The value is what it contributes if it gets picked. The score between two positions is the dot product of one's query with the other's key, divided by the square root of the dimension to keep the numbers in a sane range. Those scores go through a softmax across all earlier positions, and the output is the weighted average of their values.

Concretely: "Priya poured the water into the bottle until it was full." At the position of it, the model has to decide whether it refers to the water or the bottle, and the answer changes what the sentence means. Below, the four dimensions are hand-labelled so you can read the arithmetic, and the query is built to look for something that can be full.

Text
import mathdef softmax(xs):    m = max(xs); e = [math.exp(x - m) for x in xs]    return [v / sum(e) for v in e]# Four made-up dimensions: [person, liquid, container, action].key = {"Priya": [1, 0, 0, 0], "poured": [0, 0, 0, 1], "the": [0, 0, 0, 0],       "water": [0, 1.0, 0.1, 0], "into": [0, 0, 0, 0.2],       "bottle": [0, 0.1, 1.0, 0], "until": [0, 0, 0, 0]}sentence = ["Priya", "poured", "the", "water", "into", "the", "bottle", "until"]# The query the model builds at "it": find a thing that can be full.query, d = [0, 1.6, 5.0, 0], 4scores = [sum(q * k for q, k in zip(query, key[w])) / math.sqrt(d) for w in sentence]for i, (w, a) in enumerate(zip(sentence, softmax(scores))):    print(f"{i}  {w:>7} score={scores[i]:5.2f}  attention={a:.3f}")

bottle takes 0.598 of the attention, water takes 0.130, and the six positions that offer nothing relevant split the remainder at 0.045 each. The output vector at it is now mostly the bottle's value vector, so downstream layers work with a representation of it that carries the bottle's properties. That is coreference resolution, and it is a dot product and a weighted average.

Two honest caveats. The labelled dimensions are a teaching device; in a real model the query, key and value projections are learned matrices and the individual dimensions are not interpretable. And a layer does not run one of these, it runs many in parallel as separate heads with separate projections, so one head can be tracking subject-verb agreement while another tracks the referent. Attention is also masked so a position can only see positions before it, which is exactly what lets the model be trained on every position of a document at once.

What the stack of layers builds up

A useful way to picture the rest is a residual stream: each position carries a running vector, and every layer reads from it, computes something, and adds the result back. Attention moves information between positions. The feed-forward block that follows transforms each position's vector in place, and it holds the bulk of the parameters. Then the next layer does it again, thirty to a hundred times depending on the model.

What changes as you go up is easy to overclaim, but the broad finding from probing work is consistent enough to be worth knowing. Early layers are dominated by local and surface properties: which token this is, its part of speech, immediate syntax. Middle layers are where the abstract work concentrates, including tracking entities across a passage and pulling in the facts that turn out to be relevant. Later layers turn that back into something concrete, because the network's last job is a narrow one.

That last job is one matrix multiply. The final vector for the last position is multiplied by an unembedding matrix with one row per vocabulary entry, producing one raw score per entry. Those scores are the logits. There are as many as there are vocabulary entries, most of them nonsense in context, and the model has committed to nothing yet. Everything it knows about what comes next is now a list of numbers.

From logits to one chosen token

Logits are unbounded real numbers, so they get pushed through a softmax to become a probability distribution. Temperature is a divisor applied to the logits before that softmax. Divide by a small number and the differences between logits get amplified, so the distribution sharpens. Divide by a large number and everything flattens. This code takes six plausible next tokens after "She picked up" and shows what temperature does to them.

Text
import math# Raw scores the final layer produced for six candidate next tokens.vocab  = ["the", "a", "my", "her", "some", "aardvark"]logits = [3.4,  2.9,  1.8,  1.6,  0.4,  -4.2]def softmax(xs, temperature=1.0):    scaled = [x / temperature for x in xs]    m = max(scaled)                      # subtract the max for stability    exps = [math.exp(x - m) for x in scaled]    return [e / sum(exps) for e in exps]for t in (0.2, 1.0, 1.8):    probs = softmax(logits, t)    print(f"T={t}: " + "  ".join(f"{w}={p:.3f}" for w, p in zip(vocab, probs)))def top_p(probs, p=0.9):    order = sorted(range(len(probs)), key=lambda i: -probs[i])    kept, run = [], 0.0    for i in order:        kept.append(i); run += probs[i]        if run >= p: break    z = sum(probs[i] for i in kept)    return {vocab[i]: round(probs[i] / z, 3) for i in kept}print("top_p(0.9) at T=1.0:", top_p(softmax(logits, 1.0), 0.9))

At T=0.2 the top token holds 0.924 and everything below second place is effectively dead. At T=1.0 it holds 0.494 with a live tail. At T=1.8 the leader is down to 0.365 and even aardvark becomes reachable at 0.005, which is how high temperature produces sentences that fall apart. Note what temperature does not do: it never reorders the candidates. It only decides how much the order matters.

Top-p, or nucleus sampling, does something different. Sort the candidates, keep the shortest prefix whose probabilities sum to p, renormalise, and sample from that. The last line of the output keeps four tokens and drops the rest. The reason to prefer this over top-k is that it adapts to the shape of the distribution: where the model is confident, top-p keeps one or two options; where it is genuinely torn, it keeps many. Top-k keeps the same fixed count either way, offering forty alternatives to a token the model is 99% sure about.

Which raises the question most explanations skip. If the model has a best guess, why not always take it? Three reasons.

  • Greedy is locally optimal, not globally. The highest-probability token now can walk you into a region where every continuation is poor. Nothing in the loop backtracks; the token is already in the context and the model will build on it.
  • Greedy decoding degenerates on open-ended text. It falls into repetition loops, because once a phrase has appeared twice the most probable continuation is that phrase a third time. Sampling is what breaks the cycle.
  • It is not as deterministic as it looks. Requests at temperature 0 are only approximately reproducible on production infrastructure, because batching changes the order of floating-point reductions and the top two logits are sometimes within that noise.

The practical split is clean enough. Extraction, classification, structured output, and code with one correct form all want a low temperature. Anything where several outputs would be acceptable produces better text with real sampling, and forcing greedy decoding there makes the prose flat and the phrasing repetitive.

Three training stages, three different things changed

The loop described so far is inference. It is worth being precise about which stage of training put which behaviour there, because it explains a lot about what you can and cannot expect to fix.

Pretraining is next-token prediction over an enormous corpus. Take a document, hide everything after position n, ask for the token at n+1, compare the predicted distribution against the actual token with cross-entropy loss, update the weights. No human labels are needed because the text is its own supervision, which is the entire reason this scaled. Essentially all the capability is built here: grammar, world knowledge, code, translation, the ability to continue a pattern. It also burns the overwhelming majority of the compute. What you have at the end is not an assistant. It is a document continuer, and if you hand it a question it may well produce three more questions, because that is what a page containing one question tends to look like.

Supervised fine-tuning continues training on a curated set of prompt-and-good-response pairs, with the same loss function on a comparatively tiny amount of data. What changes here is mostly form, not knowledge: the model learns that the continuation of a question is an answer, in a certain register, at a certain length, with a certain structure. It is the cheapest stage and often the highest-leverage one, and it is a poor tool for installing new facts. If a capability was not built during pretraining, fine-tuning rarely conjures it.

Preference-based alignment works on comparisons rather than examples. Show a human two responses to the same prompt and record which one they prefer. In RLHF, those comparisons train a reward model that scores responses, and then the language model is optimised to score well under it using reinforcement learning, with a penalty term that keeps it from drifting too far from the fine-tuned model. DPO reaches a closely related objective without the separate reward model, deriving a loss you can apply directly to the preference pairs, which makes it dramatically simpler to run. Either way, what this stage changes is the ranking among outputs the model could already produce. It shifts probability mass toward responses that raters liked.

That framing predicts the failure modes, and they are real. Optimising for rater approval produces sycophancy: a model that concedes when a confident user pushes back on a correct answer, because agreement scored well. It produces length inflation, because raters reward thorough-looking answers. It produces hedging on questions that deserve a straight answer. And pushing hard on it can narrow the output distribution enough to hurt tasks where variety is the point. None of that is a bug in the reward model. It is an accurate reflection of what humans clicked.

What a long context actually costs

Two distinct costs hide behind a context window number, and they behave differently.

The first is the KV cache. When generating one token at a time, the keys and values computed for earlier positions never change, because attention only looks backwards. So you compute them once and keep them. Without this, producing token 500 would mean re-running the entire prefix through every layer. With it, processing the prompt is one large parallel pass (prefill), and each generated token is a small step against a cache that grows by one entry per layer per head. The cache is pure memory, and it grows linearly.

Text
def kv_cache_bytes(tokens, layers=32, kv_heads=8, head_dim=128, bytes_per_value=2):    # One key vector and one value vector per layer, per KV head, per token.    return 2 * tokens * layers * kv_heads * head_dim * bytes_per_valueprint(f"per token: {kv_cache_bytes(1) / 1024:.0f} KiB")for n in (1_000, 32_000, 128_000):    print(f"{n:>7} tokens: {kv_cache_bytes(n) / 2**30:6.2f} GiB")# Attention score pairs during prefill: this is the quadratic term.for n in (1_000, 2_000, 4_000):    print(f"{n:>7} tokens: {n * (n + 1) // 2:,} score pairs per head per layer")

For that configuration the cache costs 128 KiB per token, so a 32,000-token conversation holds 3.91 GiB and a 128,000-token one holds 15.62 GiB, on top of the weights and before any other session. This is why long-context requests are priced differently and why a server's throughput falls when many users each hold a long history: the binding constraint is memory that cannot be reclaimed until the session ends.

The second cost is the quadratic term in the second loop. Prefill computes a score for every pair of positions, so 1,000 tokens is about 500,000 pairs per head per layer and 4,000 tokens is about 8,000,000. Double the prompt and quadruple that work. FlashAttention and its relatives are genuine wins, but they reduce memory traffic and improve arithmetic intensity; they do not reduce the number of pairs. Once you are generating, the per-token cost is linear in context length rather than quadratic, since one new query attends to all cached keys, which is why a long conversation gets steadily slower per token rather than suddenly falling over.

Then there is the part that surprises people who paid for the big window. Retrieval accuracy from a long context depends on where in the context the needed information sits, and the consistent finding is a U shape: strong at the beginning, strong at the end, weakest in the middle. Two mechanisms plausibly contribute. A softmax over tens of thousands of positions has a fixed budget of attention to spread, so no single distant position gets much of it. And the training distribution contains very few examples where the crucial fact was 80,000 tokens back, so the ability to use that range is undertrained relative to the advertised limit. A 200,000-token window is a hard ceiling on what fits, not a promise that all of it is equally usable.

Why writing out steps makes answers better

Return to the compute ceiling, because it is the key to the whole "reasoning" story. One forward pass is a fixed amount of computation: a fixed depth, a fixed width, no loops, no recursion. Any problem requiring more sequential steps than the network has layers cannot be solved inside a single token.

Intermediate tokens are the way out. Once a token is emitted it sits in the context, and every subsequent position can attend to it. The model has moved its working memory onto the page, where it is durable and re-readable. Ten tokens of visible arithmetic buys ten more forward passes and ten more places to park a partial result. That is the mechanism, and it explains the shape of the benefit precisely: writing out steps reliably helps on multi-step problems and does nothing at all for a single lookup like the capital of a country, because there was never a serial-depth problem to solve.

Reasoning-tuned models industrialise this. Post-training rewards long traces that arrive at verifiably correct answers, so the model learns to spend hundreds or thousands of tokens enumerating cases, checking its own work, and abandoning dead ends before it commits to a final answer. You are buying accuracy with test-time compute, and on problems with enough structure the exchange rate is favourable.

What it is not is introspection. The trace comes out of the same next-token machinery as the answer, sampled from the same distribution. It is not a readout of a separate hidden process. Faithfulness studies find models that produce a coherent-looking chain and then state an answer inconsistent with it, and models measurably swayed by a cue in the prompt that never appears anywhere in their stated reasoning. Treat a trace as scratch work you are free to audit, not as testimony about why the model answered as it did. It also inherits a specific weakness from the loop: a wrong step early on tends to get built upon rather than corrected, because the model's job at every position is to produce a plausible continuation of what is already written, and a plausible continuation of a wrong step is usually the next wrong step.

Reading a failure back to the mechanism that caused it

The point of all this is diagnostic. When a model does something odd, the mechanism usually tells you whether the fix is available to you at all.

What you observeWhere it comes fromWhat actually changes it
Miscounts letters, botches a reversal, invents a rhymeCharacters are not visible inside a tokenPass the string with separators, or do it in code
Arithmetic on long numbers is confidently wrongDigit chunking plus fixed compute per tokenA tool call; a calculator is exact and cheaper
The same content costs far more in a non-English scriptMerge table learned mostly on English bytesNothing at your layer; budget for it
Ignores a constraint stated in the middle of a long inputPosition-dependent retrieval over long contextsShorter inputs; constraints at an edge, not buried
Same prompt, different answer each timeSampling from a distribution, not lookupLower temperature and constrained output formats
Prose is repetitive and oddly flatGreedy or near-greedy decodingRaise temperature; use top-p rather than top-k
Reverses a correct answer when you push backPreference tuning rewarded agreeablenessRecognise it as a training artefact, not a rethink
Long conversations get slower and pricier per turnKV cache growth and linear per-token attentionTrim history; start fresh sessions deliberately

You have the mental model working when you can look at a bad output and predict which category it lands in before you start experimenting. A tokenisation failure will not be fixed by rewording, however many attempts you make. A sampling artefact will not be fixed by a longer instruction. A middle-of-context miss is a signal to shorten the input, not to insist harder.

One closing caution, because the confident version of this material is usually overselling. Everything above is architecture and arithmetic, and it is genuinely well understood: you could implement all of it. Why a specific set of billions of weights produces a specific behaviour is not well understood, and interpretability research is still early. Knowing the machinery tells you the shape of the failures to expect. It does not let you predict what a given model will say next, and anyone claiming otherwise is selling something.