Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Walk through autoregressive text generation step by step, from first token to stop?


The generation loopprefill theprompt, fill KV cachelast-positionlogitssample withtemperature, top-pstream andappend the tokendecode one tokenover the cachesets time tofirst tokensets streaming speedExit on end-of-turn, a stop string, max_tokens or a full context window.
The model only turns tokens into one distribution; the loop, the sampler and the stop rules are serving code.

What you need to know

  1. Tokenize — the prompt, including the chat template's role markers, becomes ids of shape (1, T).
  2. Prefill — one forward pass over all T tokens builds the KV cache in every layer; keep only logits[:, -1, :].
  3. Choose a token — apply the decoding rule to that row.
  4. Emit — append the id to the sequence, and stream its text to the user.
  5. Decode — run the model on just the new token; its query attends to the cache; its K and V are appended; get the next logits.
  6. Check stop conditions — if none apply, go back to step 3.

The decoding rule

Python
import torchdef next_token(logits, temperature=0.8, top_p=0.9):    if temperature == 0:        return int(logits.argmax())                   # greedy    probs = torch.softmax(logits / temperature, dim=-1)    sorted_p, idx = probs.sort(descending=True)    keep = sorted_p.cumsum(-1) - sorted_p < top_p     # smallest set covering top_p    sorted_p = sorted_p * keep    choice = torch.multinomial(sorted_p / sorted_p.sum(), 1)    return int(idx[choice])torch.manual_seed(0)logits = torch.tensor([3.0, 2.5, 1.0, -1.0, -2.0])   # 5-token toy vocabularyprint(next_token(logits, temperature=0), [next_token(logits) for _ in range(8)])# 0 [1, 1, 1, 0, 0, 0, 0, 1]

Temperature below 1 sharpens the distribution; top-p keeps only the most likely tokens whose probabilities add up to p, and samples among them. Here only tokens 0 and 1 survive, so the low-probability tail can never be picked.

Stop conditions, in the order servers usually check them

  • The end-of-sequence or chat end-of-turn token is sampled.
  • A stop string set by the caller appears in the decoded text.
  • max_tokens is reached — the response is marked as cut off (for example finish_reason="length").
  • The context window is full.

Things people miss

  • State is only the tokens and the KV cache. Nothing else carries over between steps.
  • Same seed does not always mean same output in production. At temperature 0 on a single request, output is reproducible; on a busy server, requests are batched differently each time and GPU kernels may add numbers in a different order, so tiny differences can flip a token.
  • Streaming needs an incremental decoder. A token can be part of a multi-byte character; printing token by token can show broken characters in Hindi or emoji.
  • Speculative decoding changes step 5: a small draft model proposes several tokens and the main model verifies them all in one pass, keeping the main model's output distribution but taking fewer slow steps.

A real-life example

A chatbot serving a million users gets "मेरा ऑर्डर कहाँ है?" ("Where is my order?"). The server tokenizes the chat template plus 1,200 tokens of order history. Prefill takes about 80 ms and produces the first token. Then each decode step takes about 20 ms, because the GPU is shared with 60 other conversations in the same batch, and the answer streams at about 50 tokens per second.

Two bugs from their launch week illustrate the stop rules. The first: some answers ended mid-sentence, because max_tokens was 150 and the code ignored finish_reason. The second: Hindi answers sometimes showed "�" characters in the app, because the streaming code decoded each token separately, splitting Devanagari characters across tokens. They raised the limit and checked the finish reason, and switched to an incremental decoder that holds back bytes until a full character is formed.

Follow-up questions to expect

  • "Why is the first token slower than the rest?" — It waits for the whole prompt's prefill; later tokens need only one small decode step each.
  • "What does top-p do that top-k does not?" — Top-p adapts the number of candidates to how confident the model is; top-k always keeps exactly k.
  • "How does the model know to stop?" — It learned to produce an end-of-turn token during fine-tuning; the server stops when it samples that token or hits a limit.