Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Trace how vocab links tokenizer, embedding layer, and LM head into one system?


Tokens for one news sentence, two vocabularies131121320EnglishHindigpt2 (50,257)o200k (200,019)Measured with tiktoken on the lesson's sentence pair.
The vocabulary decides what a language costs: the same Hindi sentence is over five times longer to a GPT-2 tokenizer.

What you need to know

The loop

Text
tokenizer.encode   text -> ids in [0, V)wte                (V, d)   row i  = input vector for id iblocks             (B, T, d)lm_head            (d, V)   column i = output direction for id isoftmax + sample   -> idtokenizer.decode   id -> text   (and the id is fed back in)

What the vocabulary choice decides

The tokenizer decides how many tokens a text costs. More tokens mean more compute, more KV cache and less text per context window.

Python
import tiktokenen = "The state government announced new flood relief measures for farmers on Monday."hi = "राज्य सरकार ने सोमवार को किसानों के लिए बाढ़ राहत के नए उपायों की घोषणा की।"for name in ["gpt2", "o200k_base"]:    enc = tiktoken.get_encoding(name)    print(name, enc.n_vocab, len(enc.encode(en)), len(enc.encode(hi)))# gpt2       50257  13 112# o200k_base 200019 13 20

GPT-2's byte-level BPE was trained mostly on English. The Hindi sentence falls apart into 112 tokens — mostly single UTF-8 bytes. The newer 200K vocabulary (o200k_base) has learned Devanagari pieces and needs 20. English costs 13 tokens with both. A larger vocabulary means a bigger embedding and LM head, but far fewer tokens for many languages.

Keeping the contract

  • Adding special tokens (for example chat markers) needs model.resize_token_embeddings(len(tokenizer)) in Hugging Face. The new rows start random, so they must be trained before they mean anything.
  • Padded vocabularies. Implementations often round V up to a multiple of 64 or 128 for GPU efficiency — nanoGPT uses 50,304 instead of 50,257. The extra ids are never produced by the tokenizer; some stacks set their logits to minus infinity so they can never be sampled.
  • Mismatched tokenizers give no error if both vocabularies have the same size. The ids simply mean different things, and the output is plausible-looking nonsense.

A real-life example

An Indian-language news app started with a GPT-2-era model for Hindi summaries. With 112 tokens for a one-line sentence, a 1,024-token context held only a short paragraph of Hindi, and generation was slow because every syllable took several steps.

They move to a model whose tokenizer covers Indian scripts. The same Hindi article now costs about a fifth as many tokens, so it fits in the context, generates faster and costs less per article. During the switch, one service still loaded the old tokenizer with the new model. Nothing crashed — the new vocabulary was larger, so all ids were valid — but the summaries were garbled. They added a startup check that compares the tokenizer's vocabulary size and a few known token ids against the model's config.

Follow-up questions to expect

  • "What happens if you add tokens but do not resize?" — The new ids are out of range for the embedding table, which raises an index error on lookup.
  • "Why does a bigger vocabulary help multilingual models?" — More languages get whole-word or syllable tokens, so texts are shorter; the cost is a larger embedding and LM head.
  • "Why can streaming output show broken characters?" — One Devanagari character is several bytes and may span several byte-level tokens; the decoder must buffer until a full character is formed.