Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Trace how vocab links tokenizer, embedding layer, and LM head into one system?
What you need to know
The loop
tokenizer.encode text -> ids in [0, V)wte (V, d) row i = input vector for id iblocks (B, T, d)lm_head (d, V) column i = output direction for id isoftmax + sample -> idtokenizer.decode id -> text (and the id is fed back in)What the vocabulary choice decides
The tokenizer decides how many tokens a text costs. More tokens mean more compute, more KV cache and less text per context window.
1import tiktoken2en = "The state government announced new flood relief measures for farmers on Monday."3hi = "राज्य सरकार ने सोमवार को किसानों के लिए बाढ़ राहत के नए उपायों की घोषणा की।"4for name in ["gpt2", "o200k_base"]:5 enc = tiktoken.get_encoding(name)6 print(name, enc.n_vocab, len(enc.encode(en)), len(enc.encode(hi)))7# gpt2 50257 13 1128# o200k_base 200019 13 20GPT-2's byte-level BPE was trained mostly on English. The Hindi sentence falls apart into 112 tokens — mostly single UTF-8 bytes. The newer 200K vocabulary (o200k_base) has learned Devanagari pieces and needs 20. English costs 13 tokens with both. A larger vocabulary means a bigger embedding and LM head, but far fewer tokens for many languages.
Keeping the contract
- Adding special tokens (for example chat markers) needs
model.resize_token_embeddings(len(tokenizer))in Hugging Face. The new rows start random, so they must be trained before they mean anything. - Padded vocabularies. Implementations often round V up to a multiple of 64 or 128 for GPU efficiency — nanoGPT uses 50,304 instead of 50,257. The extra ids are never produced by the tokenizer; some stacks set their logits to minus infinity so they can never be sampled.
- Mismatched tokenizers give no error if both vocabularies have the same size. The ids simply mean different things, and the output is plausible-looking nonsense.
A real-life example
An Indian-language news app started with a GPT-2-era model for Hindi summaries. With 112 tokens for a one-line sentence, a 1,024-token context held only a short paragraph of Hindi, and generation was slow because every syllable took several steps.
They move to a model whose tokenizer covers Indian scripts. The same Hindi article now costs about a fifth as many tokens, so it fits in the context, generates faster and costs less per article. During the switch, one service still loaded the old tokenizer with the new model. Nothing crashed — the new vocabulary was larger, so all ids were valid — but the summaries were garbled. They added a startup check that compares the tokenizer's vocabulary size and a few known token ids against the model's config.
Follow-up questions to expect
- "What happens if you add tokens but do not resize?" — The new ids are out of range for the embedding table, which raises an index error on lookup.
- "Why does a bigger vocabulary help multilingual models?" — More languages get whole-word or syllable tokens, so texts are shorter; the cost is a larger embedding and LM head.
- "Why can streaming output show broken characters?" — One Devanagari character is several bytes and may span several byte-level tokens; the decoder must buffer until a full character is formed.