Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is tokenization, and why is it important in LLMs?
What you need to know
Why not words or characters?
- Whole words need a huge vocabulary and still meet unknown words (new product names, typos).
- Characters never meet an unknown word, but sequences get 4–5 times longer, and attention cost grows with length.
- Subwords sit in the middle: a vocabulary of about 30,000 to 200,000 pieces covers any text. BERT used about 30,000 WordPiece tokens; GPT-4o's tokenizer and Llama 3 use vocabularies of roughly 200,000 and 128,000.
How byte-pair encoding (BPE) learns its pieces
- Start small — the vocabulary is every single byte or character.
- Count pairs — find the most frequent adjacent pair in the training text, say "e" + "r".
- Merge — add "er" as a new token and rewrite the text with it.
- Repeat — keep merging ("er" + " " , "low" + "er" ...) until the vocabulary reaches its target size.
At the end, frequent strings such as " the" or " account" are single tokens, and rare strings are built from smaller pieces. WordPiece and SentencePiece's unigram method differ in how they choose pieces, but the idea is the same.
Why engineers care
- Cost — APIs charge per input and output token.
- Context — the window is a token budget; more tokens per sentence means less room for documents.
- Latency — each output token is one forward pass, so a longer answer is a slower answer.
- Language fairness — tokenizers are trained mostly on English, so other scripts split into more pieces and cost more for the same message.
- Odd failures — "How many r's in strawberry?" is hard because "strawberry" arrives as two or three chunks, not ten letters.
You can measure this directly:
1import tiktoken23enc = tiktoken.get_encoding("o200k_base") # GPT-4o-family tokenizer4texts = [5 "Please block my debit card.",6 "Mera debit card block kar do please.",7 "कृपया मेरा डेबिट कार्ड ब्लॉक करें।",8]9for t in texts:10 print(len(enc.encode(t)), t)The code prints the token count for the same request in English, Romanised Hinglish and Devanagari Hindi. Other providers ship their own tokenizers and token-counting endpoints, so always count with the tokenizer of the model you actually call.
A real-life example
A bank's support bot serves customers who write in English, Hinglish ("mera card block kar do") and Hindi. The team budgets using an English average: 1,500 tokens per chat, 100,000 chats a day, at an assumed $2.50 per million input tokens:
100,000 x 1,500 = 150 million tokens/day150 x $2.50 = $375/dayAfter launch, the bill is about 40% higher. Logs show that Devanagari chats use far more tokens than the English estimate, and Hinglish with creative spellings ("plz", "kiya h") splits into many small pieces. The fixes: budget per language from real traffic, move to a model whose newer tokenizer handles Indic scripts more efficiently, and trim the long system prompt that was repeated in every turn.
Follow-up questions to expect
- "Roughly how many tokens is a page of English?" — About 500 words, so around 650–700 tokens, using the 0.75 words-per-token rule.
- "Why do LLMs struggle with arithmetic on long numbers?" — Numbers are split into irregular chunks ("12345" might become "123" + "45"), so digits do not line up the way they do on paper. Tools or code are the fix.
- "Can you change a model's tokenizer?" — Not without retraining the embeddings; the token IDs are baked into the weights. You can add a few new tokens and train their embeddings.