LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

How do LLMs handle out-of-vocabulary (OOV) words?


What you need to know

The old problem

Word-level models had a fixed list of, say, 50,000 words. Anything else became a special <UNK> token. "PhonePe", a customer's name or a typo all turned into the same <UNK>, and the meaning was gone before the model even started.

How subwords solve it

A tokenizer greedily matches the longest known pieces. An invented example of a split:

Text
"cryptoeconomics" -> "crypto" + "econom" + "ics""acount"          -> "ac" + "ount"          (typo, still encoded)"Zyphora"         -> "Z" + "yph" + "ora"    (new brand name)

The exact splits depend on the tokenizer, but every string gets some encoding.

What still goes wrong

  • More tokens — a rare word in 5 pieces costs 5 times a common word.
  • Weaker meaning — the model must assemble meaning from pieces. "Metformin" is familiar from medical text; a new drug brand split into 4 fragments is less reliably understood.
  • Scripts — each Devanagari character is 3 bytes in UTF-8. If the tokenizer learned few Hindi merges, a word can fall back to many byte-level pieces.
  • Spelling tasks — the model sees chunks, not letters, so spelling out or reversing a rare word is error-prone.

What you can do

  • Give context: "Zyphora (our new savings account)" helps the model a lot.
  • For a domain full of special terms, fine-tune so the model sees them often, or, for open-weight models, add new tokens and train their embeddings.
  • Pick a model whose tokenizer is efficient for your languages.

A real-life example

A bank's chatbot receives: "mera acount blok ho gya, UPI pin bhi nahi chal raha 😟". This mixes Hindi and English in Roman script, has two typos and an emoji. A 2015-style word-level model would turn "mera", "acount", "blok", "gya" and the emoji into <UNK> and understand almost nothing.

A current LLM splits the typos into pieces, still recognises "acount blok" as "account block" from surrounding context ("UPI pin", "nahi chal raha"), and routes the chat to the blocked-account flow. The team notices the message uses about twice the tokens of the clean English version and accepts it, but adds the bank's product names ("Suraksha FD", "Yuva Card") with short descriptions to the system prompt so the model does not guess what they are.

Follow-up questions to expect

  • "Does subword tokenization hurt quality for rare words?" — Somewhat: meaning must be composed from fragments the model has seen in other contexts. Context and fine-tuning reduce the effect.
  • "How do you add domain-specific tokens?" — Add them to the tokenizer, resize the embedding matrix, and fine-tune so the new rows are learned. Initialising them as the average of their old subword embeddings helps.
  • "Is there still an unknown token anywhere?" — Some WordPiece models such as BERT keep an [UNK], but byte-level tokenizers never need one.