Natural Language Processing Basics

Tokenization & Text Cleaning


Here is a sentence pulled from a real product review:

Text
Dr. Patel's U.S.-based team said it "isn't cheap" — $9.99/month — but it works!!!

You need to turn that into a list of words so a model can work with it. The obvious move is to split on spaces. Try it, and look carefully at what comes back.

Python
text = 'Dr. Patel\'s U.S.-based team said it "isn\'t cheap" — $9.99/month — but it works!!!'print(text.split())
Text
['Dr.', "Patel's", 'U.S.-based', 'team', 'said', 'it', '"isn\'t', 'cheap"', '—', '$9.99/month', '—', 'but', 'it', 'works!!!']

Count the damage. works!!! is now a different word from works, and from works!, and from works. — your model will treat all four as unrelated. "isn't carries a stray quotation mark. $9.99/month has fused a price and a billing period into one token that will appear exactly once in your entire corpus and teach the model nothing. Meanwhile U.S.-based got lucky and stayed whole, but if a naive rule had split on punctuation it would have shattered into U, S, based.

Splitting on whitespace looks like a solved problem and is not. Everything in this lesson exists because that one line of code is wrong in about nine different ways, and each of those ways costs you accuracy downstream.

A word the tokeniser has never seen, still readableunhappiness012negationnounendingWord-level tokenisation returns UNK here and the sentiment of the sentence is lost with it.
Subwords trade one clean token for three reusable ones, and that is what removes the out-of-vocabulary hole.

What a token actually is

A token is the smallest unit of text your model is allowed to see. Tokenisation is the process of cutting text into those units.

Notice the definition says nothing about words. That is deliberate. A token might be a word, a piece of a word, a punctuation mark, or a single character. Which one you choose is a design decision, and it is yours to make.

Your model never sees text. It sees whatever your tokeniser hands it. Every distinction the tokeniser destroys is a distinction the model can never learn, and every spurious distinction it creates is noise the model must spend capacity ignoring.

That is why this step deserves more care than it usually gets. It is the only part of the pipeline that is genuinely irreversible.

Word tokenisation done properly

A real word tokeniser does not split on spaces. It applies a set of learned or hand-tuned rules that separate punctuation from words while protecting the cases where punctuation is part of the word.

Python
import nltkfor pkg in ("punkt_tab", "stopwords", "wordnet"):   # 'punkt_tab' replaced the old 'punkt'    nltk.download(pkg)from nltk.tokenize import word_tokenizetext = 'Dr. Patel\'s U.S.-based team said it "isn\'t cheap" but it works!!!'print(word_tokenize(text))
Text
['Dr.', 'Patel', "'s", 'U.S.-based', 'team', 'said', 'it', '``', 'is', "n't", 'cheap', "''", 'but', 'it', 'works', '!', '!', '!']

Three things improved. Dr. kept its full stop because the tokeniser holds a list of known abbreviations. isn't became is + n't, which matters enormously — the negation is now its own token that a model can learn to react to, rather than being buried inside a contraction. And works is finally just works.

Different tokenisers make different calls, and none of them is universally right:

InputNaive splitNLTK word_tokenizespaCy
don'tdon'tdo, n'tdo, n't
New YorkNew, YorkNew, YorkNew, York
state-of-the-artone tokenone tokenstate, -, of, -, the, -, art
you@example.comone tokenyou, @, example.comone token
:-)one token:, -, ):-)

The last three rows are where the tokenisers disagree: spaCy splits the hyphenated compound, while NLTK shreds the email address and the emoticon. The emoticon is the real trap. If you are classifying sentiment on social media, a tokeniser that shreds emoticons into meaningless punctuation has thrown away one of your strongest signals. NLTK ships TweetTokenizer for exactly this reason.

Sentence tokenisation and the full-stop problem

Splitting into sentences sounds easier than splitting into words. It is not.

The naive rule is "split on ., !, ?". Now feed it this:

Text
Dr. Patel joined in Jan. 2019. He paid $9.99 for the U.S. plan. It worked.

The naive rule produces eight pieces, and only the last, It worked, is a whole sentence; the rest are fragments like Dr, 2019 and S. The full stop is doing two completely different jobs in the same text — ending a sentence, and marking an abbreviation — and nothing about the character itself tells you which.

Proper sentence splitters (NLTK's Punkt, spaCy's parser-driven splitter) handle this with a mixture of abbreviation lists and statistics about which tokens tend to start sentences.

Python
from nltk.tokenize import sent_tokenizetext = "Dr. Patel joined in Jan. 2019. He paid $9.99 for the U.S. plan. It worked."for s in sent_tokenize(text):    print(repr(s))
Text
'Dr. Patel joined in Jan. 2019.''He paid $9.99 for the U.S. plan.''It worked.'

You need sentence tokenisation whenever your unit of analysis is a sentence rather than a document — sentence-level sentiment, summarisation, or feeding long documents to a model with a fixed input limit.

Subword tokenisation: the fix for words you have never seen

Word-level tokenisation has a hard ceiling, and it is worth seeing the ceiling clearly.

You build a vocabulary from your training data — say the 30,000 most common words. At inference time a user writes unfriendliness, which never appeared in training. Your tokeniser maps it to <UNK>, the unknown token. So does cryptocurrencies. So does Bhattacharya. So does every typo anyone ever makes. All of them collapse to the same meaningless symbol, and the model has no way to tell them apart.

This is the out-of-vocabulary problem, and it is not rare. English has an effectively unbounded number of words: every language keeps inventing them, and morphology (prefixes, suffixes, compounds) multiplies them.

Subword tokenisation solves it by keeping common words whole and breaking rare words into pieces that are in the vocabulary.

Text
unfriendliness  ->  un ##friend ##li ##nesstokenization    ->  token ##izationthe             ->  the

The ## prefix marks "this piece continues the previous one", so the sequence can be joined back into the original string. Now unfriendliness is not an unknown blob — it is built from un (which the model has learned negates things) and friend (which it has plenty of evidence about). Nothing is ever truly unknown, because in the worst case a word decomposes into individual characters.

How byte-pair encoding builds the vocabulary

Byte-pair encoding (BPE) is the most common way to learn these pieces, and the algorithm is short enough to follow by hand. Start with a corpus where every word is split into characters, plus an end-of-word marker:

Text
Corpus counts:  l o w </w>        5  l o w e r </w>    2  n e w e s t </w>  6  w i d e s t </w>  3

Now repeat: find the most frequent adjacent pair of symbols, and merge it into one symbol.

StepMost frequent pairCountNew symbol
1e s6 + 3 = 9es
2es t9est
3est </w>9est</w>
4l o5 + 2 = 7lo
5lo w7low

After five merges the tokeniser has discovered the suffix est and the stem low without anyone telling it that English has superlatives. Stop after a fixed number of merges — typically enough to reach a vocabulary of 30,000 to 50,000 symbols — and you have a tokeniser that covers any input.

Subword tokenisation is why modern language models have no unknown-word problem at all. The cost is that a token is no longer a word, so you can never assume one token equals one word when you reason about lengths or costs.

That last point has a practical bite. If a model has a 512-token limit, that is roughly 380 English words, not 512 — and far fewer for text full of names, code, or non-English script.

Cleaning: every step throws information away

Cleaning is where most people go wrong, because the standard advice — lowercase everything, strip punctuation, remove stopwords — is presented as a checklist rather than as a set of trade-offs. Each of those steps deletes information. Sometimes the information was noise. Sometimes it was your signal.

StepWhat it buys youWhat it destroysSkip it when
LowercasingThe and the stop being separate features; smaller vocabularyApple vs apple, US vs us, ALL-CAPS shoutingDoing named entity recognition, or capitalisation carries emotion
Removing punctuationFewer junk tokens!!! as an intensity signal, ? as a question marker, emoticonsSentiment, question detection, social media
Removing stopwordsSmaller feature space, faster modelsnot, no, never, and all grammarAnything where negation or word order matters
Removing numbersStops one-off tokens like 9.99 bloating the vocabularyRatings, years, quantities, pricesThe numbers are the point (reviews, finance)
Stripping HTMLRemoves <br /> noise from scraped textAlmost nothingEssentially never — do this

The stopword mistake that quietly ruins sentiment models

This one is worth staring at. Standard English stopword lists — including NLTK's — contain not, no, nor, and against.

Python
from nltk.corpus import stopwordssw = set(stopwords.words('english'))s = "this movie was not good at all"print([w for w in s.split() if w not in sw])
Text
['movie', 'good']

A scathing review has become the token sequence movie good. You have not cleaned the text; you have inverted its meaning. Every model trained on that pipeline inherits the error, and it will never show up as a crash — only as a mysterious accuracy ceiling.

The fix is to keep a custom list:

Python
NEGATIONS = {"not", "no", "nor", "never", "n't", "cannot", "without", "against"}custom_stopwords = set(stopwords.words('english')) - NEGATIONSprint([w for w in "this movie was not good at all".split()       if w not in custom_stopwords])# ['movie', 'not', 'good']

Stemming versus lemmatisation

Both aim at the same goal: collapse run, runs, running into one feature so the model does not have to learn about each separately. They go about it in opposite ways.

Stemming chops suffixes off using rules. It is fast, language-specific, and produces strings that are often not real words.

Lemmatisation looks the word up in a dictionary and returns its base form, using the part of speech to disambiguate. It is slower, needs linguistic resources, and always produces a real word.

Python
from nltk.stem import PorterStemmer, WordNetLemmatizerstem = PorterStemmer().stemlem = WordNetLemmatizer().lemmatizefor w in ["studies", "studying", "better", "was", "flies", "universal"]:    print(f"{w:12} stem={stem(w):10} lemma(verb)={lem(w, pos='v')}")
WordPorter stemLemmaComment
studiesstudistudyStem is not a word, but it is consistent
studyingstudistudyBoth collapse correctly
betterbetterbetter as a verb; good with pos='a'Only the lemmatiser can know this, and only when told the part of speech
waswabeStemmer produces nonsense
fliesfliflyBoth collapse noun and verb senses
universaluniversuniversalStemmer merges it with university — a real error

That final row is the classic stemming failure: universal, universe and university all stem to univers, so a search engine using Porter stemming treats a query about universities as a query about the universe.

The lemmatiser has its own trap. Without a part-of-speech tag it assumes everything is a noun:

Python
lem = WordNetLemmatizer().lemmatizeprint(lem("running"))            # 'running'  <- wrong, treated as a nounprint(lem("running", pos="v"))   # 'run'      <- correct

If you lemmatise without tagging, roughly every verb in your corpus is left untouched and you have paid the cost of lemmatisation for a fraction of the benefit.

StemmingLemmatisation
SpeedVery fast (pure string rules)10–100× slower; dictionary lookups, needs POS tagging
OutputOften not a real wordAlways a real word
ErrorsOver-merges (universal/university)Under-merges if POS tag is missing or wrong
Good forSearch indexing, huge corpora, bag-of-words featuresAnything a human will read; smaller corpora; linguistic analysis

Order of operations

Preprocessing steps are not commutative. Running them in the wrong order silently breaks things.

  1. Strip HTML first. If you lowercase and remove punctuation before stripping tags, <br /> becomes the token br, which will end up in your top-20 most frequent words for any scraped corpus.
  2. Normalise whitespace and unicode next. Curly quotes, non-breaking spaces and em dashes should become their plain equivalents before anything counts characters.
  3. Tokenise before removing punctuation. The tokeniser needs punctuation to find token boundaries. Strip it first and isn't becomes isnt, which is in nobody's vocabulary.
  4. Lowercase after tokenising if you are lowercasing at all — some tokenisers use capitalisation to detect sentence starts.
  5. Stem or lemmatise last, and only one of the two. Doing both is pointless: the stemmer's output is not a dictionary word, so the lemmatiser cannot find it.

A pipeline you can read

Python
import reimport unicodedatafrom nltk.tokenize import word_tokenizefrom nltk.corpus import stopwordsfrom nltk.stem import WordNetLemmatizerNEGATIONS = {"not", "no", "nor", "never", "n't", "cannot"}STOPWORDS = set(stopwords.words("english")) - NEGATIONSLEMMATIZER = WordNetLemmatizer()def preprocess(text, drop_stopwords=False, lemmatise=True):    # 1. HTML    text = re.sub(r"<[^>]+>", " ", text)    # 2. URLs and emails become type markers, not one-off tokens    text = re.sub(r"http\S+|www\.\S+", " URL ", text)    text = re.sub(r"\S+@\S+\.\S+", " EMAIL ", text)    # 3. Unicode normalisation, then whitespace    text = unicodedata.normalize("NFKC", text)    text = re.sub(r"\s+", " ", text).strip()    # 4. Tokenise, then lowercase    tokens = [t.lower() for t in word_tokenize(text)]    # 5. Drop pure punctuation, keep words and numbers    tokens = [t for t in tokens if re.search(r"[a-z0-9]", t)]    if drop_stopwords:        tokens = [t for t in tokens if t not in STOPWORDS]    if lemmatise:        tokens = [LEMMATIZER.lemmatize(t, pos="v") for t in tokens]    return tokensprint(preprocess("<p>Dr. Patel said it isn't cheap!!! See https://x.com</p>"))# ['dr.', 'patel', 'say', 'it', 'be', "n't", 'cheap', 'see', 'url']

The same thing with spaCy

spaCy does tokenisation, part-of-speech tagging and lemmatisation in one pass, which means its lemmas are correct without you supplying tags by hand.

Python
import spacynlp = spacy.load("en_core_web_sm")doc = nlp("Dr. Patel was running the U.S.-based studies in Jan. 2019.")for tok in doc:    print(f"{tok.text:12} {tok.lemma_:10} {tok.pos_:6} stop={tok.is_stop}")
Text
Dr.          Dr.        PROPN  stop=FalsePatel        Patel      PROPN  stop=Falsewas          be         AUX    stop=Truerunning      run        VERB   stop=Falsethe          the        DET    stop=TrueU.S.-based   u.s.-base  VERB   stop=Falsestudies      study      NOUN   stop=Falsein           in         ADP    stop=TrueJan.         Jan.       PROPN  stop=False2019         2019       NUM    stop=False.            .          PUNCT  stop=False

Note was → be and running → run, both of which the tag-free NLTK lemmatiser gets wrong. The small model is not perfect either: it tags U.S.-based as a verb and lemmatises it to u.s.-base. The cost is speed: spaCy's full pipeline is far slower than a regex. Disable the components you do not need with nlp.select_pipes(disable=["ner", "parser"]) and you get most of the speed back.

Choosing a pipeline for the job in front of you

There is no default. The right pipeline is a function of what your model is and what the task rewards.

TaskTokenisationLowercaseStopwordsStem/lemma
Topic classification with TF-IDFWordYesRemoveLemmatise
Sentiment analysisWordYesKeep negationsLemmatise, keep punctuation
Named entity recognitionWordNoKeep allNone
Search indexingWordYesRemoveStem (speed wins)
Fine-tuning a pretrained transformerIts own subword tokeniserWhatever it was trained withKeep allNone

That last row is the one people break most often. A pretrained model learned its embeddings against a specific tokeniser with a specific vocabulary. If you lowercase text before feeding a case-sensitive model, or strip punctuation before a tokeniser that was trained expecting it, every token index shifts away from what the model saw during pretraining. Accuracy drops and there is no error message. Load the tokeniser that ships with the checkpoint and hand it raw text.

Aggressive cleaning is a habit inherited from an era when models were small and every feature was expensive. The stronger your model, the less you should touch the text.

One operational rule holds regardless of which pipeline you pick: write it once, as a function, and call the same function everywhere — training, validation, evaluation, and the live service. The most common preprocessing bug in production is not a bad choice of stemmer. It is a training script that lowercases and an inference endpoint that forgets to, producing a model that scores 91% offline and 68% for real users. Save the preprocessing configuration next to the model weights, and treat them as a single artefact that cannot be separated.