Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is Next Sentence Prediction (NSP)?
What you need to know
How NSP training pairs are built
Positive (IsNext):A: "The customer reported a failed UPI payment."B: "The amount was debited but not credited to the merchant."Negative (NotNext):A: "The customer reported a failed UPI payment."B: "Mix the flour and sugar before adding eggs."BERT puts both sentences in one input as [CLS] A [SEP] B [SEP], and a small classifier on the [CLS] vector predicts IsNext or NotNext.
Why it was too easy
The negative sentence is random, so it almost always comes from a different document and a different topic. "UPI payment" and "flour and sugar" share no words or subject. The model can score well just by detecting a topic change, which teaches little about whether B logically follows A.
What replaced it
Next Sentence Prediction (BERT)
- Negative B is a random sentence from another document
- Solvable by topic matching
- RoBERTa removed it and scored better
Sentence Order Prediction (ALBERT)
- Negative is the same two sentences, swapped
- Topic is identical, so only order and logic help
- Teaches real discourse coherence
RoBERTa's result was the turning point: training on longer contiguous text without NSP, with more data and longer training, gave better downstream scores.
Where coherence comes from today
Decoder LLMs learn inter-sentence coherence implicitly: predicting the next token across long documents rewards knowing what should come next at every level — word, sentence and paragraph. No separate sentence-pair objective is needed.
A real-life example
A legal-document team fine-tuned an early BERT model to check whether a clause logically continued the previous one, hoping NSP pretraining would help. It flagged obvious junk — a recipe pasted into a contract — but missed real errors, such as clause 7.2 ("the buyer shall then pay the balance") appearing before clause 7.1 ("the buyer shall inspect the goods"). Both clauses were on the same topic, so the NSP-style signal was useless.
A model with order-based training, or a modern LLM prompted with "Does clause B follow logically from clause A? Answer yes or no with a reason", catches the swapped order because it must reason about sequence, not topic.
Follow-up questions to expect
- "What were BERT's two pretraining objectives?" — Masked language modelling and Next Sentence Prediction, trained together.
- "Why did RoBERTa drop NSP?" — Experiments showed removing it did not hurt and often helped, especially when training on full contiguous text instead of sentence pairs.
- "What is the
[CLS]token?" — A special token at the start of BERT's input whose final vector is used as a summary of the whole input for classification.