Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is continual pre-training, and when is it a better choice than standard fine-tuning?
What you need to know
Continued pretraining
- Raw documents, no labels
- Hundreds of millions to billions of tokens
- Teaches words, notation, domain facts
- Loss on every token
Supervised fine-tuning
- Prompt and ideal answer pairs
- Hundreds to tens of thousands of examples
- Teaches format, tone, task behaviour
- Loss usually only on the answer
When it beats SFT alone
- Language or script the base model saw little of. Such text also tends to split into many more tokens than English, so the model reads it slowly and poorly.
- Dense jargon or notation — clinical shorthand, chip-design specs, a firm's internal code.
- Lots of raw text, few labels — you cannot write enough SFT examples to teach the vocabulary.
If the model already understands the domain's words and you have a few thousand good examples, go straight to SFT.
A typical recipe
- Start from the base model — not the instruct model, whose chat behaviour continued pretraining will wear away.
- Mix the data — domain text plus 5–30% general text ("replay") to limit forgetting.
- Train gently — a short warm-up, then a learning rate well below the original pretraining peak, decaying to near zero.
- Follow with SFT — and preference tuning if needed, to restore chat behaviour on top of the new knowledge.
For a large shift, full-parameter training (or LoRA at high rank on all layers, including embeddings) is usually needed; a small LoRA cannot hold that much new knowledge.
1from trl import SFTConfig, SFTTrainer23# dataset with one column "text": raw Odia documents mixed with general text4args = SFTConfig(output_dir="cpt-odia", dataset_text_field="text",5 packing=True, max_length=4096,6 learning_rate=2e-5, num_train_epochs=1)7trainer = SFTTrainer(model="Qwen/Qwen2.5-7B", args=args, train_dataset=corpus)8trainer.train()With a plain-text dataset, TRL's SFTTrainer computes loss on every token, which is exactly the pretraining objective. packing=True joins short documents into full 4,096-token blocks so no compute is wasted on padding.
A real-life example
The e-commerce support team from section 1 wants to serve Odia-speaking customers. LoRA on a few thousand Odia chats gave fluent Hindi but broken Odia: the base model had seen too little Odia.
They collect about 400 million tokens of Odia text — news, government publications and books with suitable licences — and mix in 20% Hindi and English text. One epoch of continued pretraining on the base model follows. Then they run the same support SFT as before. Odia replies are now grammatical, and the Hindi and English support scores stay within a point of the original. The extra stage took a week of GPU time — worth it for a new market, not for a style fix.
Follow-up questions to expect
- "Should you extend the tokenizer for a new script?" — It can cut token counts and cost a lot, but new token embeddings start untrained and need a lot of continued pretraining to become useful. Many teams keep the original tokenizer unless token counts are very high.
- "Why start from the base model?" — Continued pretraining on an instruct model erodes its chat and safety behaviour, and you would have to rebuild them anyway.
- "How do you know it helped?" — Lower perplexity on held-out domain text, and — more importantly — better scores on the downstream task after SFT, compared with SFT alone.