Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is instruction tuning, and why is it critical for chat assistants?
What you need to know
What the data looks like
1{"messages": [2 {"role": "system", "content": "You are Sahayak. Reply in casual Hinglish."},3 {"role": "user", "content": "Mera order kab aayega?"},4 {"role": "assistant", "content": "Aapka order kal shaam tak pahunch jayega."}5]}The chat template
Every chat model turns those messages into a single string with special tokens. For Qwen2.5 it looks like this:
<|im_start|>systemYou are Sahayak. Reply in casual Hinglish.<|im_end|><|im_start|>userMera order kab aayega?<|im_end|><|im_start|>assistantLlama 3 uses different tokens (<|start_header_id|>, <|eot_id|>). The model learned its own format in instruction tuning, and it only works well with that format.
1from transformers import AutoTokenizer2tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")3text = tok.apply_chat_template(messages, tokenize=False,4 add_generation_prompt=True)5print(text) # always look at one rendered example before trainingadd_generation_prompt=True adds the opening of the assistant turn, which you want at inference time. In TRL, SFTTrainer applies the template for you when the dataset has a messages column.
Loss only on the answer
You want the model to learn to write answers, not to predict users' questions. In TRL 1.13, SFTConfig(assistant_only_loss=True) computes loss only on assistant turns (it needs a chat template that marks assistant spans). For prompt/completion datasets, loss is on the completion by default.
Why it matters for chat assistants
- The model learns where its turn ends — the end-of-turn token. Without it, it keeps writing and invents the next user message.
- Roles give a structure for multi-turn memory and for treating system instructions differently from user text.
- Research on instruction tuning (such as Google's FLAN work) showed that training on many varied instructions improves answers to instructions the model never saw.
A real-life example
A team fine-tunes a Qwen instruct model for Hindi support. They format training data by hand as ### User: and ### Assistant: lines. In production, vLLM uses the model's built-in template instead.
The results look strange: some replies end with a made-up ### User: line, and the Hinglish style is weaker than in offline tests. The model learned the new style under one format and is being asked under another. They re-render the training data with apply_chat_template, retrain, and the problem disappears.
Follow-up questions to expect
- "Why mask the prompt tokens?" — Without masking, part of the training effort goes into predicting users' messages and long system prompts, which wastes capacity and can make the model repeat boilerplate.
- "Should the system prompt be in the training data?" — Yes, if you will use one in production. Train with the same system prompt you will serve with, or the model's behaviour will shift when it appears.
- "Base or instruct as the starting point?" — Instruct in almost all product work: it already follows instructions and has safety training. Start from a base model only for continued pretraining or when you control the whole post-training pipeline.