Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is instruction tuning, and why is it critical for chat assistants?


Where the loss is computed in one training chatsystemReply inHinglishuserMeraorderkab?assistantKalshaamtakend ofturn0123456maskedmaskedloss hereand hereassistant_only_loss=True trains only the highlighted tokens.
The end-of-turn token is part of the target — a model that never learns it keeps writing and invents the customer's next message.

What you need to know

What the data looks like

JSON
{"messages": [  {"role": "system", "content": "You are Sahayak. Reply in casual Hinglish."},  {"role": "user", "content": "Mera order kab aayega?"},  {"role": "assistant", "content": "Aapka order kal shaam tak pahunch jayega."}]}

The chat template

Every chat model turns those messages into a single string with special tokens. For Qwen2.5 it looks like this:

Text
<|im_start|>systemYou are Sahayak. Reply in casual Hinglish.<|im_end|><|im_start|>userMera order kab aayega?<|im_end|><|im_start|>assistant

Llama 3 uses different tokens (<|start_header_id|>, <|eot_id|>). The model learned its own format in instruction tuning, and it only works well with that format.

Python
from transformers import AutoTokenizertok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")text = tok.apply_chat_template(messages, tokenize=False,                               add_generation_prompt=True)print(text)   # always look at one rendered example before training

add_generation_prompt=True adds the opening of the assistant turn, which you want at inference time. In TRL, SFTTrainer applies the template for you when the dataset has a messages column.

Loss only on the answer

You want the model to learn to write answers, not to predict users' questions. In TRL 1.13, SFTConfig(assistant_only_loss=True) computes loss only on assistant turns (it needs a chat template that marks assistant spans). For prompt/completion datasets, loss is on the completion by default.

Why it matters for chat assistants

  • The model learns where its turn ends — the end-of-turn token. Without it, it keeps writing and invents the next user message.
  • Roles give a structure for multi-turn memory and for treating system instructions differently from user text.
  • Research on instruction tuning (such as Google's FLAN work) showed that training on many varied instructions improves answers to instructions the model never saw.

A real-life example

A team fine-tunes a Qwen instruct model for Hindi support. They format training data by hand as ### User: and ### Assistant: lines. In production, vLLM uses the model's built-in template instead.

The results look strange: some replies end with a made-up ### User: line, and the Hinglish style is weaker than in offline tests. The model learned the new style under one format and is being asked under another. They re-render the training data with apply_chat_template, retrain, and the problem disappears.

Follow-up questions to expect

  • "Why mask the prompt tokens?" — Without masking, part of the training effort goes into predicting users' messages and long system prompts, which wastes capacity and can make the model repeat boilerplate.
  • "Should the system prompt be in the training data?" — Yes, if you will use one in production. Train with the same system prompt you will serve with, or the model's behaviour will shift when it appears.
  • "Base or instruct as the starting point?" — Instruct in almost all product work: it already follows instructions and has safety training. Start from a base model only for continued pretraining or when you control the whole post-training pipeline.