Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

How do you pick the right base model before you fine-tune?


What you need to know

1. Size from the serving budget

Weights in bf16 take 2 bytes per parameter: an 8B model is about 16 GB, a 70B model about 140 GB (two 80 GB GPUs before any KV cache). Quantized to 4-bit for serving, an 8B model is about 5 GB. Fine-tuning rarely makes a too-small model reason well, so pick the smallest model whose few-shot score is already close to what you need.

2. Licence

This removes more candidates than benchmarks do. Read the actual licence for the exact version:

Family (examples)Licence typeThings to check
Qwen3, Mistral Small, gpt-ossApache 2.0Permissive; keep notices
Llama 3.xLlama Community LicenceExtra terms above 700M monthly users, acceptable-use policy, naming and attribution rules
Gemma 3Gemma Terms of UseProhibited-use policy passes to your users

3. Base or instruct

Use the instruct version when your data is conversational and you are refining behaviour — it already follows instructions and uses a chat template. Use the base version for heavy continual pretraining, or when you want full control of the format.

4. Language and tokenizer fit

The tokenizer decides how many tokens your text becomes, which sets training cost, serving cost and how much fits in context. This matters a lot for Hindi and other Indian languages, where older tokenizers split words into many pieces.

Python
from transformers import AutoTokenizertext = open("hindi_tickets_sample.txt", encoding="utf-8").read()words = len(text.split())for name in ["Qwen/Qwen3-8B", "google/gemma-3-12b-it", "meta-llama/Llama-3.1-8B-Instruct"]:    tok = AutoTokenizer.from_pretrained(name)    print(f"{name}: {len(tok.encode(text)) / words:.2f} tokens per word")

Run it on a sample of your text. A model that needs 1.5× more tokens per word costs about 1.5× more to train and serve.

5. Ecosystem support

Does vLLM serve it with LoRA? Does PEFT know its layer names? Is the chat template stable? Is there a good quantized path? A model without tooling can cost weeks.

6. The bake-off

Test two or three candidates on your own held-out set, zero-shot and few-shot, with the same prompts. Public leaderboards are partly contaminated and do not match your data.

A real-life example

An e-commerce company wants a brand-voice product-description writer in English and Hindi. Constraints: one L4 GPU (24 GB) per region, under 3 seconds for a 150-word description, and a licence that allows commercial use.

The budget rules out anything above about 14B in bf16. Three 7–12B instruct models pass the licence check. The tokenizer test on 2,000 of their Hindi descriptions shows one model using about 1.6× more tokens per word than the best one, so it is dropped for cost. The remaining two are scored few-shot on 200 held-out products by the brand team. One wins clearly on Hindi fluency, and it becomes the base. The whole selection took two days and no training runs.

Follow-up questions to expect

  • "Is a bigger model always better to fine-tune?" — Better quality, often, but it costs more to train and serve. If a smaller model gets close few-shot, fine-tuning often closes the gap.
  • "What about Mixture-of-Experts models?" — Cheap to run per token, but all experts must be in memory, so fine-tuning is memory-heavy (see Section 6).
  • "Why not trust public leaderboards?" — Benchmark items leak into training data, and leaderboards measure general skills, not your task, language or format.