Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
How do you pick the right base model before you fine-tune?
What you need to know
1. Size from the serving budget
Weights in bf16 take 2 bytes per parameter: an 8B model is about 16 GB, a 70B model about 140 GB (two 80 GB GPUs before any KV cache). Quantized to 4-bit for serving, an 8B model is about 5 GB. Fine-tuning rarely makes a too-small model reason well, so pick the smallest model whose few-shot score is already close to what you need.
2. Licence
This removes more candidates than benchmarks do. Read the actual licence for the exact version:
| Family (examples) | Licence type | Things to check |
|---|---|---|
| Qwen3, Mistral Small, gpt-oss | Apache 2.0 | Permissive; keep notices |
| Llama 3.x | Llama Community Licence | Extra terms above 700M monthly users, acceptable-use policy, naming and attribution rules |
| Gemma 3 | Gemma Terms of Use | Prohibited-use policy passes to your users |
3. Base or instruct
Use the instruct version when your data is conversational and you are refining behaviour — it already follows instructions and uses a chat template. Use the base version for heavy continual pretraining, or when you want full control of the format.
4. Language and tokenizer fit
The tokenizer decides how many tokens your text becomes, which sets training cost, serving cost and how much fits in context. This matters a lot for Hindi and other Indian languages, where older tokenizers split words into many pieces.
1from transformers import AutoTokenizer23text = open("hindi_tickets_sample.txt", encoding="utf-8").read()4words = len(text.split())5for name in ["Qwen/Qwen3-8B", "google/gemma-3-12b-it", "meta-llama/Llama-3.1-8B-Instruct"]:6 tok = AutoTokenizer.from_pretrained(name)7 print(f"{name}: {len(tok.encode(text)) / words:.2f} tokens per word")Run it on a sample of your text. A model that needs 1.5× more tokens per word costs about 1.5× more to train and serve.
5. Ecosystem support
Does vLLM serve it with LoRA? Does PEFT know its layer names? Is the chat template stable? Is there a good quantized path? A model without tooling can cost weeks.
6. The bake-off
Test two or three candidates on your own held-out set, zero-shot and few-shot, with the same prompts. Public leaderboards are partly contaminated and do not match your data.
A real-life example
An e-commerce company wants a brand-voice product-description writer in English and Hindi. Constraints: one L4 GPU (24 GB) per region, under 3 seconds for a 150-word description, and a licence that allows commercial use.
The budget rules out anything above about 14B in bf16. Three 7–12B instruct models pass the licence check. The tokenizer test on 2,000 of their Hindi descriptions shows one model using about 1.6× more tokens per word than the best one, so it is dropped for cost. The remaining two are scored few-shot on 200 held-out products by the brand team. One wins clearly on Hindi fluency, and it becomes the base. The whole selection took two days and no training runs.
Follow-up questions to expect
- "Is a bigger model always better to fine-tune?" — Better quality, often, but it costs more to train and serve. If a smaller model gets close few-shot, fine-tuning often closes the gap.
- "What about Mixture-of-Experts models?" — Cheap to run per token, but all experts must be in memory, so fine-tuning is memory-heavy (see Section 6).
- "Why not trust public leaderboards?" — Benchmark items leak into training data, and leaderboards measure general skills, not your task, language or format.