Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
Which public benchmarks are used for fine-tuned models, and where do they fall short?
What you need to know
The main benchmarks
| Area | Benchmarks | Note |
|---|---|---|
| Knowledge | MMLU, MMLU-Pro, GPQA Diamond | Original MMLU is saturated for strong models |
| Maths | GSM8K, MATH, AIME | GSM8K is near ceiling; AIME uses new yearly problems |
| Code | HumanEval, MBPP, LiveCodeBench, SWE-bench Verified | LiveCodeBench uses problems released after a date to limit leaks |
| Instructions | IFEval | Checks verifiable rules such as "use exactly 3 bullet points" |
| Tool calling | BFCL, tau-bench | Function-call accuracy and multi-turn tool use |
| Chat quality | MT-Bench, Arena-Hard, LMArena | Judge-scored or human votes |
| Embeddings | MTEB, BEIR | Retrieval, clustering, similarity |
| Indian languages | AI4Bharat's MILU, IndicGenBench | Knowledge and generation in Indian languages |
Where they fall short
- Contamination. Test questions leak into pretraining data, so scores rise without ability rising. You often cannot check whether a model saw them.
- Saturation. When top models score near 90%+, differences are noise.
- Format sensitivity. The prompt template, few-shot count, chat template and answer-extraction rule can move scores by several points. Numbers from different papers often are not comparable.
- Gaming. Benchmarks become targets. Leaderboards have been criticised for letting labs test many private variants and publish the best; Hugging Face retired its Open LLM Leaderboard in 2025.
- Distribution mismatch. No public benchmark measures your documents, your language mix, your format or your tools.
Using them well: a regression suite
Run the same public tasks on the base and the fine-tuned model, with identical settings. A drop tells you what fine-tuning broke.
1# lm-evaluation-harness 0.4.x2lm-eval run --model hf \3 --model_args pretrained=meta-llama/Llama-3.1-8B-Instruct,peft=./support-hi-v3,dtype=bfloat16 \4 --tasks ifeval,mmlu_pro,gsm8k --apply_chat_template \5 --batch_size 8 --output_path results/support-hi-v3Run it once without peft=... for the base, then compare. The harness loads the LoRA adapter on top of the base directly.
A real-life example
A telecom company fine-tunes its Hindi customer-support model. On its own 600-ticket eval set, correct resolutions go from 71% (base, few-shot) to 86%. The regression suite shows:
| Task | Base | Fine-tuned |
|---|---|---|
| IFEval | 80 | 72 |
| GSM8K | 84 | 83 |
| MMLU-Pro | 47 | 46 |
The 8-point IFEval drop matters, because agents sometimes give the bot formatting instructions. The team adds 1,500 general instruction-following examples to the training mix, retrains, and gets IFEval back to 78 with the task score still at 85%. The launch decision rests on the ticket eval and a two-week A/B test; the public tasks only caught the side effect. (Made-up numbers for illustration.)
Follow-up questions to expect
- "How do you check for contamination?" — Look for n-gram overlap between training data and test items, compare scores on rephrased or newer versions of a benchmark, and prefer benchmarks with dated or held-back questions.
- "Why doesn't my score match the paper?" — Different prompt format, few-shot count, chat template or answer extraction. Fix the settings and compare only within your own runs.
- "Which benchmark would you use for tool calling?" — BFCL for single calls and tau-bench for multi-turn tool use, plus your own tool schemas.