Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

Which public benchmarks are used for fine-tuned models, and where do they fall short?


What you need to know

The main benchmarks

AreaBenchmarksNote
KnowledgeMMLU, MMLU-Pro, GPQA DiamondOriginal MMLU is saturated for strong models
MathsGSM8K, MATH, AIMEGSM8K is near ceiling; AIME uses new yearly problems
CodeHumanEval, MBPP, LiveCodeBench, SWE-bench VerifiedLiveCodeBench uses problems released after a date to limit leaks
InstructionsIFEvalChecks verifiable rules such as "use exactly 3 bullet points"
Tool callingBFCL, tau-benchFunction-call accuracy and multi-turn tool use
Chat qualityMT-Bench, Arena-Hard, LMArenaJudge-scored or human votes
EmbeddingsMTEB, BEIRRetrieval, clustering, similarity
Indian languagesAI4Bharat's MILU, IndicGenBenchKnowledge and generation in Indian languages

Where they fall short

  • Contamination. Test questions leak into pretraining data, so scores rise without ability rising. You often cannot check whether a model saw them.
  • Saturation. When top models score near 90%+, differences are noise.
  • Format sensitivity. The prompt template, few-shot count, chat template and answer-extraction rule can move scores by several points. Numbers from different papers often are not comparable.
  • Gaming. Benchmarks become targets. Leaderboards have been criticised for letting labs test many private variants and publish the best; Hugging Face retired its Open LLM Leaderboard in 2025.
  • Distribution mismatch. No public benchmark measures your documents, your language mix, your format or your tools.

Using them well: a regression suite

Run the same public tasks on the base and the fine-tuned model, with identical settings. A drop tells you what fine-tuning broke.

Bash
# lm-evaluation-harness 0.4.xlm-eval run --model hf \  --model_args pretrained=meta-llama/Llama-3.1-8B-Instruct,peft=./support-hi-v3,dtype=bfloat16 \  --tasks ifeval,mmlu_pro,gsm8k --apply_chat_template \  --batch_size 8 --output_path results/support-hi-v3

Run it once without peft=... for the base, then compare. The harness loads the LoRA adapter on top of the base directly.

A real-life example

A telecom company fine-tunes its Hindi customer-support model. On its own 600-ticket eval set, correct resolutions go from 71% (base, few-shot) to 86%. The regression suite shows:

TaskBaseFine-tuned
IFEval8072
GSM8K8483
MMLU-Pro4746

The 8-point IFEval drop matters, because agents sometimes give the bot formatting instructions. The team adds 1,500 general instruction-following examples to the training mix, retrains, and gets IFEval back to 78 with the task score still at 85%. The launch decision rests on the ticket eval and a two-week A/B test; the public tasks only caught the side effect. (Made-up numbers for illustration.)

Follow-up questions to expect

  • "How do you check for contamination?" — Look for n-gram overlap between training data and test items, compare scores on rephrased or newer versions of a benchmark, and prefer benchmarks with dated or held-back questions.
  • "Why doesn't my score match the paper?" — Different prompt format, few-shot count, chat template or answer extraction. Fix the settings and compare only within your own runs.
  • "Which benchmark would you use for tool calling?" — BFCL for single calls and tau-bench for multi-turn tool use, plus your own tool schemas.