Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
Which fine-tuning APIs exist, and key tradeoffs in control, cost, and portability?
What you need to know
The three tiers
| Hosted provider API | Managed platform | Self-hosted open stack | |
|---|---|---|---|
| Examples | OpenAI, Google Vertex AI, AWS Bedrock, Mistral | Together, Fireworks, cloud ML platforms | TRL + PEFT, Unsloth, Axolotl, LLaMA-Factory |
| Base models | The provider's own (and some partners') | Open-weight models | Any open-weight model |
| Control | Epochs, learning-rate multiplier, batch size | Most training settings | Everything: loss, masking, precision, method |
| Weights | Usually not downloadable | Adapter or merged weights usually downloadable | Yours |
| Data | Leaves your environment | Leaves your environment | Stays with you |
| Effort | Lowest | Low to medium | Highest |
A hosted job, end to end
1# openai Python SDK (current releases)2from openai import OpenAI3client = OpenAI()45f = client.files.create(file=open("train.jsonl", "rb"), purpose="fine-tune")6job = client.fine_tuning.jobs.create(7 model="gpt-4.1-mini-2025-04-14", # pick a fine-tunable snapshot from the current list8 training_file=f.id,9 method={"type": "supervised", "supervised": {"hyperparameters": {"n_epochs": 3}}},10)11print(job.id, job.status)The method field also accepts "dpo" and, for some reasoning models, "reinforcement". Vertex AI offers supervised tuning of Gemini models, and Bedrock offers fine-tuning, continued pretraining and distillation for selected models. Model lists change often, so check each provider's current documentation.
The self-hosted tools
- TRL + PEFT (Hugging Face) — the reference trainers: SFT, DPO, GRPO and more.
- Unsloth — faster, lower-memory training on a single GPU, with a TRL-compatible interface.
- Axolotl — one YAML file per run, good defaults, multi-GPU.
- LLaMA-Factory — CLI and web UI over many models and methods.
How to choose
- Data rules first. Client-confidential, health or financial data may not be allowed to leave your environment or the country. India's DPDP Act 2023 and client contracts both matter here.
- Portability. If you cannot download the weights, you are renting behaviour. When the provider retires the base model, you must retrain on a new one.
- Total cost, not training cost. Hosted APIs bill per training token: 5,000 examples × 800 tokens × 3 epochs = 12 million billed tokens. Fine-tuned models are also often priced higher per token than the base model at inference, which can dominate the bill. Self-hosting costs GPU-hours plus engineering time, and wins once fine-tuning is routine.
- Method. Custom losses, unusual masking, new preference methods, or a model nobody hosts force you onto the open stack.
A real-life example
A Mumbai law firm wants a legal-clause classifier. To test the idea in one day, an engineer fine-tunes a hosted model on 300 clauses taken only from public, already-published contracts. It reaches a useful accuracy on a small test, which proves the task is learnable.
For the real system, the partners refuse to send client contracts to any third party — they are confidential and privileged. The firm runs QLoRA on an 8B open model on one in-house GPU with TRL, trained on 6,000 internal clauses. The adapter, data and logs never leave the office, the firm owns the weights, and when a newer base model appears it can retrain on its own schedule. The hosted API was the right tool for the one-day test, and the wrong one for production.
Follow-up questions to expect
- "Can you download a model fine-tuned through OpenAI's API?" — No. You get a private model ID to call through the API.
- "What happens when a hosted base model is retired?" — Your fine-tuned model is eventually retired with it, and you must retrain on a newer base. Keep your training data and eval set ready.
- "Unsloth or plain TRL?" — Unsloth for fast single-GPU runs on supported models; plain TRL (with Accelerate) when you need multi-GPU setups or methods Unsloth does not cover.