Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

Which fine-tuning APIs exist, and key tradeoffs in control, cost, and portability?


What you need to know

The three tiers

Hosted provider APIManaged platformSelf-hosted open stack
ExamplesOpenAI, Google Vertex AI, AWS Bedrock, MistralTogether, Fireworks, cloud ML platformsTRL + PEFT, Unsloth, Axolotl, LLaMA-Factory
Base modelsThe provider's own (and some partners')Open-weight modelsAny open-weight model
ControlEpochs, learning-rate multiplier, batch sizeMost training settingsEverything: loss, masking, precision, method
WeightsUsually not downloadableAdapter or merged weights usually downloadableYours
DataLeaves your environmentLeaves your environmentStays with you
EffortLowestLow to mediumHighest

A hosted job, end to end

Python
# openai Python SDK (current releases)from openai import OpenAIclient = OpenAI()f = client.files.create(file=open("train.jsonl", "rb"), purpose="fine-tune")job = client.fine_tuning.jobs.create(    model="gpt-4.1-mini-2025-04-14",   # pick a fine-tunable snapshot from the current list    training_file=f.id,    method={"type": "supervised", "supervised": {"hyperparameters": {"n_epochs": 3}}},)print(job.id, job.status)

The method field also accepts "dpo" and, for some reasoning models, "reinforcement". Vertex AI offers supervised tuning of Gemini models, and Bedrock offers fine-tuning, continued pretraining and distillation for selected models. Model lists change often, so check each provider's current documentation.

The self-hosted tools

  • TRL + PEFT (Hugging Face) — the reference trainers: SFT, DPO, GRPO and more.
  • Unsloth — faster, lower-memory training on a single GPU, with a TRL-compatible interface.
  • Axolotl — one YAML file per run, good defaults, multi-GPU.
  • LLaMA-Factory — CLI and web UI over many models and methods.

How to choose

  • Data rules first. Client-confidential, health or financial data may not be allowed to leave your environment or the country. India's DPDP Act 2023 and client contracts both matter here.
  • Portability. If you cannot download the weights, you are renting behaviour. When the provider retires the base model, you must retrain on a new one.
  • Total cost, not training cost. Hosted APIs bill per training token: 5,000 examples × 800 tokens × 3 epochs = 12 million billed tokens. Fine-tuned models are also often priced higher per token than the base model at inference, which can dominate the bill. Self-hosting costs GPU-hours plus engineering time, and wins once fine-tuning is routine.
  • Method. Custom losses, unusual masking, new preference methods, or a model nobody hosts force you onto the open stack.

A real-life example

A Mumbai law firm wants a legal-clause classifier. To test the idea in one day, an engineer fine-tunes a hosted model on 300 clauses taken only from public, already-published contracts. It reaches a useful accuracy on a small test, which proves the task is learnable.

For the real system, the partners refuse to send client contracts to any third party — they are confidential and privileged. The firm runs QLoRA on an 8B open model on one in-house GPU with TRL, trained on 6,000 internal clauses. The adapter, data and logs never leave the office, the firm owns the weights, and when a newer base model appears it can retrain on its own schedule. The hosted API was the right tool for the one-day test, and the wrong one for production.

Follow-up questions to expect

  • "Can you download a model fine-tuned through OpenAI's API?" — No. You get a private model ID to call through the API.
  • "What happens when a hosted base model is retired?" — Your fine-tuned model is eventually retired with it, and you must retrain on a newer base. Keep your training data and eval set ready.
  • "Unsloth or plain TRL?" — Unsloth for fast single-GPU runs on supported models; plain TRL (with Accelerate) when you need multi-GPU setups or methods Unsloth does not cover.