Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

Why is LoRA attractive for enterprise rollouts (cost, safety, iteration speed)?


What you need to know

Cost, with numbers

Full fine-tuning with AdamW in mixed precision needs about 16 bytes per parameter (weights, gradients, a 32-bit master copy and two Adam states). For 8B parameters that is about 128 GB before activations, so you need several 80 GB GPUs and sharding (FSDP or DeepSpeed ZeRO).

LoRA freezes the base (16 GB in bf16) and trains about 42M adapter parameters, whose gradients and optimizer states take under 1 GB. It fits on one GPU.

A worked estimate for one run:

Text
20,000 examples x 600 tokens x 3 epochs = 36M training tokensassumed throughput: 3,000 tokens/s on one H100 (measure yours)36,000,000 / 3,000 = 12,000 s = about 3.3 GPU-hoursat an assumed 3 dollars per GPU-hour = about 10 dollars

The exact numbers depend on sequence length, packing and hardware, but the order of magnitude — hours and tens of dollars, not days and thousands — is the point.

Iteration speed

When a data change becomes an evaluated model in an afternoon, teams actually fix problems found in production each week. This is the benefit teams most underestimate.

Safety and reversibility

  • The base weights never change, so the damage from a bad adapter is limited to that adapter. Rollback means pointing the router at the previous version.
  • Each adapter has clear lineage: which base, which dataset version, which config, which eval scores. Store this in a model registry.

Compliance and data isolation

Adapters trained on client A's data can be stored separately and deleted when a contract ends or a person asks for erasure — India's DPDP Act 2023 gives people a right to have their data erased. Deleting one adapter is a much simpler story for an auditor than "their data is somewhere in our shared model".

Fleet economics

One base model in memory can serve many adapters (see the multi-LoRA serving lesson). Fifty tenants at about 84 MB each is about 4 GB of adapters, not fifty model deployments.

The honest counterweights

  • Adapters are pinned to one base checkpoint. A base upgrade means retraining the whole fleet.
  • Unmerged serving adds a little latency.
  • LoRA learns large new knowledge less well than full fine-tuning (the 2024 paper "LoRA Learns Less and Forgets Less" found this for code and maths), though it also forgets less.

A real-life example

A hospital chain in Hyderabad runs a medical-report summariser. Patient data may not leave its own servers, so it trains on-premises on two GPUs. It keeps one adapter per department — radiology, pathology, cardiology — on one shared 8B base.

When radiology-v4 goes live, doctors report that summaries sometimes leave out the "Impression" line. The platform team points the radiology route back to radiology-v3 within five minutes; no other department is affected. The data team finds that 300 new training reports came from a template without an Impression section, fixes them, retrains overnight for a few GPU-hours, and ships radiology-v5 after it passes the eval gate. With a single fully fine-tuned model, the same problem would have meant rolling back every department.

Follow-up questions to expect

  • "When is LoRA not enough?" — For large shifts in knowledge or language, such as teaching a new language, continual pretraining or full fine-tuning usually does better.
  • "How do you govern fifty adapters?" — A registry with owner, base revision, dataset version and eval scores for each adapter, plus an automatic eval gate before any adapter can be served.
  • "Does LoRA make inference cheaper?" — No. Each request still runs the full base model. LoRA saves training cost and the number of deployments.