Course Content
Fine-Tuning LLMs with LoRA, QLoRA and PEFT
4 sections · 10 lessons
Project Structure and Configuration Management
Three weeks into a project, a model in production starts producing subtly worse summaries than the one from the previous Friday. Someone asks the obvious question: what changed? The training notebook has been edited forty times since. Cell 12 says r=32 but the comment above it says 16. Cell 7, which loaded the dataset, was re-run after cell 9 at some point. The learning rate appears in three places with two different values. Nobody can say which combination produced the good model, and nobody can reproduce it.
The model is not the problem. The problem is that the experiment was never written down in a form a computer could re-execute. A notebook records what you did in an order you cannot recover; a project records what you decided in a form you can re-run.
Moving from one to the other is the difference between a demo and something a team can operate. Here is what that structure looks like and, more importantly, why each boundary is drawn where it is.
What notebooks make impossible
| Requirement | In a notebook | In a project |
|---|---|---|
| Reproduce a run from six weeks ago | Guesswork — cell state is invisible | Check out the commit, run the config |
| Change one hyperparameter | Find and edit every place it appears | One line in one YAML file |
| Run 12 variants overnight | Twelve notebook copies | A loop over twelve config files |
| Review a change | Diff of JSON with embedded outputs | Diff of readable code |
| Test the data pipeline | Not really possible | Ordinary unit tests |
| Run on a scheduler or CI | Fragile | Normal script invocation |
| Two people working at once | Merge conflicts on every cell | Normal git workflow |
Notebooks are excellent for the first afternoon — inspecting data, checking a tokeniser, trying a generation. The moment you intend to run the same thing twice with a variation, move.
A run you cannot reproduce is not a result, it is an anecdote. Structure exists so that "what produced this model?" has an answer.
The layout
finetune-project/ configs/ model.yaml # base model, quantisation, LoRA data.yaml # paths, template, lengths, splits training.yaml # optimiser, schedule, batching, checkpoints experiment.yaml # run name, seed, tracking, output dir src/ config.py # load + validate + merge YAML into typed objects data_loader.py # raw source -> formatted text records data_processor.py # text records -> tensors with masked labels model_setup.py # quantised base + adapter trainer.py # arguments, callbacks, the training call evaluate.py # metrics on held-out data utils.py # seeding, logging, device checks scripts/ prepare_data.py train.py merge_adapter.py tests/ test_data_processor.py test_config.py data/ raw/ processed/ outputs/ {run_name}/ config_snapshot.yaml adapter/ checkpoints/ metrics.json train.log requirements.txt README.mdTwo conventions carry most of the value. First, nothing in src/ contains a literal hyperparameter — every number arrives from config. Second, every run writes a snapshot of its resolved configuration into its own output directory. That snapshot, not the YAML in your working tree, is what tells you six weeks later exactly what produced a given adapter.
Configuration as data, not code
Why not Python variables?
A file of module-level constants seems equivalent and is not:
- You cannot serialise it. A YAML file can be copied into the output directory verbatim and logged to an experiment tracker. A Python module can contain arbitrary logic and there is no faithful way to record its resolved state.
- You cannot generate it. A sweep that produces twelve configs is trivial with data files and awkward with modules.
- It invites conditionals. Once
config.pyis executable, someone addsif os.environ.get("BIG"): LR = 1e-4, and now the effective learning rate depends on the shell that launched the job. - It is unreadable to non-Python tooling. Dashboards, schedulers and diff viewers all handle YAML.
The four files
# configs/model.yamlbase_model: meta-llama/Llama-2-7b-hfquantization: load_in_4bit: true quant_type: nf4 double_quant: true compute_dtype: bfloat16lora: r: 16 alpha: 32 dropout: 0.05 bias: none target_modules: [q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj]# configs/data.yamltrain_file: data/processed/train.jsonleval_file: data/processed/eval.jsonlprompt_template: | ### Instruction: {instruction} ### Input: {input} ### Response:max_length: 1024mask_prompt_loss: truedrop_over_length: true # drop rather than truncate mid-answer# configs/training.yamlnum_train_epochs: 3per_device_train_batch_size: 4gradient_accumulation_steps: 4 # effective batch = 16gradient_checkpointing: truelearning_rate: 2.0e-4lr_scheduler_type: cosinewarmup_steps: 0.03 # below 1 = fraction of all stepsweight_decay: 0.01max_grad_norm: 0.3optim: paged_adamw_8bitbf16: trueeval_strategy: stepssave_strategy: stepseval_steps: 50save_steps: 50save_total_limit: 3load_best_model_at_end: truemetric_for_best_model: eval_loss# configs/experiment.yamlrun_name: support-r16-lr2e4seed: 42output_dir: outputstracking: backend: none # none | wandb | mlflow project: support-finetunenotes: "baseline: all seven modules, rank 16"The split is by rate of change, not by topic. During a sweep you edit training.yaml and model.yaml constantly, data.yaml occasionally, and experiment.yaml once per run. Files that change together belong together.
Loading, with validation that fails early
1# src/config.py2from dataclasses import dataclass, field, asdict3from pathlib import Path4import yaml56@dataclass7class LoraCfg:8 r: int = 169 alpha: int = 3210 dropout: float = 0.0511 bias: str = "none"12 target_modules: list = field(default_factory=list)1314 def __post_init__(self):15 if self.r < 1 or self.r > 512:16 raise ValueError(f"lora.r={self.r} outside sane range 1-512")17 if not self.target_modules:18 raise ValueError("lora.target_modules must not be empty")1920@dataclass21class TrainingCfg:22 num_train_epochs: int = 323 per_device_train_batch_size: int = 424 gradient_accumulation_steps: int = 425 learning_rate: float = 2e-426 warmup_steps: float = 0.0327 # ... remaining fields2829 def __post_init__(self):30 if not (1e-6 <= self.learning_rate <= 1e-2):31 raise ValueError(f"learning_rate={self.learning_rate} implausible")32 if not (0.0 <= self.warmup_steps <= 0.5):33 raise ValueError("warmup_steps must be a fraction between 0 and 0.5")3435 @property36 def effective_batch_size(self):37 return self.per_device_train_batch_size * self.gradient_accumulation_steps3839def load(config_dir="configs", overrides=None):40 raw = {}41 for name in ("model", "data", "training", "experiment"):42 raw[name] = yaml.safe_load(Path(config_dir, f"{name}.yaml").read_text())4344 for dotted, value in (overrides or {}).items(): # e.g. "model.lora.r"45 *path, key = dotted.split(".")46 node = raw47 for part in path:48 node = node[part]49 node[key] = value5051 LoraCfg(**raw["model"]["lora"]) # validate after overrides: raises on bad values52 TrainingCfg(**raw["training"])53 return rawThe validation is the point. A learning rate of 2e4 instead of 2e-4 is a single missing hyphen, entirely plausible in YAML, and without a range check it destroys the model over four hours of GPU time. Ten lines of __post_init__ catch it in ten milliseconds.
The effective_batch_size property matters too, because it is the number that governs training dynamics and it appears nowhere in the config. Deriving it once, in code, prevents someone halving the per-device batch to fit memory and unknowingly halving the effective batch as well.
Validate configuration at load time, not at use time. A run that crashes in the first second is a minor annoyance; a run that completes with a corrupted setting is a day gone.
The data pipeline: two jobs, two modules
The most common structural mistake is one data.py that reads files, formats prompts, tokenises, masks labels and batches. It becomes untestable, because you cannot exercise the tokenisation logic without touching disk. Split it at the boundary where the representation changes.
data_loader.py — getting formatted text records
Its job: read whatever the source is, normalise into a uniform record shape, apply the prompt template, and return plain text. No tokeniser is involved, so it is trivially testable.
1# src/data_loader.py2import json3from pathlib import Path45def read_jsonl(path):6 for line in Path(path).read_text().splitlines():7 line = line.strip()8 if line:9 yield json.loads(line)1011def build_records(path, template, eos_token):12 """Returns dicts with 'prompt' and 'full'. No tokenisation here."""13 records = []14 for row in read_jsonl(path):15 if not row.get("output", "").strip():16 continue # silently useless example17 prompt = template.format(instruction=row["instruction"],18 input=row.get("input", ""))19 records.append({20 "id": row.get("id"),21 "prompt": prompt,22 "full": prompt + row["output"] + eos_token,23 })24 return recordsdata_processor.py — turning text into tensors
Its job: tokenise, mask the prompt region, enforce length policy, and collate. It takes text in and returns tensors, so a unit test needs a tokeniser and nothing else.
1# src/data_processor.py2def encode(record, tokenizer, max_length, mask_prompt=True):3 enc = tokenizer(record["full"], truncation=True, max_length=max_length)4 labels = list(enc["input_ids"])56 if mask_prompt:7 n_prompt = len(tokenizer(record["prompt"], truncation=True,8 max_length=max_length)["input_ids"])9 labels[:n_prompt] = [-100] * n_prompt1011 enc["labels"] = labels12 return enc1314def length_report(records, tokenizer):15 lens = sorted(len(tokenizer(r["full"])["input_ids"]) for r in records)16 n = len(lens)17 return {"p50": lens[n // 2], "p90": lens[int(n * 0.9)],18 "p99": lens[int(n * 0.99)], "max": lens[-1]}Now a test is straightforward, and it checks the thing that silently breaks runs:
1# tests/test_data_processor.py2def test_prompt_tokens_are_masked(tokenizer):3 rec = {"prompt": "### Instruction:\nSummarise.\n\n### Response:\n",4 "full": "### Instruction:\nSummarise.\n\n### Response:\nA short summary.</s>"}5 enc = encode(rec, tokenizer, max_length=128)67 n_prompt = len(tokenizer(rec["prompt"])["input_ids"])8 assert all(l == -100 for l in enc["labels"][:n_prompt])9 assert any(l != -100 for l in enc["labels"][n_prompt:])10 assert len(enc["labels"]) == len(enc["input_ids"])model_setup.py — quantise, then adapt, in that order
1# src/model_setup.py2import torch3from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig4from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training56DTYPES = {"bfloat16": torch.bfloat16, "float16": torch.float16}78def build(cfg):9 m = cfg["model"]10 tokenizer = AutoTokenizer.from_pretrained(m["base_model"])11 if tokenizer.pad_token is None:12 tokenizer.pad_token = tokenizer.eos_token13 tokenizer.padding_side = "right"1415 q = m["quantization"]16 bnb = BitsAndBytesConfig(17 load_in_4bit=q["load_in_4bit"],18 bnb_4bit_quant_type=q["quant_type"],19 bnb_4bit_use_double_quant=q["double_quant"],20 bnb_4bit_compute_dtype=DTYPES[q["compute_dtype"]],21 )2223 model = AutoModelForCausalLM.from_pretrained(24 m["base_model"], quantization_config=bnb, device_map="auto")25 model.config.use_cache = False26 model = prepare_model_for_kbit_training(model, use_gradient_checkpointing=True)2728 l = m["lora"]29 model = get_peft_model(model, LoraConfig(30 r=l["r"], lora_alpha=l["alpha"], lora_dropout=l["dropout"],31 bias=l["bias"], target_modules=l["target_modules"],32 task_type="CAUSAL_LM"))3334 trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)35 if trainable == 0:36 raise RuntimeError("no trainable parameters - check target_modules names")37 print(f"trainable: {trainable:,}")3839 return model, tokenizerThe order is not stylistic. Quantisation replaces linear layers with 4-bit modules; adapters must attach to those replacements. Attach the adapter first and you get adapters pointing at layers that are about to be swapped out. The trainable == 0 guard is cheap insurance against the most expensive silent failure in the pipeline: training a model with nothing to train. Current PEFT already raises an error when no name in target_modules matches, but one misspelt name among seven is skipped without a word, so for full protection compare trainable against the count you expect from r(d+k).
trainer.py and train.py
1# src/trainer.py2from transformers import Trainer, TrainingArguments, DataCollatorForSeq2Seq3from transformers.trainer_callback import TrainerCallback45class ThroughputCallback(TrainerCallback):6 def on_log(self, args, state, control, logs=None, **kwargs):7 if logs and "loss" in logs and state.global_step > 0:8 steps_left = state.max_steps - state.global_step9 print(f"step {state.global_step}/{state.max_steps} "10 f"loss={logs['loss']:.4f} remaining={steps_left}")1112def build_trainer(model, tokenizer, train_ds, eval_ds, cfg, out_dir):13 t = cfg["training"]14 args = TrainingArguments(output_dir=str(out_dir), report_to="none", **t)15 return Trainer(16 model=model, args=args,17 train_dataset=train_ds, eval_dataset=eval_ds,18 data_collator=DataCollatorForSeq2Seq(19 tokenizer, padding=True, label_pad_token_id=-100),20 callbacks=[ThroughputCallback()],21 )1# scripts/train.py2import json, argparse3from pathlib import Path4import yaml5from datasets import Dataset67from src import config, data_loader, data_processor, model_setup, trainer, utils89def parse_value(v):10 for cast in (int, float): # YAML reads "1e-4" as a string, so try these first11 try:12 return cast(v)13 except ValueError:14 pass15 return yaml.safe_load(v) # true/false, lists, plain strings1617def main():18 ap = argparse.ArgumentParser()19 ap.add_argument("--config-dir", default="configs")20 ap.add_argument("--set", action="append", default=[],21 help="override, e.g. --set training.learning_rate=1e-4")22 a = ap.parse_args()2324 overrides = {}25 for item in a.set:26 k, v = item.split("=", 1)27 overrides[k] = parse_value(v)2829 cfg = config.load(a.config_dir, overrides)30 utils.set_seed(cfg["experiment"]["seed"])3132 out = Path(cfg["experiment"]["output_dir"]) / cfg["experiment"]["run_name"]33 out.mkdir(parents=True, exist_ok=True)34 # snapshot the RESOLVED config, overrides included35 (out / "config_snapshot.yaml").write_text(yaml.safe_dump(cfg))3637 model, tokenizer = model_setup.build(cfg)3839 d = cfg["data"]40 train_recs = data_loader.build_records(d["train_file"], d["prompt_template"],41 tokenizer.eos_token)42 eval_recs = data_loader.build_records(d["eval_file"], d["prompt_template"],43 tokenizer.eos_token)44 print("token lengths:", data_processor.length_report(train_recs, tokenizer))4546 enc = lambda r: data_processor.encode(r, tokenizer, d["max_length"],47 d["mask_prompt_loss"])48 train_ds = Dataset.from_list([enc(r) for r in train_recs])49 eval_ds = Dataset.from_list([enc(r) for r in eval_recs])5051 tr = trainer.build_trainer(model, tokenizer, train_ds, eval_ds, cfg, out)52 result = tr.train()5354 tr.save_model(out / "adapter")55 tokenizer.save_pretrained(out / "adapter")56 (out / "metrics.json").write_text(json.dumps(result.metrics, indent=2))5758if __name__ == "__main__":59 main()Notice that scripts/train.py contains no decisions. It reads config, calls modules in order, and writes artefacts. Every judgement lives either in a config file (values) or in a module (logic). That is what makes a sweep a shell loop:
1for r in 8 16 32 64; do2 python scripts/train.py \3 --set model.lora.r=$r \4 --set model.lora.alpha=$((r * 2)) \5 --set experiment.run_name=sweep-r$r6donePre-flight checks
Everything that can be verified in seconds should be verified before you commit hours of GPU time:
1# src/utils.py2import os, random, torch, numpy as np34def set_seed(seed):5 random.seed(seed); np.random.seed(seed); torch.manual_seed(seed)6 torch.cuda.manual_seed_all(seed)7 os.environ["PYTHONHASHSEED"] = str(seed)89def preflight(cfg, model, tokenizer, train_ds):10 problems = []1112 if not torch.cuda.is_available():13 problems.append("no CUDA device")14 elif cfg["training"].get("bf16") and not torch.cuda.is_bf16_supported():15 problems.append("bf16 requested but unsupported - use fp16")1617 trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)18 if trainable == 0:19 problems.append("zero trainable parameters")2021 sample = train_ds[0]22 if all(l == -100 for l in sample["labels"]):23 problems.append("all labels masked - nothing to learn from")24 if tokenizer.eos_token_id not in sample["input_ids"]:25 problems.append("no EOS token in the first example")2627 free = torch.cuda.mem_get_info()[0] / 1e9 if torch.cuda.is_available() else 028 print(f"free GPU memory: {free:.1f} GB; trainable params: {trainable:,}")2930 if problems:31 raise RuntimeError("pre-flight failed:\n - " + "\n - ".join(problems))Those four checks correspond to the four ways runs most often waste a day: wrong hardware assumption, adapters not attached, labels fully masked, and no stop token. All are invisible in a loss curve and all are one line to detect.
Beyond that, three production habits pay for themselves quickly. Log the git commit hash into metrics.json so a checkpoint maps to source. Pin exact dependency versions, because a minor transformers release changing a default is a real and recurring cause of "the same config produced a different model". And write the resolved config snapshot before training starts, not after — a crashed run is exactly the one you most want to inspect.
Where this goes wrong
Config that is really code. Once a YAML file contains learning_rate: base_lr times 2 or your loader executes expressions, you have reinvented Python badly and lost serialisability. Keep values literal; compute derived quantities in typed config objects.
Snapshotting the wrong thing. Copying configs/ into the output directory records the files, not the run — command-line overrides are missing. Serialise the merged, post-validation object.
Over-structuring on day one. Eleven modules and an abstract base class for a single experiment is its own failure. Start with config.py, data.py, train.py; split when a file genuinely does two jobs.
Untracked data. A perfectly versioned config pointing at data/train.jsonl, which someone regenerated on Tuesday, is not reproducible. Record a content hash of the dataset in the run metadata.
What this means for the next project you start
The concrete test of whether your structure is doing its job: pick a model you trained two weeks ago and try to reproduce it from its output directory alone. You should need exactly the config snapshot, the git commit, and the dataset hash. If reproducing it requires asking a colleague what they remember, the structure has a gap, and the gap is almost always a value that lived in code instead of in config.
Build this before the first serious run, not after the first painful one. It is roughly two hours of work and it converts every subsequent experiment from a manual procedure into a command you can put in a loop — which is what turns "we tried rank 16" into "we swept rank across eight values overnight and here is the curve".