Fine-Tuning LLMs with LoRA, QLoRA and PEFT

Project Structure and Configuration Management


Three weeks into a project, a model in production starts producing subtly worse summaries than the one from the previous Friday. Someone asks the obvious question: what changed? The training notebook has been edited forty times since. Cell 12 says r=32 but the comment above it says 16. Cell 7, which loaded the dataset, was re-run after cell 9 at some point. The learning rate appears in three places with two different values. Nobody can say which combination produced the good model, and nobody can reproduce it.

The model is not the problem. The problem is that the experiment was never written down in a form a computer could re-execute. A notebook records what you did in an order you cannot recover; a project records what you decided in a form you can re-run.

Moving from one to the other is the difference between a demo and something a team can operate. Here is what that structure looks like and, more importantly, why each boundary is drawn where it is.

Four modules, one config file, no notebookconfig.yaml — every hyperparameter as datadata_loader.py — records to formatted textdata_processor.py — text to masked tensorsmodel_setup.py — quantise, then attach adaptertrainer.py and main.py — run it, log the config
When the run is a file rather than forty edited cells, 'what changed since Friday' is a diff instead of an argument.

What notebooks make impossible

RequirementIn a notebookIn a project
Reproduce a run from six weeks agoGuesswork — cell state is invisibleCheck out the commit, run the config
Change one hyperparameterFind and edit every place it appearsOne line in one YAML file
Run 12 variants overnightTwelve notebook copiesA loop over twelve config files
Review a changeDiff of JSON with embedded outputsDiff of readable code
Test the data pipelineNot really possibleOrdinary unit tests
Run on a scheduler or CIFragileNormal script invocation
Two people working at onceMerge conflicts on every cellNormal git workflow

Notebooks are excellent for the first afternoon — inspecting data, checking a tokeniser, trying a generation. The moment you intend to run the same thing twice with a variation, move.

A run you cannot reproduce is not a result, it is an anecdote. Structure exists so that "what produced this model?" has an answer.

The layout

Text
finetune-project/  configs/    model.yaml           # base model, quantisation, LoRA    data.yaml            # paths, template, lengths, splits    training.yaml        # optimiser, schedule, batching, checkpoints    experiment.yaml      # run name, seed, tracking, output dir  src/    config.py            # load + validate + merge YAML into typed objects    data_loader.py       # raw source  ->  formatted text records    data_processor.py    # text records ->  tensors with masked labels    model_setup.py       # quantised base + adapter    trainer.py           # arguments, callbacks, the training call    evaluate.py          # metrics on held-out data    utils.py             # seeding, logging, device checks  scripts/    prepare_data.py    train.py    merge_adapter.py  tests/    test_data_processor.py    test_config.py  data/    raw/ processed/  outputs/    {run_name}/      config_snapshot.yaml      adapter/      checkpoints/      metrics.json      train.log  requirements.txt  README.md

Two conventions carry most of the value. First, nothing in src/ contains a literal hyperparameter — every number arrives from config. Second, every run writes a snapshot of its resolved configuration into its own output directory. That snapshot, not the YAML in your working tree, is what tells you six weeks later exactly what produced a given adapter.

Configuration as data, not code

Why not Python variables?

A file of module-level constants seems equivalent and is not:

  • You cannot serialise it. A YAML file can be copied into the output directory verbatim and logged to an experiment tracker. A Python module can contain arbitrary logic and there is no faithful way to record its resolved state.
  • You cannot generate it. A sweep that produces twelve configs is trivial with data files and awkward with modules.
  • It invites conditionals. Once config.py is executable, someone adds if os.environ.get("BIG"): LR = 1e-4, and now the effective learning rate depends on the shell that launched the job.
  • It is unreadable to non-Python tooling. Dashboards, schedulers and diff viewers all handle YAML.

The four files

Text
# configs/model.yamlbase_model: meta-llama/Llama-2-7b-hfquantization:  load_in_4bit: true  quant_type: nf4  double_quant: true  compute_dtype: bfloat16lora:  r: 16  alpha: 32  dropout: 0.05  bias: none  target_modules: [q_proj, k_proj, v_proj, o_proj,                   gate_proj, up_proj, down_proj]
Text
# configs/data.yamltrain_file: data/processed/train.jsonleval_file:  data/processed/eval.jsonlprompt_template: |  ### Instruction:  {instruction}  ### Input:  {input}  ### Response:max_length: 1024mask_prompt_loss: truedrop_over_length: true      # drop rather than truncate mid-answer
Text
# configs/training.yamlnum_train_epochs: 3per_device_train_batch_size: 4gradient_accumulation_steps: 4      # effective batch = 16gradient_checkpointing: truelearning_rate: 2.0e-4lr_scheduler_type: cosinewarmup_steps: 0.03                  # below 1 = fraction of all stepsweight_decay: 0.01max_grad_norm: 0.3optim: paged_adamw_8bitbf16: trueeval_strategy: stepssave_strategy: stepseval_steps: 50save_steps: 50save_total_limit: 3load_best_model_at_end: truemetric_for_best_model: eval_loss
Text
# configs/experiment.yamlrun_name: support-r16-lr2e4seed: 42output_dir: outputstracking:  backend: none        # none | wandb | mlflow  project: support-finetunenotes: "baseline: all seven modules, rank 16"

The split is by rate of change, not by topic. During a sweep you edit training.yaml and model.yaml constantly, data.yaml occasionally, and experiment.yaml once per run. Files that change together belong together.

Loading, with validation that fails early

Python
# src/config.pyfrom dataclasses import dataclass, field, asdictfrom pathlib import Pathimport yaml@dataclassclass LoraCfg:    r: int = 16    alpha: int = 32    dropout: float = 0.05    bias: str = "none"    target_modules: list = field(default_factory=list)    def __post_init__(self):        if self.r < 1 or self.r > 512:            raise ValueError(f"lora.r={self.r} outside sane range 1-512")        if not self.target_modules:            raise ValueError("lora.target_modules must not be empty")@dataclassclass TrainingCfg:    num_train_epochs: int = 3    per_device_train_batch_size: int = 4    gradient_accumulation_steps: int = 4    learning_rate: float = 2e-4    warmup_steps: float = 0.03    # ... remaining fields    def __post_init__(self):        if not (1e-6 <= self.learning_rate <= 1e-2):            raise ValueError(f"learning_rate={self.learning_rate} implausible")        if not (0.0 <= self.warmup_steps <= 0.5):            raise ValueError("warmup_steps must be a fraction between 0 and 0.5")    @property    def effective_batch_size(self):        return self.per_device_train_batch_size * self.gradient_accumulation_stepsdef load(config_dir="configs", overrides=None):    raw = {}    for name in ("model", "data", "training", "experiment"):        raw[name] = yaml.safe_load(Path(config_dir, f"{name}.yaml").read_text())    for dotted, value in (overrides or {}).items():   # e.g. "model.lora.r"        *path, key = dotted.split(".")        node = raw        for part in path:            node = node[part]        node[key] = value    LoraCfg(**raw["model"]["lora"])      # validate after overrides: raises on bad values    TrainingCfg(**raw["training"])    return raw

The validation is the point. A learning rate of 2e4 instead of 2e-4 is a single missing hyphen, entirely plausible in YAML, and without a range check it destroys the model over four hours of GPU time. Ten lines of __post_init__ catch it in ten milliseconds.

The effective_batch_size property matters too, because it is the number that governs training dynamics and it appears nowhere in the config. Deriving it once, in code, prevents someone halving the per-device batch to fit memory and unknowingly halving the effective batch as well.

Validate configuration at load time, not at use time. A run that crashes in the first second is a minor annoyance; a run that completes with a corrupted setting is a day gone.

The data pipeline: two jobs, two modules

The most common structural mistake is one data.py that reads files, formats prompts, tokenises, masks labels and batches. It becomes untestable, because you cannot exercise the tokenisation logic without touching disk. Split it at the boundary where the representation changes.

data_loader.py — getting formatted text records

Its job: read whatever the source is, normalise into a uniform record shape, apply the prompt template, and return plain text. No tokeniser is involved, so it is trivially testable.

Python
# src/data_loader.pyimport jsonfrom pathlib import Pathdef read_jsonl(path):    for line in Path(path).read_text().splitlines():        line = line.strip()        if line:            yield json.loads(line)def build_records(path, template, eos_token):    """Returns dicts with 'prompt' and 'full'. No tokenisation here."""    records = []    for row in read_jsonl(path):        if not row.get("output", "").strip():            continue                          # silently useless example        prompt = template.format(instruction=row["instruction"],                                 input=row.get("input", ""))        records.append({            "id": row.get("id"),            "prompt": prompt,            "full": prompt + row["output"] + eos_token,        })    return records

data_processor.py — turning text into tensors

Its job: tokenise, mask the prompt region, enforce length policy, and collate. It takes text in and returns tensors, so a unit test needs a tokeniser and nothing else.

Python
# src/data_processor.pydef encode(record, tokenizer, max_length, mask_prompt=True):    enc = tokenizer(record["full"], truncation=True, max_length=max_length)    labels = list(enc["input_ids"])    if mask_prompt:        n_prompt = len(tokenizer(record["prompt"], truncation=True,                                 max_length=max_length)["input_ids"])        labels[:n_prompt] = [-100] * n_prompt    enc["labels"] = labels    return encdef length_report(records, tokenizer):    lens = sorted(len(tokenizer(r["full"])["input_ids"]) for r in records)    n = len(lens)    return {"p50": lens[n // 2], "p90": lens[int(n * 0.9)],            "p99": lens[int(n * 0.99)], "max": lens[-1]}

Now a test is straightforward, and it checks the thing that silently breaks runs:

Python
# tests/test_data_processor.pydef test_prompt_tokens_are_masked(tokenizer):    rec = {"prompt": "### Instruction:\nSummarise.\n\n### Response:\n",           "full":   "### Instruction:\nSummarise.\n\n### Response:\nA short summary.</s>"}    enc = encode(rec, tokenizer, max_length=128)    n_prompt = len(tokenizer(rec["prompt"])["input_ids"])    assert all(l == -100 for l in enc["labels"][:n_prompt])    assert any(l != -100 for l in enc["labels"][n_prompt:])    assert len(enc["labels"]) == len(enc["input_ids"])

model_setup.py — quantise, then adapt, in that order

Python
# src/model_setup.pyimport torchfrom transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfigfrom peft import LoraConfig, get_peft_model, prepare_model_for_kbit_trainingDTYPES = {"bfloat16": torch.bfloat16, "float16": torch.float16}def build(cfg):    m = cfg["model"]    tokenizer = AutoTokenizer.from_pretrained(m["base_model"])    if tokenizer.pad_token is None:        tokenizer.pad_token = tokenizer.eos_token    tokenizer.padding_side = "right"    q = m["quantization"]    bnb = BitsAndBytesConfig(        load_in_4bit=q["load_in_4bit"],        bnb_4bit_quant_type=q["quant_type"],        bnb_4bit_use_double_quant=q["double_quant"],        bnb_4bit_compute_dtype=DTYPES[q["compute_dtype"]],    )    model = AutoModelForCausalLM.from_pretrained(        m["base_model"], quantization_config=bnb, device_map="auto")    model.config.use_cache = False    model = prepare_model_for_kbit_training(model, use_gradient_checkpointing=True)    l = m["lora"]    model = get_peft_model(model, LoraConfig(        r=l["r"], lora_alpha=l["alpha"], lora_dropout=l["dropout"],        bias=l["bias"], target_modules=l["target_modules"],        task_type="CAUSAL_LM"))    trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)    if trainable == 0:        raise RuntimeError("no trainable parameters - check target_modules names")    print(f"trainable: {trainable:,}")    return model, tokenizer

The order is not stylistic. Quantisation replaces linear layers with 4-bit modules; adapters must attach to those replacements. Attach the adapter first and you get adapters pointing at layers that are about to be swapped out. The trainable == 0 guard is cheap insurance against the most expensive silent failure in the pipeline: training a model with nothing to train. Current PEFT already raises an error when no name in target_modules matches, but one misspelt name among seven is skipped without a word, so for full protection compare trainable against the count you expect from r(d+k)r(d+k).

trainer.py and train.py

Python
# src/trainer.pyfrom transformers import Trainer, TrainingArguments, DataCollatorForSeq2Seqfrom transformers.trainer_callback import TrainerCallbackclass ThroughputCallback(TrainerCallback):    def on_log(self, args, state, control, logs=None, **kwargs):        if logs and "loss" in logs and state.global_step > 0:            steps_left = state.max_steps - state.global_step            print(f"step {state.global_step}/{state.max_steps} "                  f"loss={logs['loss']:.4f} remaining={steps_left}")def build_trainer(model, tokenizer, train_ds, eval_ds, cfg, out_dir):    t = cfg["training"]    args = TrainingArguments(output_dir=str(out_dir), report_to="none", **t)    return Trainer(        model=model, args=args,        train_dataset=train_ds, eval_dataset=eval_ds,        data_collator=DataCollatorForSeq2Seq(            tokenizer, padding=True, label_pad_token_id=-100),        callbacks=[ThroughputCallback()],    )
Python
# scripts/train.pyimport json, argparsefrom pathlib import Pathimport yamlfrom datasets import Datasetfrom src import config, data_loader, data_processor, model_setup, trainer, utilsdef parse_value(v):    for cast in (int, float):       # YAML reads "1e-4" as a string, so try these first        try:            return cast(v)        except ValueError:            pass    return yaml.safe_load(v)        # true/false, lists, plain stringsdef main():    ap = argparse.ArgumentParser()    ap.add_argument("--config-dir", default="configs")    ap.add_argument("--set", action="append", default=[],                    help="override, e.g. --set training.learning_rate=1e-4")    a = ap.parse_args()    overrides = {}    for item in a.set:        k, v = item.split("=", 1)        overrides[k] = parse_value(v)    cfg = config.load(a.config_dir, overrides)    utils.set_seed(cfg["experiment"]["seed"])    out = Path(cfg["experiment"]["output_dir"]) / cfg["experiment"]["run_name"]    out.mkdir(parents=True, exist_ok=True)    # snapshot the RESOLVED config, overrides included    (out / "config_snapshot.yaml").write_text(yaml.safe_dump(cfg))    model, tokenizer = model_setup.build(cfg)    d = cfg["data"]    train_recs = data_loader.build_records(d["train_file"], d["prompt_template"],                                           tokenizer.eos_token)    eval_recs  = data_loader.build_records(d["eval_file"], d["prompt_template"],                                           tokenizer.eos_token)    print("token lengths:", data_processor.length_report(train_recs, tokenizer))    enc = lambda r: data_processor.encode(r, tokenizer, d["max_length"],                                          d["mask_prompt_loss"])    train_ds = Dataset.from_list([enc(r) for r in train_recs])    eval_ds  = Dataset.from_list([enc(r) for r in eval_recs])    tr = trainer.build_trainer(model, tokenizer, train_ds, eval_ds, cfg, out)    result = tr.train()    tr.save_model(out / "adapter")    tokenizer.save_pretrained(out / "adapter")    (out / "metrics.json").write_text(json.dumps(result.metrics, indent=2))if __name__ == "__main__":    main()

Notice that scripts/train.py contains no decisions. It reads config, calls modules in order, and writes artefacts. Every judgement lives either in a config file (values) or in a module (logic). That is what makes a sweep a shell loop:

Bash
for r in 8 16 32 64; do  python scripts/train.py \    --set model.lora.r=$r \    --set model.lora.alpha=$((r * 2)) \    --set experiment.run_name=sweep-r$rdone

Pre-flight checks

Everything that can be verified in seconds should be verified before you commit hours of GPU time:

Python
# src/utils.pyimport os, random, torch, numpy as npdef set_seed(seed):    random.seed(seed); np.random.seed(seed); torch.manual_seed(seed)    torch.cuda.manual_seed_all(seed)    os.environ["PYTHONHASHSEED"] = str(seed)def preflight(cfg, model, tokenizer, train_ds):    problems = []    if not torch.cuda.is_available():        problems.append("no CUDA device")    elif cfg["training"].get("bf16") and not torch.cuda.is_bf16_supported():        problems.append("bf16 requested but unsupported - use fp16")    trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)    if trainable == 0:        problems.append("zero trainable parameters")    sample = train_ds[0]    if all(l == -100 for l in sample["labels"]):        problems.append("all labels masked - nothing to learn from")    if tokenizer.eos_token_id not in sample["input_ids"]:        problems.append("no EOS token in the first example")    free = torch.cuda.mem_get_info()[0] / 1e9 if torch.cuda.is_available() else 0    print(f"free GPU memory: {free:.1f} GB; trainable params: {trainable:,}")    if problems:        raise RuntimeError("pre-flight failed:\n  - " + "\n  - ".join(problems))

Those four checks correspond to the four ways runs most often waste a day: wrong hardware assumption, adapters not attached, labels fully masked, and no stop token. All are invisible in a loss curve and all are one line to detect.

Beyond that, three production habits pay for themselves quickly. Log the git commit hash into metrics.json so a checkpoint maps to source. Pin exact dependency versions, because a minor transformers release changing a default is a real and recurring cause of "the same config produced a different model". And write the resolved config snapshot before training starts, not after — a crashed run is exactly the one you most want to inspect.

Where this goes wrong

Config that is really code. Once a YAML file contains learning_rate: base_lr times 2 or your loader executes expressions, you have reinvented Python badly and lost serialisability. Keep values literal; compute derived quantities in typed config objects.

Snapshotting the wrong thing. Copying configs/ into the output directory records the files, not the run — command-line overrides are missing. Serialise the merged, post-validation object.

Over-structuring on day one. Eleven modules and an abstract base class for a single experiment is its own failure. Start with config.py, data.py, train.py; split when a file genuinely does two jobs.

Untracked data. A perfectly versioned config pointing at data/train.jsonl, which someone regenerated on Tuesday, is not reproducible. Record a content hash of the dataset in the run metadata.

What this means for the next project you start

The concrete test of whether your structure is doing its job: pick a model you trained two weeks ago and try to reproduce it from its output directory alone. You should need exactly the config snapshot, the git commit, and the dataset hash. If reproducing it requires asking a colleague what they remember, the structure has a gap, and the gap is almost always a value that lived in code instead of in config.

Build this before the first serious run, not after the first painful one. It is roughly two hours of work and it converts every subsequent experiment from a manual procedure into a command you can put in a loop — which is what turns "we tried rank 16" into "we swept rank across eight values overnight and here is the curve".