Parameter-efficient fine-tuning (PEFT) methods adapt a pretrained large model to a new task by training only a small fraction of its parameters rather than the full model. The most widely used method is LoRA (Low-Rank Adaptation), which freezes the original weights and adds trainable low-rank matrices to specific layers. Hugging Face’s PEFT library provides a unified interface for LoRA, QLoRA, and other PEFT methods, integrating directly with Transformers and the training ecosystem. This guide covers the full workflow: LoRA configuration, training, saving, loading, and merging adapters.
Installation
pip install peft transformers accelerate bitsandbytes trl
How LoRA Works
LoRA adds two small trainable matrices A and B to each targeted weight matrix W. Instead of updating W directly, it computes W + BA during the forward pass, where A is (rank × input_dim) and B is (output_dim × rank). With rank r=16 on a 4096×4096 weight matrix, the adapter adds 2 × 4096 × 16 = 131,072 parameters instead of 4096² = 16.7M. Training only these adapter parameters instead of W reduces memory for gradients and optimiser state by 100× or more. The pretrained weights W stay frozen — only A and B receive gradients — so the base model’s knowledge is preserved and the adapter captures the task-specific change.
Basic LoRA Fine-Tuning
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments
from peft import LoraConfig, get_peft_model, TaskType
from trl import SFTTrainer
import torch
model_name = "meta-llama/Llama-3.1-8B"
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16, device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token
lora_config = LoraConfig(
task_type=TaskType.CAUSAL_LM,
r=16, # rank — higher = more capacity, more memory
lora_alpha=32, # scaling factor; typically 2× rank
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"], # all linear layers
lora_dropout=0.05,
bias="none",
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# trainable: ~83M / 8B total = ~1%
training_args = TrainingArguments(
output_dir="./lora-output",
num_train_epochs=3,
per_device_train_batch_size=2,
gradient_accumulation_steps=8,
learning_rate=2e-4,
bf16=True,
logging_steps=10,
save_strategy="epoch",
)
trainer = SFTTrainer(
model=model,
args=training_args,
train_dataset=train_dataset,
tokenizer=tokenizer,
dataset_text_field="text",
max_seq_length=2048,
)
trainer.train()
Choosing LoRA Rank and Target Modules
Rank (r) controls the adapter’s capacity. r=4 is minimal — sufficient for narrow task adaptation on simple tasks. r=16 is a solid default for most instruction fine-tuning. r=64 approaches full fine-tuning quality for tasks that require significant knowledge injection. Higher rank means more trainable parameters and slightly higher memory use, but also better adaptation capability. For target modules: targeting only attention layers (q_proj, v_proj) is the minimum effective configuration; targeting all linear layers (attention + MLP) gives better quality at the cost of more adapter parameters. A practical heuristic: start with r=16 on all linear layers and reduce if memory is tight.
QLoRA: 4-bit Base + LoRA Adapters
from transformers import BitsAndBytesConfig
from peft import prepare_model_for_kbit_training
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
)
model = AutoModelForCausalLM.from_pretrained(model_name,
quantization_config=bnb_config, device_map="auto")
# prepare_model_for_kbit_training is required before adding LoRA to a quantised model
model = prepare_model_for_kbit_training(model)
model = get_peft_model(model, lora_config)
Saving and Loading Adapters
from peft import PeftModel
# Save only the adapter — tiny file (~50-200MB for r=16)
model.save_pretrained("./adapter-checkpoint")
tokenizer.save_pretrained("./adapter-checkpoint")
# Load adapter on top of base model later
base = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "./adapter-checkpoint")
model.eval()
# Run inference
inputs = tokenizer("Summarize this article:", return_tensors="pt").to("cuda")
with torch.no_grad():
out = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(out[0], skip_special_tokens=True))
Merging Adapters into the Base Model
# Merge for deployment — creates a standard model with no PEFT dependency
merged = model.merge_and_unload()
merged.save_pretrained("./merged-model")
# Verify merged model works
merged_loaded = AutoModelForCausalLM.from_pretrained("./merged-model",
torch_dtype=torch.bfloat16, device_map="auto")
Multiple Adapters with Multiplexing
PEFT supports loading multiple adapters onto a single base model and switching between them at inference time — useful for serving many fine-tuned variants without storing multiple full models:
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "./adapter-summarisation", adapter_name="summarise")
model.load_adapter("./adapter-qa", adapter_name="qa")
model.load_adapter("./adapter-code", adapter_name="code")
# Switch adapter at inference time
model.set_adapter("summarise")
out1 = model.generate(**inputs_doc)
model.set_adapter("qa")
out2 = model.generate(**inputs_question)
This pattern is effective for multi-task serving where you want to minimise GPU memory — one base model copy plus several small adapters instead of several full model copies.
Figure 1 — LoRA: low-rank adapter added to a frozen pretrained weight matrix
Other PEFT Methods
The PEFT library supports several methods beyond LoRA. Prefix Tuning prepends trainable virtual tokens to the input at each layer — good for generation tasks but less flexible than LoRA. Prompt Tuning adds trainable embeddings only to the input layer — the smallest adapter, but less effective for large model-task gaps. IA³ (Infused Adapter by Inhibiting and Amplifying Inner Activations) scales activations with learned vectors rather than adding weight matrices — very parameter-efficient but less widely tested. In practice, LoRA covers the large majority of use cases and has the best library support, documentation, and benchmark comparisons. Use the alternatives only if LoRA is not meeting your efficiency target for a specific hardware constraint.
PEFT with LoRA is the standard method for adapting large language models to new tasks without full fine-tuning. Configure the LoraConfig, wrap the model with get_peft_model, train as normal, and save only the adapter. The adapter is typically 50–200MB regardless of base model size, can be swapped at runtime for multi-task serving, and merges cleanly into the base model for deployment. Paired with bitsandbytes 4-bit quantisation, it makes fine-tuning 7B–70B models accessible on hardware that previously could not even load them.
Figure 1 — LoRA adds two small trainable matrices to each frozen weight
Understanding lora_alpha and the Scaling Factor
The lora_alpha parameter controls how much the adapter output is scaled before being added to the frozen weight output. The effective scaling factor is lora_alpha / r — so with r=16 and lora_alpha=32, the adapter output is scaled by 2.0. Setting lora_alpha=r keeps the scale at 1.0; doubling lora_alpha to 2×r (the common default) doubles the adapter’s influence. In practice, the learning rate and lora_alpha interact: if you double the learning rate, you can halve lora_alpha and get similar convergence. A safe default is lora_alpha=2×r. If the model is underfitting (loss not decreasing), increase lora_alpha or learning rate; if it is unstable (loss spiking), decrease lora_alpha or add dropout.
Flash Attention Compatibility
PEFT LoRA adapters are fully compatible with FlashAttention-2 when using Hugging Face Transformers. Enable both by passing attn_implementation="flash_attention_2" to from_pretrained() before adding the LoRA config — the adapter wraps the attention projection matrices which are separate from the attention algorithm used:
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from peft import get_peft_model, LoraConfig, prepare_model_for_kbit_training
import torch
bnb_config = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
quantization_config=bnb_config,
attn_implementation="flash_attention_2", # FA2 + QLoRA
device_map="auto",
)
model = prepare_model_for_kbit_training(model)
model = get_peft_model(model, LoraConfig(r=16, lora_alpha=32,
target_modules=["q_proj","k_proj","v_proj","o_proj"], task_type="CAUSAL_LM"))
Monitoring Training with PEFT
The SFTTrainer from TRL logs training loss by default. Add a validation step to monitor generalisation during training, and use the EarlyStoppingCallback to stop when val loss plateaus:
from transformers import EarlyStoppingCallback
from trl import SFTTrainer
trainer = SFTTrainer(
model=model,
args=TrainingArguments(
evaluation_strategy="steps",
eval_steps=50,
load_best_model_at_end=True,
metric_for_best_model="eval_loss",
greater_is_better=False,
...
),
train_dataset=train_dataset,
eval_dataset=val_dataset,
callbacks=[EarlyStoppingCallback(early_stopping_patience=3)],
...
)
DoRA and LoRA Variants
DoRA (Weight-Decomposed Low-Rank Adaptation) decomposes each weight matrix into magnitude and direction components and applies LoRA only to the direction, keeping the magnitude trainable as a separate scalar. This improves performance on tasks where the magnitude of weight changes matters, such as reasoning tasks. DoRA is available in PEFT as a drop-in replacement for LoRA:
lora_config = LoraConfig(
r=16,
lora_alpha=32,
target_modules=["q_proj", "v_proj"],
use_dora=True, # DoRA instead of standard LoRA
)
model = get_peft_model(model, lora_config)
In practice, DoRA gives 1–3% quality improvement over LoRA at the same rank on instruction following and reasoning benchmarks, with minimal additional compute. It is worth enabling when you are pushing quality boundaries at a fixed adapter size. LoftQ is another variant that initialises LoRA weights to minimise quantisation error when used with 4-bit bitsandbytes — use it when QLoRA quality is noticeably below full fine-tuning on your task.
Evaluating Adapter Quality
Always evaluate on a held-out validation set before merging or deploying an adapter. A simple approach for instruction-tuned models is to compare outputs against a reference using a reward model or human evaluation; for task-specific fine-tunes (classification, NER, summarisation), use task-appropriate metrics:
model.eval()
correct = 0
with torch.no_grad():
for batch in val_loader:
inputs = tokenizer(batch['prompt'], return_tensors='pt', padding=True).to('cuda')
out = model.generate(**inputs, max_new_tokens=50)
preds = [tokenizer.decode(o, skip_special_tokens=True) for o in out]
correct += sum(p.strip() == r.strip() for p, r in zip(preds, batch['label']))
print(f"Accuracy: {correct/len(val_dataset):.3f}")
PEFT for Non-Language Models
The PEFT library is not limited to language models. It supports LoRA for vision transformers (ViT), stable diffusion U-Net weights (for fine-tuning image generation models), and sequence-to-sequence models like BART and T5. The configuration is the same — specify target_modules to match the architecture’s linear layer names. For vision transformers, target the attention projection layers (query, value); for diffusion models, target the attention layers in the U-Net’s cross-attention blocks. The adapter files are small regardless of model type, making PEFT a practical method for storing many task-specific adaptations of a shared base model in production.
Common PEFT Errors
A few issues that come up frequently. KeyError on target_modules — the module names vary by architecture; print [n for n, _ in model.named_modules()] to see the actual layer names, then set target_modules accordingly. AttributeError: ‘NoneType’ has no attribute ‘weight’ after merge_and_unload() — the model was in training mode; call model.eval() before merging. PEFT adapter not reducing trainable parameters — prepare_model_for_kbit_training() was not called before adding LoRA to a quantised model; this step is required to cast LayerNorm layers to float32 and enable input gradients. Inference much slower than base model — the adapter was not merged; call merge_and_unload() for production to eliminate the runtime adapter overhead.
PEFT with LoRA is the standard method for fine-tuning large models efficiently. Configure LoraConfig, wrap the model, train as usual, and save only the adapter. The adapter is typically 50–200MB regardless of base model size, can be swapped at runtime for multi-task serving, and merges cleanly into the base model for deployment. Combined with 4-bit bitsandbytes quantisation, it brings 7B–70B model fine-tuning within reach of single-GPU hardware that previously could not load these models at all.
PEFT with LoRA is the standard method for adapting large models to new tasks without full fine-tuning. Configure LoraConfig, wrap the model, train as usual, and save only the adapter — typically 50–200MB regardless of base model size. Combined with 4-bit bitsandbytes, it brings 7B–70B model fine-tuning within reach of single-GPU hardware that previously could not load these models at all.