bitsandbytes makes it possible to load large language models in 8-bit or 4-bit precision, slashing the GPU memory required without retraining. A 7B parameter model in float16 needs ~14GB VRAM; in 4-bit it needs ~4GB. This is the primary reason practitioners can now fine-tune LLaMA, Mistral, and similar models on a single consumer GPU. This guide covers 8-bit and 4-bit loading, the QLoRA pattern for fine-tuning, and the tradeoffs involved.
Installation
pip install bitsandbytes transformers accelerate
# Verify CUDA is detected
python -c "import bitsandbytes as bnb; print(bnb.__version__)"
8-bit Loading
8-bit quantisation with LLM.int8() uses a mixed-precision scheme: outlier features (a small fraction of values with large magnitude) are kept in float16 while the majority of weights are quantised to int8. This preserves model quality close to float16 while halving memory usage:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
load_in_8bit=True, # bitsandbytes 8-bit
device_map="auto",
torch_dtype=torch.float16,
)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B")
inputs = tokenizer("Explain quantization:", return_tensors="pt").to("cuda")
with torch.no_grad():
out = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(out[0], skip_special_tokens=True))
4-bit Loading with NF4
4-bit NormalFloat (NF4) quantisation, introduced with QLoRA, achieves better quality than standard int4 by using a quantisation grid optimised for normally-distributed weights. Combined with double quantisation (quantising the quantisation constants themselves), it reduces memory further:
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4", # NormalFloat4 — best quality
bnb_4bit_compute_dtype=torch.bfloat16, # compute in bf16 for speed
bnb_4bit_use_double_quant=True, # double quantisation saves ~0.4 bits/param
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
quantization_config=bnb_config,
device_map="auto",
)
print(f"Memory: {model.get_memory_footprint()/1e9:.1f} GB")
QLoRA: Fine-Tuning a 4-bit Model
QLoRA combines 4-bit quantisation with LoRA adapters — the base model stays frozen in 4-bit, and only the small LoRA adapter weights are trained in full precision. This is the standard recipe for fine-tuning 7B–70B models on a single GPU:
from transformers import AutoModelForCausalLM, BitsAndBytesConfig, TrainingArguments
from peft import get_peft_model, LoraConfig, prepare_model_for_kbit_training
from trl import SFTTrainer
import torch
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
quantization_config=bnb_config,
device_map="auto",
)
# Required step: prepares quantised model for training
model = prepare_model_for_kbit_training(model)
lora_config = LoraConfig(
r=16,
lora_alpha=32,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# Trainable: ~0.1% of total parameters
training_args = TrainingArguments(
output_dir="./qlora-output",
num_train_epochs=3,
per_device_train_batch_size=4,
gradient_accumulation_steps=4,
optim="paged_adamw_8bit", # bitsandbytes paged optimiser
bf16=True,
learning_rate=2e-4,
lr_scheduler_type="cosine",
warmup_ratio=0.05,
logging_steps=10,
save_steps=100,
)
trainer = SFTTrainer(
model=model,
args=training_args,
train_dataset=dataset,
tokenizer=tokenizer,
dataset_text_field="text",
max_seq_length=2048,
)
trainer.train()
Memory Requirements at Different Precisions
Rough memory footprint for a 7B parameter model at different precisions (weights only, excluding activations and KV cache):
- float32: ~28 GB — does not fit on most single GPUs
- float16 / bfloat16: ~14 GB — fits on A100 40GB or RTX 4090
- int8 (LLM.int8): ~7 GB — fits on RTX 3080 / 4070
- 4-bit NF4: ~4 GB — fits on RTX 3060 / 4060
For fine-tuning, you also need memory for gradients, optimiser state, and activations — QLoRA with 4-bit keeps the base model in 4-bit and trains only LoRA adapters, so gradient and optimiser memory is proportional to the adapter size (roughly 0.1–1% of total parameters), not the full model.
Figure 1 — GPU memory required for a 7B model by precision
Paged Optimisers
bitsandbytes also provides paged optimisers (paged_adamw_8bit, paged_adamw_32bit) that store optimiser state in CPU RAM and page it to GPU only when needed. For QLoRA fine-tuning where the adapter is small, this typically saves 1–2GB of VRAM with negligible speed impact:
from transformers import TrainingArguments
training_args = TrainingArguments(
optim="paged_adamw_8bit", # 8-bit paged Adam — best for QLoRA
# or:
# optim="paged_adamw_32bit" # 32-bit paged — slightly better stability
...
)
Quality Tradeoffs
Quantisation degrades model quality, but the degree depends on the quantisation method and model size. LLM.int8() preserves quality very close to float16 — perplexity differences are typically under 0.1%. 4-bit NF4 with double quantisation degrades quality slightly more, but the gap is small enough that QLoRA fine-tuning can often recover it. Larger models quantise better than smaller ones — a 70B model in 4-bit typically outperforms a 7B model in float16 on most benchmarks, making quantisation a strictly positive tradeoff for practitioners choosing between models of different sizes. The main quality loss in QLoRA comes not from quantisation but from LoRA’s low-rank approximation; increasing LoRA rank (r=64 instead of r=16) recovers most of the quality gap at the cost of more trainable parameters.
bitsandbytes is the most practical tool for fitting large language models into limited VRAM. The 4-bit NF4 path combined with QLoRA is now the standard approach for single-GPU fine-tuning — it makes 7B model fine-tuning accessible on an RTX 3060, and 13B fine-tuning on an RTX 4090. Install it, set the BitsAndBytesConfig, and the rest of your Transformers and PEFT code works unchanged.
Figure 1 — GPU memory for a 7B model by precision
Merging LoRA Adapters After Fine-Tuning
After QLoRA training, the adapter weights live separately from the quantised base model. For deployment you can merge them into a single full-precision model, or keep them separate and load dynamically. Merging requires dequantising the base model first:
from peft import PeftModel
from transformers import AutoModelForCausalLM
import torch
# Load base model in float16 (not 4-bit) for merging
base_model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
torch_dtype=torch.float16,
device_map="cpu", # merge on CPU to avoid VRAM issues
)
# Load and merge adapter
model = PeftModel.from_pretrained(base_model, "./qlora-output/checkpoint-best")
model = model.merge_and_unload() # merges adapter into base weights
# Save merged model
model.save_pretrained("./merged-model")
tokenizer.save_pretrained("./merged-model")
The merged model is a standard float16 checkpoint with no bitsandbytes dependency — it can be loaded with any serving framework including vLLM, TGI, and Ollama. Keep the separate adapter if you need to switch between different fine-tuned versions of the same base model at runtime.
Checking Model Memory Usage
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch
bnb_config = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B",
quantization_config=bnb_config, device_map="auto")
print(f"Memory footprint: {model.get_memory_footprint()/1e9:.2f} GB")
print(f"GPU allocated: {torch.cuda.memory_allocated()/1e9:.2f} GB")
print(f"GPU reserved: {torch.cuda.memory_reserved()/1e9:.2f} GB")
# Breakdown by layer dtype
for name, param in model.named_parameters():
print(f"{name}: {param.dtype}, {param.numel()/1e6:.1f}M params")
How LLM.int8() Works
The original bitsandbytes 8-bit quantisation (LLM.int8(), introduced by Tim Dettmers) observed that transformer weights have a bimodal distribution — most values cluster near zero, but a small fraction of “outlier” values have much larger magnitudes. Quantising outliers to int8 causes large rounding errors that accumulate into noticeable quality loss. LLM.int8() solves this by decomposing the matrix multiplication: it identifies outlier feature dimensions (typically less than 0.1% of dimensions), keeps those in float16, and quantises everything else to int8. The int8 matrix multiply runs fast on NVIDIA hardware via CUDA’s int8 tensor cores, and the float16 outlier computation is added back. The result is a near-lossless 8-bit model with halved memory at roughly the same inference speed as float16 on Ampere GPUs.
How NF4 Works
NormalFloat4 (NF4) quantises weights to 4-bit using a non-uniform quantisation grid that is information-theoretically optimal for normally distributed data. Standard int4 uses a uniform grid (equal spacing between quantisation levels) which wastes most levels on rarely-occurring extreme values. NF4 places quantisation levels where the data density is highest — near zero for normally distributed weights — so the same 4 bits represent the data more accurately. The double quantisation step quantises the per-block quantisation constants (the scale factors) themselves from float32 to float8, saving an additional ~0.4 bits per parameter. Together these make NF4 the highest-quality 4-bit format available without changing the model architecture.
GPTQ vs bitsandbytes: When to Use Which
bitsandbytes performs quantisation on-the-fly when loading the model — no calibration data required, runs in seconds. GPTQ is a post-training quantisation method that requires a small calibration dataset and takes minutes to hours to quantise a model, but produces slightly better quality at the same bit width. For fine-tuning workflows, bitsandbytes is the standard choice because it integrates directly with the training loop. For inference-only deployment where you want the best possible quality at 4-bit and are willing to do a one-time quantisation pass, GPTQ (via AutoGPTQ or llama.cpp’s GGUF format) may give marginally better results. AWQ (Activation-aware Weight Quantisation) is a newer method that outperforms both at 4-bit for inference, but as of 2026 bitsandbytes remains the default for training because it is the only method with native PEFT integration for QLoRA.
When Not to Use Quantisation
Quantisation is a memory-saving tradeoff, not a free lunch. Avoid it when: you have sufficient VRAM for float16 and inference latency is critical — quantised models are slightly slower per token due to dequantisation overhead; you need to fine-tune the full model rather than LoRA adapters (4-bit frozen weights cannot receive gradients directly); or you are working with very small models (under 3B parameters) where the quality loss from 4-bit is proportionally larger and float16 already fits in memory. For models above 7B where fitting in memory is the constraint, 4-bit NF4 is almost always the right default.
bitsandbytes’ 4-bit NF4 path with QLoRA is the standard recipe for single-GPU fine-tuning of large language models. Load the model with BitsAndBytesConfig, wrap it with prepare_model_for_kbit_training, add LoRA adapters with PEFT, and train with the paged_adamw_8bit optimiser. That four-step pattern fits 7B fine-tuning into 8GB VRAM and 13B fine-tuning into 16GB — hardware accessible to individual practitioners without cloud GPU costs. The fact that this is now standard practice rather than a research curiosity is a direct result of bitsandbytes and the QLoRA paper that combined it with LoRA — two tools that changed what is possible on a single GPU more than any other development in the past three years of practical LLM tooling.
For most practitioners the decision is simple: if you need to run or fine-tune a model larger than what fits in float16 on your GPU, use bitsandbytes. The installation is a single pip command, the integration with Transformers and PEFT is seamless, and the quality tradeoff at 4-bit NF4 is small enough to be irrelevant for the majority of fine-tuning use cases.