Fine-tuning adapts a pre-trained large language model to your specific use case. Maybe you want a model that follows your company's coding conventions, answers domain-specific questions with precision, or produces output in a particular format. The base model already understands language and code -- fine-tuning teaches it your preferences and domain knowledge.
The problem is that full fine-tuning of a 7B or 8B parameter model requires enormous amounts of GPU memory. You'd need multiple A100s just to hold the model, optimizer states, and gradients in memory. QLoRA (Quantized Low-Rank Adaptation) solves this by combining two techniques: 4-bit quantization compresses the base model to fit in a fraction of the memory, and LoRA (Low-Rank Adaptation) injects small trainable matrices into the model's layers instead of updating all parameters. Together, they let you fine-tune an 8B parameter model on a single consumer GPU.
Instead of training all billions of parameters, LoRA decomposes weight updates into two small matrices. For a weight matrix of size d x d, LoRA replaces the full update with two matrices of size d x r and r x d, where r (the rank) is typically 8-64. This means you're training roughly 0.1-1% of the total parameters while achieving results close to full fine-tuning.
This guide walks through the entire process end-to-end using real, working code.
Prerequisites and Installation
You'll need an NVIDIA GPU with CUDA support. QLoRA dramatically reduces memory requirements: an 8B parameter model needs approximately 6-8GB of VRAM with 4-bit quantization, making it feasible on an RTX 3060 (12GB), RTX 4060 (8GB), or better. For context, loading the same model in full fp16 precision would require around 16GB.
Install the required libraries:
pip install torch transformers peft trl datasets bitsandbytes accelerate
Here's what each package does:
- torch: PyTorch, the underlying deep learning framework
- transformers: Hugging Face's library for loading and working with pre-trained models
- peft: Parameter-Efficient Fine-Tuning -- provides the LoRA implementation
- trl: Transformer Reinforcement Learning -- provides
SFTTrainerfor supervised fine-tuning - datasets: Hugging Face's dataset loading and processing library
- bitsandbytes: NVIDIA GPU quantization library -- enables 4-bit model loading
- accelerate: Handles device placement and distributed training
Make sure your NVIDIA drivers and CUDA toolkit are installed. You can verify with nvidia-smi.
Dataset Preparation
Fine-tuning datasets for instruction-following models use a chat/conversation format. The standard structure is a list of messages with roles (system, user, assistant), stored as JSONL (one JSON object per line).
Create a file called training_data.jsonl where each line looks like this:
{"messages": [{"role": "system", "content": "You are a helpful coding assistant."}, {"role": "user", "content": "Write a function to reverse a linked list in Python."}, {"role": "assistant", "content": "Here's a function to reverse a singly linked list:\n\n```python\ndef reverse_linked_list(head):\n prev = None\n current = head\n while current:\n next_node = current.next\n current.next = prev\n prev = current\n current = next_node\n return prev\n```\n\nThis iterates through the list, reversing each node's pointer to point to the previous node instead of the next one. Time complexity is O(n), space complexity is O(1)."}]}
Each line is a complete conversation. Multi-turn conversations work too -- just include multiple user/assistant pairs in the messages array.
Load the dataset with Hugging Face's datasets library:
from datasets import load_dataset
dataset = load_dataset("json", data_files="training_data.jsonl", split="train")
You can also use existing datasets from the Hugging Face Hub. For example, load_dataset("mlabonne/guanaco-llama2-1k") provides a pre-formatted instruction dataset. The key is that the data matches what you want the model to learn -- quality matters far more than quantity.
Loading the Base Model with 4-bit Quantization
Load the base model in 4-bit precision using BitsAndBytesConfig. This compresses the model weights from 16 bits to 4 bits per parameter, cutting memory usage by roughly 4x:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
model_name = "meta-llama/Llama-3.1-8B-Instruct"
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
)
model = AutoModelForCausalLM.from_pretrained(
model_name,
quantization_config=bnb_config,
device_map="auto",
torch_dtype=torch.bfloat16,
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "right"
Breaking down the BitsAndBytesConfig parameters:
load_in_4bit=True: Loads model weights in 4-bit precision instead of the default 16-bit or 32-bit. This is what makes fitting a large model on a consumer GPU possible.bnb_4bit_quant_type="nf4": Uses NormalFloat4 quantization rather than uniform int4. NF4 is information-theoretically optimal for normally distributed weights, which neural network weights tend to be. This produces better quality than naive 4-bit quantization.bnb_4bit_compute_dtype=torch.bfloat16: While the weights are stored in 4-bit, actual matrix multiplications are performed in bfloat16. This preserves numerical stability and takes advantage of hardware-accelerated bf16 on modern GPUs (Ampere and newer).bnb_4bit_use_double_quant=True: Quantizes the quantization constants themselves, saving an additional ~0.4 bits per parameter. A small but meaningful memory reduction at no quality cost.
Setting device_map="auto" lets accelerate handle placing model layers across available GPUs (or on a single GPU). The tokenizer configuration ensures padding works correctly during batched training.
Note: Llama models require accepting Meta's license on Hugging Face and authenticating with huggingface-cli login. You can substitute any other causal LM model (Mistral, Qwen, Gemma, etc.) by changing model_name.
Configuring LoRA
With the quantized model loaded, configure the LoRA adapters using the peft library:
from peft import LoraConfig, prepare_model_for_kbit_training, get_peft_model
model = prepare_model_for_kbit_training(model)
lora_config = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
bias="none",
task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
prepare_model_for_kbit_training prepares the quantized model for training by handling gradient computation for quantized weights and enabling input gradient computation for the layers that need it.
The key LoRA parameters:
r=16: The rank of the low-rank decomposition. This controls the capacity of the adapter -- higher rank means more trainable parameters and more expressiveness, but also more memory. Values of 8-64 are typical. Start with 16 and increase if the model underfits.lora_alpha=32: A scaling factor applied to the LoRA updates. The effective learning rate for LoRA layers is proportional tolora_alpha / r. A common convention is to setlora_alphato 2x the rank.lora_dropout=0.05: Dropout applied to the LoRA layers for regularization. Helps prevent overfitting, especially with small datasets.target_modules: Which linear layers receive LoRA adapters. The list above covers all the attention projections (q_proj,k_proj,v_proj,o_proj) and the MLP layers (gate_proj,up_proj,down_proj) in the Llama architecture. Targeting all linear layers generally gives the best results.bias="none": Don't train bias terms -- keeping this as "none" is standard for QLoRA.task_type="CAUSAL_LM": Tells PEFT this is a causal language model fine-tuning task.
The print_trainable_parameters() call shows exactly how few parameters you're actually training. For an 8B model with rank 16 targeting all linear layers, expect output like:
trainable params: 41,943,040 || all params: 8,071,442,432 || trainable%: 0.5196
That's about 42 million trainable parameters out of 8 billion -- roughly 0.5% of the model.
Training with SFTTrainer
Set up the training loop using SFTTrainer from the trl library. SFTTrainer handles tokenization and formatting of chat-style datasets automatically:
from trl import SFTTrainer
from transformers import TrainingArguments
training_args = TrainingArguments(
output_dir="./output",
num_train_epochs=3,
per_device_train_batch_size=4,
gradient_accumulation_steps=4,
learning_rate=2e-4,
weight_decay=0.01,
warmup_ratio=0.03,
lr_scheduler_type="cosine",
logging_steps=10,
save_strategy="epoch",
bf16=True,
gradient_checkpointing=True,
max_grad_norm=0.3,
optim="paged_adamw_8bit",
)
trainer = SFTTrainer(
model=model,
args=training_args,
train_dataset=dataset,
tokenizer=tokenizer,
max_seq_length=2048,
)
trainer.train()
trainer.save_model("./fine-tuned-adapter")
The training arguments worth understanding:
gradient_checkpointing=True: Instead of storing all intermediate activations in memory for backpropagation, this recomputes them during the backward pass. It uses more compute but dramatically reduces memory -- essential for fitting training on consumer GPUs.optim="paged_adamw_8bit": Uses an 8-bit version of the AdamW optimizer with paged memory management. The optimizer states for AdamW normally require 8 bytes per parameter (momentum + variance). The 8-bit version reduces this to 2 bytes per parameter, and paging handles memory spikes by offloading to CPU RAM if GPU memory runs out.gradient_accumulation_steps=4: Accumulates gradients over 4 forward passes before performing a weight update. Withper_device_train_batch_size=4, this gives an effective batch size of 16 without needing the VRAM to hold 16 samples at once.bf16=True: Enables bfloat16 mixed-precision training. Computations use bf16 where possible for speed, while keeping critical operations in fp32 for stability.learning_rate=2e-4: A good starting point for QLoRA fine-tuning. This is higher than typical full fine-tuning rates because LoRA adapters are initialized near zero and need larger updates to learn.max_grad_norm=0.3: Clips gradients to prevent training instability, which can happen with quantized models.
Training time depends on your dataset size, sequence length, and GPU. A dataset of 1,000 examples with max_seq_length=2048 on an RTX 4090 typically completes in 15-30 minutes for 3 epochs.
Merging and Exporting
After training, the adapter weights are saved separately from the base model (only about 80-160MB for a typical LoRA config). For deployment, you typically want to merge the adapter back into the base model to get a single, self-contained model:
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base_model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B-Instruct",
torch_dtype=torch.bfloat16,
device_map="auto",
)
model = PeftModel.from_pretrained(base_model, "./fine-tuned-adapter")
merged_model = model.merge_and_unload()
merged_model.save_pretrained("./merged-model")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
tokenizer.save_pretrained("./merged-model")
Note that for merging, you load the base model in full precision (bf16, not 4-bit quantized). The merging process adds the LoRA weight deltas directly into the base model weights, then discards the LoRA machinery. The result is a standard Hugging Face model that can be loaded without the peft library.
The merged model can be served with any inference framework: vLLM, TGI, llama.cpp (after conversion to GGUF), or Ollama.
Alternatively, you can skip merging and serve the adapter separately. vLLM supports this directly with the --lora-modules flag, which lets you hot-swap LoRA adapters at runtime. This is useful when you have multiple fine-tuned variants of the same base model -- you load the base model once and switch adapters per request.
Tips and Common Pitfalls
Start small and validate. Begin with 100-500 high-quality examples. Fine-tune, evaluate the outputs manually, and iterate on your dataset before scaling to thousands of examples. Data quality dominates data quantity for most use cases.
Watch the training loss curve. A healthy training run shows loss decreasing steadily then leveling off. If loss plateaus immediately and barely drops, your learning rate is probably too low -- try increasing it by 2-5x. If loss spikes or oscillates wildly, the learning rate is too high -- reduce it or increase warmup_ratio.
Evaluate on held-out data. Training loss alone tells you if the model is memorizing your data, not if it generalizes. Always set aside 10-20% of your dataset for evaluation and check the model's outputs on those examples after training.
Handle OOM errors systematically. If you run out of GPU memory, try these in order: reduce per_device_train_batch_size to 1 or 2, reduce max_seq_length to 1024, increase gradient_accumulation_steps to maintain effective batch size, and make sure gradient_checkpointing is enabled. If none of that works, reduce the LoRA rank r.
Don't overtrain. With small datasets (under 1,000 examples), 1-3 epochs is usually sufficient. More epochs risk overfitting, where the model memorizes training examples rather than learning general patterns. If training loss reaches near zero, you've almost certainly overfit.
Wrapping Up
QLoRA makes fine-tuning accessible on hardware that most developers already own or can rent cheaply. The workflow is straightforward: prepare your dataset, load a quantized base model, attach LoRA adapters, train with SFTTrainer, and merge for deployment.
Most developers doing fine-tuning work on remote GPU servers since CUDA is required and many development machines are Apple Silicon. If you're SSHing into remote GPU instances regularly, tools like Agents UI can help you manage those sessions -- keeping multiple terminals organized across local and remote machines while running training jobs.