A complete LLM fine-tuning project — using LoRA to efficiently adapt a pretrained language model to a specific task on modest hardware, the standard practical approach to customizing an LLM without full fine-tuning's resource demands.
Problem Statement
Fine-tune a small pretrained open-source language model on a specific, narrow task (e.g. generating text in a particular style, or answering questions in a specific domain format) using LoRA, and evaluate the fine-tuned model's outputs against the base model's.
Dataset
A set of instruction-response pairs matching the target task/style — a few hundred to a few thousand well-curated examples is typically enough to see a meaningful behavioral shift with LoRA fine-tuning.
Architecture & Approach
This project uses the HuggingFace transformers and peft libraries to apply LoRA to a small pretrained model — freezing the base model's weights entirely and training only small, injected low-rank adapter matrices, directly applying LoRA's approach in practice.
Step-by-Step Build
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, Trainer
from peft import LoraConfig, get_peft_model
import torch
# 1. Load a small pretrained base model
model_name = "some-small-base-model" # a small open model appropriately sized for available hardware
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
# 2. Configure and apply LoRA -- only these small adapter matrices will be trained
lora_config = LoraConfig(
r=8, # rank of the low-rank matrices -- smaller = fewer trainable params
lora_alpha=16,
target_modules=["q_proj", "v_proj"], # apply LoRA to the attention projection layers
lora_dropout=0.05,
task_type="CAUSAL_LM"
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters() # confirms only a small fraction of total parameters are trainable
# 3. Prepare the instruction-response dataset
def format_example(example):
return f"### Instruction:\n{example['instruction']}\n\n### Response:\n{example['response']}"
def tokenize_function(examples):
texts = [format_example(ex) for ex in examples]
return tokenizer(texts, truncation=True, padding="max_length", max_length=256)
tokenized_dataset = raw_dataset.map(tokenize_function, batched=True)
# 4. Train using HuggingFace's Trainer
training_args = TrainingArguments(
output_dir="./lora-finetuned",
num_train_epochs=3,
per_device_train_batch_size=4,
learning_rate=2e-4,
logging_steps=10,
save_strategy="epoch"
)
trainer = Trainer(model=model, args=training_args, train_dataset=tokenized_dataset)
trainer.train()
# 5. Save just the small LoRA adapter weights (not the full model)
model.save_pretrained("./lora-adapter")
# 6. Compare base vs fine-tuned model outputs on the same prompt
def generate_response(model, prompt):
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=100)
return tokenizer.decode(output[0], skip_special_tokens=True)
test_prompt = "### Instruction:\nExplain photosynthesis simply.\n\n### Response:\n"
print("Fine-tuned output:", generate_response(model, test_prompt))
Expected Results
The fine-tuned model's outputs should visibly shift toward the target task's style and format compared to the base model's outputs on the same prompts — the specific magnitude of improvement depends heavily on dataset quality and size, but even a few hundred well-curated examples typically produce a noticeable, measurable behavioral shift with LoRA.
Key Learnings & Extensions
- Notice how small the saved adapter file is compared to the full base model — this is LoRA's core efficiency benefit made concrete: the adapter can be shared, versioned, and swapped independently of the (much larger) frozen base model.
- Extension: Try QLoRA (LoRA combined with a quantized base model) to fine-tune a larger base model within the same hardware constraints, directly applying QLoRA.
- Extension: Build a small evaluation set and compare base vs fine-tuned model performance quantitatively (not just qualitatively), applying the rigor from Statistical Significance if comparing across multiple runs.