Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +918448811320
Back to Deep Learning Notes
Topic #678

Fine-Tuning an LLM Project

A complete LLM fine-tuning project — using LoRA to efficiently adapt a pretrained language model to a specific task on modest hardware, the standard practical approach to customizing an LLM without full fine-tuning's resource demands.

Problem Statement

Fine-tune a small pretrained open-source language model on a specific, narrow task (e.g. generating text in a particular style, or answering questions in a specific domain format) using LoRA, and evaluate the fine-tuned model's outputs against the base model's.

Dataset

A set of instruction-response pairs matching the target task/style — a few hundred to a few thousand well-curated examples is typically enough to see a meaningful behavioral shift with LoRA fine-tuning.

Architecture & Approach

This project uses the HuggingFace transformers and peft libraries to apply LoRA to a small pretrained model — freezing the base model's weights entirely and training only small, injected low-rank adapter matrices, directly applying LoRA's approach in practice.

Step-by-Step Build

from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, Trainer
from peft import LoraConfig, get_peft_model
import torch

# 1. Load a small pretrained base model
model_name = "some-small-base-model"   # a small open model appropriately sized for available hardware
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)

# 2. Configure and apply LoRA -- only these small adapter matrices will be trained
lora_config = LoraConfig(
    r=8,                                 # rank of the low-rank matrices -- smaller = fewer trainable params
    lora_alpha=16,
    target_modules=["q_proj", "v_proj"],   # apply LoRA to the attention projection layers
    lora_dropout=0.05,
    task_type="CAUSAL_LM"
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()   # confirms only a small fraction of total parameters are trainable

# 3. Prepare the instruction-response dataset
def format_example(example):
    return f"### Instruction:\n{example['instruction']}\n\n### Response:\n{example['response']}"

def tokenize_function(examples):
    texts = [format_example(ex) for ex in examples]
    return tokenizer(texts, truncation=True, padding="max_length", max_length=256)

tokenized_dataset = raw_dataset.map(tokenize_function, batched=True)

# 4. Train using HuggingFace's Trainer
training_args = TrainingArguments(
    output_dir="./lora-finetuned",
    num_train_epochs=3,
    per_device_train_batch_size=4,
    learning_rate=2e-4,
    logging_steps=10,
    save_strategy="epoch"
)

trainer = Trainer(model=model, args=training_args, train_dataset=tokenized_dataset)
trainer.train()

# 5. Save just the small LoRA adapter weights (not the full model)
model.save_pretrained("./lora-adapter")

# 6. Compare base vs fine-tuned model outputs on the same prompt
def generate_response(model, prompt):
    inputs = tokenizer(prompt, return_tensors="pt")
    output = model.generate(**inputs, max_new_tokens=100)
    return tokenizer.decode(output[0], skip_special_tokens=True)

test_prompt = "### Instruction:\nExplain photosynthesis simply.\n\n### Response:\n"
print("Fine-tuned output:", generate_response(model, test_prompt))

Expected Results

The fine-tuned model's outputs should visibly shift toward the target task's style and format compared to the base model's outputs on the same prompts — the specific magnitude of improvement depends heavily on dataset quality and size, but even a few hundred well-curated examples typically produce a noticeable, measurable behavioral shift with LoRA.

Key Learnings & Extensions

  • Notice how small the saved adapter file is compared to the full base model — this is LoRA's core efficiency benefit made concrete: the adapter can be shared, versioned, and swapped independently of the (much larger) frozen base model.
  • Extension: Try QLoRA (LoRA combined with a quantized base model) to fine-tune a larger base model within the same hardware constraints, directly applying QLoRA.
  • Extension: Build a small evaluation set and compare base vs fine-tuned model performance quantitatively (not just qualitatively), applying the rigor from Statistical Significance if comparing across multiple runs.

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →