Post

Supervised Fine-Tuning (SFT) for LLMs: A Python Guide

A practical guide to supervised fine-tuning of LLMs in Python using TRL and LoRA β€” data format, training loop, and when SFT beats prompting or RAG.

Supervised Fine-Tuning (SFT) for LLMs: A Python Guide
πŸš€ Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started β†’

When Prompting Isn’t Enough

Prompting and RAG cover most needs. But when you want a model to reliably adopt a format, tone, or narrow skill β€” every time, without long instructions β€” supervised fine-tuning (SFT) is the tool. SFT trains the model on input-output pairs so the behavior is baked in. This guide runs it in Python with TRL and LoRA.

1
pip install trl peft transformers datasets

The Data Format

SFT data is pairs: a prompt and the ideal completion. Quality beats quantity β€” a few hundred clean examples often beat thousands of noisy ones.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
from datasets import Dataset

examples = [
    {"prompt": "Summarize: The meeting was rescheduled to Friday at 3pm.",
     "completion": "Meeting moved to Friday 3pm."},
    {"prompt": "Summarize: The server crashed twice due to a memory leak.",
     "completion": "Server crashed twice; cause was a memory leak."},
    # ... hundreds more, consistent in style
]

def to_text(ex):
    return {"text": f"### Instruction:\n{ex['prompt']}\n\n### Response:\n{ex['completion']}"}

dataset = Dataset.from_list(examples).map(to_text)

The template matters: use the same one at training and inference. If you train on ### Instruction: and prompt differently in production, quality drops.

LoRA: Fine-Tuning Without the Full Cost

Full fine-tuning updates every weight and needs serious GPU memory. LoRA trains small adapter matrices instead β€” a fraction of the parameters, often under 1% β€” with nearly the same result. This is what makes SFT affordable.

1
2
3
4
5
6
7
8
9
from peft import LoraConfig

peft_config = LoraConfig(
    r=16,                 # adapter rank; higher = more capacity, more memory
    lora_alpha=32,
    lora_dropout=0.05,
    task_type="CAUSAL_LM",
    target_modules=["q_proj", "v_proj"],
)

Pair LoRA with quantization (QLoRA) and you can fine-tune billion-parameter models on a single consumer GPU.

The Training Loop

TRL’s SFTTrainer wraps the whole loop β€” tokenization, batching, gradient steps.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl import SFTTrainer, SFTConfig

model_name = "meta-llama/Llama-3.2-1B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)

trainer = SFTTrainer(
    model=model,
    train_dataset=dataset,
    peft_config=peft_config,
    args=SFTConfig(
        output_dir="./sft-out",
        num_train_epochs=3,
        per_device_train_batch_size=4,
        learning_rate=2e-4,
    ),
)
trainer.train()

Watch the loss. If it drops to near zero on training data but the model repeats itself on new inputs, you overfit β€” reduce epochs or add more varied data.

SFT vs Prompting vs RAG

Reach for SFT when you need consistent behavior that prompting can’t enforce and the knowledge is a skill, not a fact. Use RAG when the model needs current or private facts β€” fine-tuning bakes knowledge in at a fixed date and is expensive to update. Many production systems use both: SFT for behavior, RAG for facts.

Frequently Asked Questions

How much data do I need for SFT?

Often a few hundred to a few thousand high-quality pairs. Consistency matters more than volume β€” every example should model the exact behavior you want repeated.

What is the difference between LoRA and full fine-tuning?

Full fine-tuning updates all model weights and needs large GPU memory. LoRA freezes the base model and trains tiny adapter matrices, cutting memory and storage while matching most of the quality.

Does fine-tuning teach the model new facts?

Poorly. Fine-tuning is strong for format and behavior but unreliable for injecting facts, which it may hallucinate or forget. Use retrieval for factual grounding.

Takeaways

  • SFT trains on prompt-completion pairs to bake in behavior.
  • Use a consistent prompt template at training and inference.
  • LoRA plus quantization makes fine-tuning cheap; use RAG for facts.

For a full worked example on real hardware, see fine-tuning NVIDIA Nemotron Nano.

Khushal Jethava
Khushal Jethava

Machine Learning Engineer at Codiste, specializing in Generative AI, NLP, and Computer Vision. Building production AI systems with Python.

πŸš€ Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started β†’
This post is licensed under CC BY 4.0 by the author.
πŸš€ Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started β†’
πŸš€ Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started β†’