
The process of further training a pretrained model—typically a large language model—on a curated set of labeled input-output examples so it learns to follow instructions or perform a specific task. It’s usually the first tuning stage after pretraining and lays the groundwork for later alignment steps like RLHF or DPO.
Where SFT Fits in the LLM Training Pipeline
- Pretraining: The model learns general language patterns from massive unlabeled text.
- Supervised Fine-Tuning: The model trains on curated (prompt, ideal response) pairs, teaching it to follow instructions rather than just predict the next token.
- Preference Alignment: Techniques like RLHF, RLAIF, or DPO further tune the model using human or AI preference signals.
Why SFT Matters
- Task Specialization: Businesses use SFT to adapt a general-purpose foundation model to a specific voice, domain, or workflow, such as customer support scripts or legal drafting.
- Full-Weight vs. Efficient Tuning: Classic SFT updates all of a model’s weights; PEFT methods like LoRA achieve similar results by updating far fewer parameters.