Markdown

SFT

SFT is the Acronym for Supervised Fine-Tuning

The process of further training a pretrained model—typically a large language model—on a curated set of labeled input-output examples so it learns to follow instructions or perform a specific task. It’s usually the first tuning stage after pretraining and lays the groundwork for later alignment steps like RLHF or DPO.

Where SFT Fits in the LLM Training Pipeline

  1. Pretraining: The model learns general language patterns from massive unlabeled text.
  2. Supervised Fine-Tuning: The model trains on curated (prompt, ideal response) pairs, teaching it to follow instructions rather than just predict the next token.
  3. Preference Alignment: Techniques like RLHF, RLAIF, or DPO further tune the model using human or AI preference signals.

Why SFT Matters

  • Task Specialization: Businesses use SFT to adapt a general-purpose foundation model to a specific voice, domain, or workflow, such as customer support scripts or legal drafting.
  • Full-Weight vs. Efficient Tuning: Classic SFT updates all of a model’s weights; PEFT methods like LoRA achieve similar results by updating far fewer parameters.

Articles Tagged SFT

View Additional Articles Tagged SFT