
Parameter-Efficient Fine-Tuning (PEFT) is a family of techniques for adapting a large pretrained model to a new task by training only a small subset of parameters—or a small set of added parameters—instead of updating the entire model. It delivers most of the benefit of full fine-tuning at a fraction of the compute, memory, and storage cost, which is why it’s become the default way most teams customize large language models.
Common PEFT Techniques
- LoRA (Low-Rank Adaptation): Freezes the original model weights and trains small low-rank matrices injected into each layer.
- QLoRA: Combines LoRA with a quantized base model to cut memory requirements further.
- Prompt/Prefix Tuning: Learns a small set of continuous “virtual tokens” prepended to the input instead of touching the model’s weights at all.
- Adapters: Inserts small trainable layers between the frozen layers of the base model.
Why PEFT Matters
- Lower Cost: Training a fraction of a model’s parameters requires far less GPU memory and compute than full fine-tuning.
- Faster Iteration: Smaller trainable footprints mean quicker experiments and faster time-to-deployment for domain-specific models.
- Multiple Task Adapters: Because the base model stays frozen, teams can swap lightweight PEFT adapters in and out to serve many use cases from a single underlying model.