Markdown

LoRA

LoRA is the Acronym for Low-Rank Adaptation

A parameter-efficient way to adapt a pretrained neural network without rewriting the original matrix of weights. Instead of updating every parameter, the method freezes the base model and learns a pair of much smaller matrices whose product is added at each chosen layer. Edward J. Hu and colleagues at Microsoft introduced the technique in 2021 for large language models, and it is now the default path for most in-house fine-tunes.

How the Adapter Is Built

A weight matrix W in a transformer stays frozen. The update is a low-rank product BA, where A and B have a rank far smaller than the original dimensions. Only A and B are trained. That is the heart of Parameter-Efficient Fine-Tuning (PEFT): you change behavior with a thin adapter instead of a second copy of the full network.

  • Frozen base: The published checkpoint does not move during the run.
  • Trainable pair: Two skinny matrices per adapted layer, typically injected into attention projections.
  • Merge at serve time: After training, W + BA can be folded back into a single matrix so inference cost matches the original model.

Compared With Full Fine-Tuning

Full Supervised Fine-Tuning (SFT) updates every parameter. That needs far more Graphics Processing Unit (GPU) memory and produces a complete new checkpoint. Hu et al. reported on the order of 10,000 times fewer trainable parameters than a full GPT-3 175B fine-tune, about three times less GPU memory, and no extra latency once the adapter is merged. Quality on their tasks was on par or better.

QLoRA

QLoRA keeps the frozen base in 4-bit precision and still trains LoRA adapters in higher precision. Tim Dettmers and colleagues showed a 65B-parameter model could be fine-tuned on a single 48 GB GPU. That is how many teams take an Apache 2.0 or MIT Large Language Model (LLM) and specialize it on one workstation instead of a training cluster.

Where Teams Use It

Use LoRA when the job is stable behavior: brand voice, extraction schemas, tool-call format, or a fixed taxonomy. Do not use it as a substitute for documents that change every week. Those facts belong in retrieval. After training, serve the merged weights or load the adapter beside the base model on your own runtime.

Additional Acronyms for LoRA

Articles Tagged LoRA

View Additional Articles Tagged LoRA