Problem. Updating every weight of a large language model (LLM) needs more memory than most teams have and leaves one full copy of the model per task.
Idea. Parameter-efficient fine-tuning (PEFT) freezes the pretrained weights and trains only a small set of added or selected parameters.
Added value. Earlier surveys [5], [6] sort methods mainly by how they change the model. This review adds two lenses, cost (memory, training time, serving) and strength of evidence, and distills 50 sources into four families, two axes and practical guidance.
| Additive | Inserts small trainable modules or vectors; all pretrained weights stay frozen. | |
| Selective | Fine-tunes only a chosen subset of the existing weights. | |
| Reparameterization | Learns a compact, often low-rank update that merges into the weights at no inference cost. | |
| Hybrid | Combines or searches over the other families. | |
| Two axes | Memory (weight and optimizer storage) and serving (tasks per base model); both span all families. |
Fig. 1 Where each PEFT family acts on a Transformer block. Gray blocks are frozen; colored blocks are trained or added. Abbreviations: MLP, multilayer perceptron; FFN, feed-forward network; BERT, bidirectional encoder representations from transformers; GLUE, General Language Understanding Evaluation; LST, ladder side-tuning.
Fig. 2 LoRA on one weight matrix (stage 1 shown). 1 Training: only A and B are trained; W0 stays frozen. 2 Inference: BA merges into W0 at no extra cost. 3 Cost: r(d + k) values replace d·k; 10,000× fewer trainable parameters and 3× less GPU memory on a 175B model [4].
| Key to Fig. 2 | |
|---|---|
| W0 | Frozen pretrained weights (d × k). |
| A, B | The only trained parts; B starts at 0. |
| x → h | Layer input and output. |
| ⊕ | Adds the frozen path W0x to the update. |
| ∂L/∂A, ∂L/∂B | Gradients (red, dashed); they reach A and B only. |
| r, α | Rank r ≪ min(d, k); α/r is a fixed scale. |
| ΔW = BA | The learned low-rank update. |
| Update rules, in one notation | |
|---|---|
| Adapters [3] | h ← h + Wupσ(Wdownh) |
| LoRA [4] | h = W0x + (α/r)BAx |
| Weight-decomposed LoRA [20] | W = m ⊙ (W0+BA)/‖W0+BA‖c |
| Quantized LoRA [40] | h = dequant(W0NF4)x + (α/r)BAx |
| Gradient low-rank projection [2] | ΔWt = ηPtAdam(PtTGt) |
Symbols: σ nonlinearity; Wup, Wdown adapter projections; m magnitude; ‖·‖c column norm; ⊙ element-wise product; dequant decodes 4-bit NormalFloat (NF4); Pt projection; Gt gradient; η learning rate.
Table I Representative PEFT methods by family. Trainable fractions are as reported by each paper and are not directly comparable, since base model and rank differ.
| Method (year) [ref] | What is trained | Trainable fraction | Inference overhead |
|---|
Rank: on code, rank 256 on all linear modules approached full fine-tuning; rank 8 on attention alone fell far short [47].
Scaling: with the usual α/r, larger ranks bring no gain; the rank-stabilized scale α/√r fixes this [7].
Placement: query plus value projections beat query alone at equal budget [4] and decide adapter results [29]; budget allocation matters as much as the method [19], [37].
Learning rate and start: a larger rate for B (ηB = ληA, λ ≫ 1) gains 1–2 points [34]; principal-component [35] and quantization-aware [41] initializations help.
Consequence: many reported gains are no larger than those from tuning a default-configured baseline.
Fig. 3 Estimated training memory of full fine-tuning, LoRA and quantized LoRA (QLoRA) [2]; bars to scale; activations excluded. The largest common GPU holds 80 GB. fp16/fp32: 16-/32-bit floating point.