Abu Dhabi University
Interactive poster

Parameter-Efficient Fine-Tuning of Large Language Models:

A Review of Methods, Evidence and Open Problems
Mohamed Ahmed Obaid Alzaabi1 (1109963) · 1109963@students.adu.ac.ae
Supervised by: Dr. Anas Altarabsheh2
1 Master of Science in Artificial Intelligence, College of Engineering, Abu Dhabi University, Abu Dhabi, United Arab Emirates
2 Department of Electrical Engineering, College of Engineering, Abu Dhabi University, Abu Dhabi, United Arab Emirates
Abstract

Problem. Updating every weight of a large language model (LLM) needs more memory than most teams have and leaves one full copy of the model per task.

Idea. Parameter-efficient fine-tuning (PEFT) freezes the pretrained weights and trains only a small set of added or selected parameters.

Added value. Earlier surveys [5], [6] sort methods mainly by how they change the model. This review adds two lenses, cost (memory, training time, serving) and strength of evidence, and distills 50 sources into four families, two axes and practical guidance.

Why PEFT: what full fine-tuning costs
  • Memory: mixed-precision training with the Adam optimizer (adaptive moment estimation) stores 16 bytes per parameter for weights, gradients and optimizer states. A 7-billion-parameter (7B) model needs 112 GB before activations [2].
  • Storage and serving: every task adds a full model copy; serving many tasks loads them all.
  • A small update can suffice: the intrinsic dimension (fewest free parameters reaching 90 % of full accuracy) is about 200 on a paraphrase task and falls as models grow [10].
Taxonomy: four families, two axes
Select a family to highlight it in the diagram
AdditiveInserts small trainable modules or vectors; all pretrained weights stay frozen.
SelectiveFine-tunes only a chosen subset of the existing weights.
ReparameterizationLearns a compact, often low-rank update that merges into the weights at no inference cost.
HybridCombines or searches over the other families.
Two axesMemory (weight and optimizer storage) and serving (tasks per base model); both span all families.

Fig. 1 Where each PEFT family acts on a Transformer block. Gray blocks are frozen; colored blocks are trained or added. Abbreviations: MLP, multilayer perceptron; FFN, feed-forward network; BERT, bidirectional encoder representations from transformers; GLUE, General Language Understanding Evaluation; LST, ladder side-tuning.

Mechanisms, comparison and evidence
AMechanism: how low-rank adaptation (LoRA) works
tap a stage
One 4,096 × 4,096 table. Full fine-tuning trains the grey square; LoRA trains only the green strips.
0.39 % of the table is trained

Fig. 2 LoRA on one weight matrix (stage 1 shown). 1 Training: only A and B are trained; W0 stays frozen. 2 Inference: BA merges into W0 at no extra cost. 3 Cost: r(d + k) values replace d·k; 10,000× fewer trainable parameters and 3× less GPU memory on a 175B model [4].

Key to Fig. 2
W0Frozen pretrained weights (d × k).
A, BThe only trained parts; B starts at 0.
x → hLayer input and output.
⊕Adds the frozen path W0x to the update.
∂L/∂A, ∂L/∂BGradients (red, dashed); they reach A and B only.
r, αRank r ≪ min(d, k); α/r is a fixed scale.
ΔW = BAThe learned low-rank update.
Update rules, in one notation
Adapters [3]
h ← h + Wupσ(Wdownh)
LoRA [4]
h = W0x + (α/r)BAx
Weight-decomposed LoRA [20]
W = m ⊙ (W0+BA)/‖W0+BA‖c
Quantized LoRA [40]
h = dequant(W0NF4)x + (α/r)BAx
Gradient low-rank projection [2]
ΔWt = ηPtAdam(PtTGt)

Symbols: σ nonlinearity; Wup, Wdown adapter projections; m magnitude; ‖·‖c column norm; ⊙ element-wise product; dequant decodes 4-bit NormalFloat (NF4); Pt projection; Gt gradient; η learning rate.

BComparison: representative methods by family

Table I Representative PEFT methods by family. Trainable fractions are as reported by each paper and are not directly comparable, since base model and rank differ.

Method (year) [ref]What is trainedTrainable fractionInference overhead
CEvidence: tuning choices and practical guidance
Hyperparameters (H1–H5) move results more than methods do
H1

Rank: on code, rank 256 on all linear modules approached full fine-tuning; rank 8 on attention alone fell far short [47].

H2

Scaling: with the usual α/r, larger ranks bring no gain; the rank-stabilized scale α/√r fixes this [7].

H3

Placement: query plus value projections beat query alone at equal budget [4] and decide adapter results [29]; budget allocation matters as much as the method [19], [37].

H4

Learning rate and start: a larger rate for B (ηB = ληA, λ ≫ 1) gains 1–2 points [34]; principal-component [35] and quantization-aware [41] initializations help.

H5

Consequence: many reported gains are no larger than those from tuning a default-configured baseline.

Practical guidance
Answer the questions; the recipe below updates.
Base model fits in memory at 16-bit?
Deployed model must stay quantized?
Target far from pretraining (code, maths)?
Large model with spare input length?
What the evidence says (E1–E5)
Trainable parameters and memory differ
fp16 weightsgradientsfp32 masterAdam moments

Fig. 3 Estimated training memory of full fine-tuning, LoRA and quantized LoRA (QLoRA) [2]; bars to scale; activations excluded. The largest common GPU holds 80 GB. fp16/fp32: 16-/32-bit floating point.

E1PEFT usually matches full fine-tuning on the authors' own benchmarks [4], [13], [40]; adapters come within 0.4 GLUE points at 3.6 % of parameters [3]. Across 100+ tasks no method dominates, and combinations help [5].
E2PEFT falls short on harder targets: on code and mathematics, LoRA trails full fine-tuning but forgets less [47].
E3LoRA learns a different solution: it adds "intruder dimensions" (new directions) [48]; an exact fit needs rank ≈ half the layer width [11]; B matters more than A [49].
E4Parameters, memory and time differ: LST cuts activation memory [9]; QLoRA fits 65B in 48 GB [40]; GaLore, 7B in 24 GB [2].
E5Safety is not preserved: ten adversarial examples remove it; benign data degrades it [50].
Conclusions and open problems
  1. Trainable fraction is a weak proxy for memory and time. Activations, optimizer states and base precision set them. E4Fig. 3
  2. LoRA matches full fine-tuning on many tasks but finds a different solution. E1E2E3
  3. Tuning choices often outweigh the choice of method. H1–H5
  4. Reporting is the most pressing open problem. State the base model and precision, adapted matrices, every hyperparameter (the baseline's too), peak memory, GPU-hours, latency and seeds. Table IH5E4
Open problems
  • Theory that predicts rank and placement.
  • Principled composition and routing of adapters [46].
  • Harder benchmarks, and evaluation beyond language.
  • Checking that efficiency has not cost safety (E5, [50]).
Scope and method of the review
  • Sources (2019–2025): leading machine-learning and language venues plus arXiv, with citation chasing from two surveys [5], [6].
  • Included: methods that later work reuses or compares against; theory and systematic evidence; memory and serving systems.
  • Excluded: distillation, inference pruning and in-context learning, except as baselines.
  • Final set: 50 sources (2019–2025): 49 peer-reviewed papers and one preprint [7], retained for its widely adopted rank-stabilized scaling rule. Two vision studies [8], [9] confirm that the mechanisms generalize beyond language.
Key references
← Home
Click any section to zoom in · tap [n] or dotted terms for details