Parameter-Efficient Fine-Tuning of Large Language Models:
A Review of Methods, Evidence and Open Problems
2 Department of Electrical Engineering, College of Engineering, Abu Dhabi University, Abu Dhabi, United Arab Emirates
How to teach a giant AI model a new skill without retraining all of it
Parameter-efficient fine-tuning (PEFT), explained from zero in twelve short steps, each with something to try. No background needed.
An AI language model is a giant pile of numbers
Chatbots like ChatGPT run on a large language model (LLM). Inside, it is just numbers, called weights or parameters.
When you type a question, the model does arithmetic with these numbers and produces an answer. What the model “knows” is stored in their values, learned by reading enormous amounts of text.
Pick a model size
Each dot stands for many millions of numbers. Bigger models know more, and they are heavier.
Specialising a model normally means a full copy for every job
A general model can be trained a little more on examples from one job, such as bank customer support. This is called fine-tuning.
The usual way, full fine-tuning, nudges every weight. So each job produces a brand-new 14 GB model, and serving many jobs means storing and loading many copies.
PEFT keeps one shared model untouched and saves only a tiny add-on per job.
Add jobs and watch the storage grow
Full fine-tuning
PEFT
Full fine-tuning: every job is a separate full model. PEFT: one base model is kept once, and each job adds only a small file on top of it. Sizes are for a 7-billion-weight model; the add-on size is a typical example.
Training needs about 8× more memory than the model itself
Using a model needs 2 bytes per weight. Training it needs much more, because the graphics card (GPU) must also keep, for every weight:
which way to change it (the gradient), a high-precision backup, and two running averages used by the Adam optimizer.
That adds up to 16 bytes per weight [2], and even the largest common GPU (80 GB) cannot hold a 7B model's training state.
Memory per weight during training
Total for a 7B model: 0 GB
Activations (temporary results kept while training) come on top of this.
A new skill needs only a tiny change
Researchers asked: what is the smallest number of weights you must train to learn a new task? On a task that checks whether two sentences mean the same thing, about 200 numbers reached 90 % of full fine-tuning's accuracy [10].
This is called the task's intrinsic dimension, and it shrinks as models grow, because a bigger model already knows more.
Weights in the model vs weights the task needed
Not to scale: at true scale the green dot would be 1 in about 1.8 million. That tiny change is the whole idea behind PEFT.
Freeze the model, train a small extra piece
Parameter-efficient fine-tuning (PEFT) locks (“freezes”) every original weight and trains only a small set of new or selected numbers, usually well under 1 %.
Gradients and Adam's extra memory are then needed only for that small set. Each job's result is a tiny file that sits on top of the shared model.
There are dozens of PEFT methods. The poster sorts them into four families, next.
Which weights change during training?
Four ways to do it, plus two ways to save more
A model is built from one block repeated many times. Each block has two main stations: attention (words look at each other) and a feed-forward network (each word is processed on its own).
Grey parts are frozen. Pick a family to see where it touches the block.
Tap a tab to see where that family touches the block.
LoRA: a small side path that later melts into the model
LoRA belongs to family 3. It zooms in on one table of numbers, W₀, and leaves it frozen.
Next to it, LoRA trains two thin tables, A and B. Multiplied together, B × A has the same shape as W₀ but is built from far fewer numbers. The thinness is set by the rank r.
Try it: the grey square is one table of numbers inside the model (4,096 × 4,096). Full fine-tuning would change all of it. LoRA changes only the two thin green strips. Move the slider to make the strips thicker or thinner.
Drawn to scale (strips shown at least 2 px thick so you can see them).
of the table is trained
Fewer trained numbers is not the same as less memory
With LoRA, the frozen model still has to sit in GPU memory at full size. That is why a 7B model still needs about 14 GB.
QLoRA squeezes the frozen model from 16-bit to 4-bit numbers, like saving a photo at lower quality, then applies LoRA on top. It fine-tuned a 65B model on a single 48 GB GPU [40].
Training memory by method
Bars drawn to the same scale. Activations and the small add-on are left out [2]. The largest common GPU holds 80 GB.
Mostly yes, with four important “buts”
The review read 50 studies and checked what the evidence really supports, not just what each paper claims about itself. The poster lists five findings, E1 to E5 (E for evidence).
Tap a card.
How you set it up matters more than which method you pick
Before training you choose hyperparameters: the rank, where to place LoRA, how fast it learns. These choices often move results more than switching methods (H1–H5, H for hyperparameters).
Size: on code, rank 256 everywhere got close to full fine-tuning; the old default, rank 8 on attention only, fell far short [47].
Scaling: a small change to how LoRA's correction is scaled lets bigger ranks actually help [7].
Placement: where you spend the budget matters as much as the method [4], [19].
Speed and start: letting B learn faster gains 1–2 points [34].
So: many “new method wins” are no bigger than what tuning the old method's settings would give.
Build your starting recipe
Four things to remember
- “Only 0.1 % trained” doesn't mean cheap. Memory and time depend on the frozen model, temporary results and the optimizer.
- LoRA often matches full fine-tuning, but it isn't the same model. It learns less on hard tasks, forgets less, and finds a different solution.
- Settings beat method choice. Tune the rank, placement and learning rate before hunting for a new method.
- Papers must report more. Model and precision, every setting (including the comparison's), peak memory, GPU-hours, speed and repeated runs.
Still open: questions for future research
- A theory that predicts the right rank and placement before training.
- Combining or switching between many add-ons correctly [46].
- Harder tests, and tests beyond language.
- Checking that saving memory hasn't removed safety [50].
Seven quick questions
No pressure, this isn't a test. It just shows which parts are clear and which are worth another look.