Glossary · 5 minute read
What Is Parameter-Efficient Fine-Tuning? PEFT Explained
Parameter-efficient fine-tuning is the family of methods that adapt a model by training a small number of new or selected parameters while leaving most of the original weights frozen. LoRA, adapter layers, prefix tuning, and prompt tuning are the main approaches, and they match full fine-tuning closely on most adaptation tasks.
Parameter-efficient fine-tuning made model customisation cheap enough to be an ordinary engineering option rather than a research project. The shift is genuinely significant, and it has also produced a wave of fine-tuning that should have been retrieval. This explainer covers the methods, their differences, and what actually determines success. It complements what is lora fine-tuning and fine-tuning vs rag, and reflects FISTA Solutions' approach in AI enablement delivery.
Why does training so few parameters work?
Because the base model already holds general linguistic and reasoning competence, and adaptation to a specific task is a comparatively small adjustment. Empirically, the update required to specialise a model appears to occupy a low-dimensional subspace, which is precisely what low-rank methods exploit.
The practical consequence is that a fraction of a percent of the parameters, trained well, achieves most of what updating all of them would.
| Method | Trains | Relative cost | Expressiveness |
|---|---|---|---|
| Full fine-tuning | All weights | Highest | Highest |
| LoRA | Low-rank matrices | Low | High |
| Adapter layers | Inserted modules | Low | High |
| Prefix tuning | Per-layer prefix vectors | Very low | Moderate |
| Prompt tuning | Soft prompt embeddings | Lowest | Lower |
| Partial freezing | Selected layers | Moderate | Varies |
What are the main approaches?
LoRA trains low-rank matrices alongside frozen weights and merges their contribution at inference. Adapter layers insert small trainable modules between existing layers. Prefix tuning prepends learned vectors to the input of each layer. Prompt tuning learns a soft prompt in embedding space, leaving the model entirely untouched.
They differ in how much capacity they add and where. Broadly, more capacity means more expressiveness and more cost, with LoRA occupying a favourable middle position.
Which should most teams choose?
LoRA, for reasons that are practical rather than theoretical. Tooling support is mature, serving frameworks handle adapter loading and multi-adapter serving natively, and adapters are portable artefacts that can be versioned and distributed.
Those ecosystem properties matter more for production work than marginal quality differences between methods, which are usually small and workload-dependent.
When are the lighter methods appropriate?
Prompt and prefix tuning are attractive when many task variants are needed and each requires only a modest behavioural shift, because the artefacts are tiny and switching is trivial. They are less effective where the adaptation is substantial.
They are also useful as a first experiment: if a soft prompt achieves the goal, heavier methods are unnecessary.
Does PEFT solve catastrophic forgetting?
It reduces it. Because the original weights are frozen, the base capabilities are structurally preserved in a way that full fine-tuning does not guarantee. But the adapter still influences behaviour everywhere, so performance on tasks outside the training distribution can degrade.
Evaluation must therefore cover capabilities the tuning was not aimed at, not only the target task. A model that formats invoices beautifully and has become worse at general reasoning is a regression nobody measured.
Does it change what fine-tuning is for?
No. PEFT changes the cost of fine-tuning, not its capability. It teaches form, style, format, and task behaviour — exactly what full fine-tuning teaches — and it is equally poor at installing factual knowledge.
The affordability has unfortunately made it easier to attempt the wrong thing cheaply. Knowledge belongs in retrieval; behaviour belongs in tuning.
What determines success?
Data quality, overwhelmingly. Consistent, correct, well-formatted examples in modest numbers outperform large inconsistent sets. Curation is the substantive work and the place effort should go.
The second determinant is honest evaluation: comparing against a well-prompted base model, on a held-out set, including tasks outside the tuning target. Teams that skip that comparison frequently ship an adapter that prompting would have matched.
What is the ongoing cost?
Lifecycle. Each adapter is tied to a base model version, so base upgrades require retraining and re-evaluation. Training data must be maintained and versioned. Evaluation must be rerun. A dozen adapters is a dozen small ongoing obligations.
That cost is manageable and should be planned rather than discovered when a provider deprecates a base model version.
How does this fit with routing?
Well. A set of task-specific adapters over a shared base model is a natural complement to model routing: route by task, apply the appropriate adapter, serve from one resident base. That architecture is considerably cheaper than separate models per task and is well supported by current serving stacks.
What about merging adapters back?
LoRA adapters can be merged into the base weights, producing a single model with no inference-time overhead. That is useful where one adaptation is permanent and serving simplicity matters more than flexibility, and unhelpful where multiple adapters share a base.
Merging also forecloses the option of removing the adaptation later without retraining, so it is worth treating as a deployment decision rather than a default step.
How FISTA Solutions helps
FISTA Solutions chooses adaptation methods on ecosystem maturity rather than theoretical marginal quality, prioritises data curation over dataset size, evaluates adapters against well-prompted base models and on out-of-distribution capabilities, and plans adapter lifecycle against base model versions, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To adapt a model without over-investing, message FISTA on WhatsApp, or read what is lora fine-tuning.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why does training few parameters work at all?
Because adaptation to a specific task appears to require far less capacity than the model already holds. The base model has learned general language competence; adaptation is a comparatively small adjustment, and a low-dimensional update captures most of it.
02What are the main PEFT methods?
LoRA and its variants, which train low-rank matrices alongside frozen weights; adapter layers inserted between existing layers; prefix tuning, which prepends trainable vectors to each layer's input; and prompt tuning, which learns a soft prompt in the embedding space.
03Which should most teams use?
LoRA, in practice. Not because it is theoretically superior in every case but because tooling support, serving infrastructure, and the portability of adapters are all mature, which matters more than marginal quality differences for production work.
04Does PEFT prevent catastrophic forgetting?
It reduces it, because most weights are untouched, but it does not eliminate it. Adaptation can still degrade performance on tasks outside the training distribution, so evaluation must cover capabilities the tuning was not aimed at.
05Does it change what fine-tuning is good for?
No. It changes the cost of fine-tuning, not its capability. PEFT teaches form, style, and task behaviour exactly as full fine-tuning does, and it is equally unsuited to installing factual knowledge that belongs in retrieval where it can be updated and cited.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.