← all articles

LoRA adapters or a full fine tune: how to actually decide

The real question isn’t accuracy, it’s what you’re allowed to touch

Every time someone asks whether to use LoRA fine tuning or a full fine tune, they’re really asking a resource question dressed up as a quality question. Both approaches can get a model to learn a new task, a new tone, a new format, or a new domain vocabulary. What separates them is how much hardware, time, and risk you’re willing to spend to get there. If you’ve ever queued a full fine tune job and watched it eat every GPU you own for a weekend, you already know which side of that tradeoff you’d rather be on by default.

LoRA, short for low rank adaptation, exists because most teams don’t have the hardware or the need to update every weight in a model. A full fine tune touches all of them. Once you understand the mechanical difference between the two, the decision stops being a vibe and starts being arithmetic.

How LoRA actually works under the hood

A transformer’s weight matrices are big. A single attention projection in a mid sized model can be a matrix with millions of parameters. Full fine tuning updates that entire matrix directly, which means you need to store gradients and optimizer state for every one of those numbers.

LoRA takes a different approach. It freezes the original weight matrix completely and instead learns two small matrices, usually called A and B, whose product approximates the change you’d want to make to the original weights. If the original matrix has dimensions d by k, LoRA replaces the update with a matrix A of size d by r and a matrix B of size r by k, where r is the rank you choose, commonly something like 8, 16, 32, or 64. Multiply A by B and you get a matrix the same shape as the original, but because r is small compared to d and k, the number of trainable parameters in A and B together is a small fraction of what the original matrix holds.

During training, only A and B get gradients. The base weights sit there untouched, in inference mode, contributing to the forward pass but never updated. At inference time you either keep A and B as a separate add-on that gets summed into the base weights on the fly, or you merge them mathematically into a new copy of the base weights so there’s zero extra latency at serving time.

This is why LoRA fine tuning is popular for anyone paying their own hardware bill. You’re not touching billions of parameters. You’re touching a rounding error’s worth of them, and the base model’s original knowledge stays intact because you never disturbed it.

What a full fine tune buys you that LoRA can’t

None of this means LoRA is strictly better. A full fine tune updates every weight, which means it has more capacity to shift the model’s behavior in ways a low rank update structurally can’t reach. If the change you need is large, spread across the whole network, or requires the model to essentially relearn how it represents a domain, a rank 16 adapter bolted onto a handful of projection matrices may not have enough degrees of freedom to get there.

Full fine tuning also avoids a subtle limitation of LoRA: you have to choose which weight matrices get adapters. Most LoRA setups target the attention projections, sometimes the feed forward layers too, but rarely every single matrix in every layer. That’s a design choice, not a law of physics, and it means a LoRA fine tune is only ever as good as the coverage you gave it. A full fine tune doesn’t have this problem because it updates everything by definition.

If your training set is large, high quality, and genuinely different from what the base model saw, full fine tuning has the raw capacity to make the model behave like it was trained on that data from the start. LoRA is closer to steering the model within the neighborhood of what it already knows.

The memory math that decides most of this argument

This is where the decision usually gets made in practice, before anyone even looks at output quality. Full fine tuning with a standard optimizer like Adam needs to store, per trainable parameter: the parameter itself, its gradient, and two additional optimizer state values (momentum and variance). In mixed precision training that adds up to roughly four times the raw parameter memory just for optimizer state and gradients, on top of the model weights themselves and activation memory during the backward pass.

LoRA sidesteps almost all of that. The base weights don’t need gradients or optimizer state at all, since they’re frozen. Only the small A and B matrices need that overhead, and they’re a tiny slice of the total parameter count. Combine LoRA with a quantized base model, the approach usually called QLoRA, and you can load the frozen weights in 4 bit precision since they’re never updated and precision loss there doesn’t accumulate through training. That combination is why a technique that would otherwise need multiple data center GPUs can instead run on a single consumer card.

This is the actual reason LoRA fine tuning took off. It’s not that engineers decided low rank updates were philosophically superior. It’s that most people don’t have a rack of A100s, and LoRA is often the difference between a fine tune that’s possible on hardware you own and one that isn’t.

Where LoRA falls apart

LoRA has real failure modes and pretending otherwise doesn’t help anyone. Pick a rank too low, or target too few modules, and the adapter simply doesn’t have enough capacity to learn the task. You’ll see training loss plateau higher than you’d like, and no amount of extra epochs fixes a capacity problem, only more rank or more coverage does.

LoRA also inherits everything about the base model’s tokenizer, architecture, and pretraining biases, because it can only nudge behavior, not restructure it. If the base model has never seen your domain’s vocabulary or formatting conventions in any meaningful way, a small adapter is unlikely to teach it from scratch as reliably as continuing to train the full network on a large enough dataset would.

There’s also a practical serving wrinkle. If you keep adapters unmerged so you can swap them per request, you’re adding a small amount of extra compute at inference time for every forward pass, since the base weights and adapter output both have to be combined. Merge the adapter into the base weights and that overhead disappears, but then you’ve traded flexibility for a static model, same as you would with a full fine tune.

Serving adapters in production

One underrated reason teams reach for LoRA has nothing to do with training cost. It’s deployment flexibility. Because the base model stays frozen and unchanged, you can train a dozen different LoRA adapters for a dozen different customers or tasks, and swap between them at inference time without hosting a dozen full copies of the model. You keep one base model resident and load small adapter weights per request. A full fine tune gives you none of that. Each fine tuned checkpoint is its own full sized model, with its own storage cost and its own deployment slot.

If your product needs per customer customization, or you’re iterating on many small variants of the same base model, that serving story alone can settle the decision before training cost even enters the conversation.

A rule of thumb that has held up

Start with LoRA. It’s cheaper to try, faster to iterate on, and reversible in the sense that you can throw away a bad adapter and keep your base model clean. Reach for a full fine tune only when you’ve hit LoRA’s actual ceiling: you’ve tried reasonable ranks, you’ve covered the relevant weight matrices, your dataset is large and high quality, and the model still isn’t absorbing the behavior you need. At that point the constraint isn’t your hardware budget anymore, it’s the structural limit of a low rank update, and that’s a real limit worth respecting rather than arguing with.

The honest tradeoff

Neither approach is a shortcut around good data. LoRA fine tuning is a way to make good data go further on hardware you can actually afford, and it keeps your base model’s general knowledge intact while you do it. A full fine tune is a bigger commitment, in compute and in risk of catastrophic forgetting on tasks you didn’t train for, but it’s the only one of the two that can genuinely rewrite what the model knows from the ground up. Pick based on what you’re trying to change and what you’re willing to pay to change it, not based on which one sounds more serious.

If you want more breakdowns like this on the tools and techniques actually worth your GPU hours, come find us at AI Tool Gazette.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →