← all articles

When fine tuning actually beats a better prompt

fine-tuning prompt-engineering llm-costs distillation

Sixty examples sitting in a folder, and I was certain that was plenty.

I wanted a model to write in a particular voice. Not a wild ask. I had months of material, I pulled out sixty pieces I liked, I paid for a training run, and what came back wrote in the average of those sixty voices. Which reads about how you would expect. A committee.

A longer prompt with four hand picked examples beat it outright, and finding that out took an afternoon. The training had taken a week.

The other time I paid for it, the thing was obviously correct inside three weeks and I would do it again tomorrow. Same technique, opposite result, and the difference had nothing to do with the model or the vendor.

Most people who ask me whether to fine tune are standing in the first situation and think they are in the second.

Training moves the shape, not the contents

Worth being blunt here, because a wrong mental model at this step produces very expensive decisions later.

Fine tuning adjusts weights so that output tends toward a particular shape. Tone. Format. Length. How a response opens. Which word it reaches for when it has a choice.

It does not install facts you can look up. Train on a thousand support tickets and you get something that writes like your support team. What you do not get is something that can answer a question by finding the correct ticket, because the information went in as a statistical nudge and comes back out the same way, smeared against everything else in the set.

So the first filter takes about ten seconds. If the complaint is that the model does not know something, you have a retrieval problem, and training is a slow and expensive way to fail at it.

The arithmetic that decides it

Here is the part almost nobody works out, and above a certain volume it settles the question on its own.

A properly engineered prompt with four or five examples in it runs to two or three thousand tokens. That rides along on every single request. Every call, forever, whether the question was hard or trivial.

Fine tune the behaviour in and the prompt collapses to a line or two. A few hundred tokens instead of a few thousand, usually on a smaller model, which moves the per token price as well.

At ten calls a day the difference is a rounding error and you can stop reading. At a million calls a month it is the entire business case, and it makes the training invoice look like a typo.

Latency follows the same logic when a person is waiting. A small model with a short prompt starts producing sooner, and in front of a user that is worth real money.

So the honest test is volume crossed with stability. Low traffic, or a task whose definition keeps moving, stay on the prompt. High traffic on something that has held still for months, run the numbers properly and they will often say train.

The ladder, and why the order is where the money goes

There is an order to this. Skipping rungs is how people lose a fortnight.

Write a better prompt. This sounds dismissive and it is the highest return activity available to you. Most prompts I get shown are a paragraph of vague intent where they should be a page of specifics, edge cases and explicit refusals.

Put examples in the prompt. Three or four demonstrations of the exact input and the exact desired output move behaviour further than people expect. This is the rung skipped most often and it is the one that most often ends the problem.

Split the task. One call doing four things badly becomes four calls each doing one thing well, and each of those is easier to instruct and far easier to check.

Add retrieval, if the gap is knowledge rather than behaviour.

Fifth, and only fifth, consider training.

The order matters because every rung is cheaper and easier to undo than the one above it. A prompt change costs a minute. A training run costs a day, and the dataset behind it costs a week.

Two options that sit between the last two rungs

The first is generating your data with the model you cannot afford to run.

You already have something that does the task well. The problem is the bill for running it on every request. So run it on a few thousand representative inputs, read the outputs, and that becomes your training set. People call this distillation, and it takes the cost problem and the data problem down together.

The reading is not optional and it is the actual project. A set produced by a large model contains that model’s mistakes, and training on those teaches the mistakes deliberately, permanently, at your expense. Somebody has to sit and go through them. Budget for that person before you budget for anything else.

The second is routing, and I would try it before training in almost every case. Easy cases go to the small cheap model, hard ones escalate to the expensive one. You get most of the cost benefit with none of the maintenance, you can build it in an afternoon, and when it goes wrong it goes wrong loudly. That last property is worth more than the savings.

The invoice is a person, not a graphics card

Cost surprises people in both directions.

The training run itself is usually cheap. A small model on a few thousand examples is often tens of dollars. That is the number everybody fixates on and it is the least interesting one on the page.

The spend that hurts is the days somebody burns assembling and checking data, the scoring rig you have to build before you can tell whether anything improved, and the retraining you will do when the base model gets retired, which is the same work again with less enthusiasm.

Budget a person for a week or two. If you were bracing for a compute bill, you have the shape of this backwards.

What you own afterwards

This is the part I regret about one of the two.

A fine tuned model is a dependency with your name on it. The base underneath gets deprecated on somebody else’s schedule, and when it does you retrain, which means the dataset has to still exist, still be understood, and still be reproducible by whoever is around at the time. I have watched a team lose the ability to rebuild their own model because the one person who assembled the data had left.

It also freezes your behaviour at the moment you trained. General models improve on their own. Yours improves when you do the work again.

Evaluation gets harder too, because you can no longer measure against a public baseline that everybody documents. You measure against your own previous version, which is a far weaker signal than it feels like.

Treat it as taking on maintenance. That framing on its own would have stopped me the second time.

Consistency beats volume, and it is not close

If you go ahead anyway, the training run is the easy part. It is a command and a wait.

The work is several hundred examples that are correct, that agree with each other, and that look like the input you actually get in production.

Two hundred examples that agree beat two thousand that contradict. Contradictions teach the model that both behaviours are acceptable, and you get an unpredictable blend of the two. That is precisely what my sixty voice samples produced. Nobody wrote them to disagree. They came from different months and different moods, and the model averaged them faithfully.

The examples also have to be messy in the way your real inputs are messy. Train on clean well formed cases and you get something that falls over on the rest, which is most of them.

Build the scoring before you train

The most self deception I see lives in this step, mine included.

Build the evaluation first. Test cases with known correct answers, scored automatically, run against the base model before you touch anything, so you have a number to beat.

Without that you will look at ten outputs, decide they seem nicer, and ship. I did exactly that. The trained model felt better, I could not demonstrate it to anyone including myself, and I had no way to notice the specific category of input it had quietly got worse at.

Check the regressions on purpose. A model pushed toward one behaviour gives something up somewhere else. The question is never whether it improved on the target. It is what it surrendered elsewhere, and you only see that if you were measuring elsewhere before you started.

Five conditions, and I have met them once

Before training, all of these should be true:

  • prompting has plateaued below a number you can state out loud
  • your call volume makes the token arithmetic favourable
  • several hundred consistent examples already exist, without anyone creating them for this project
  • the task definition has not moved in months
  • you are willing to own the retraining forever

Five conditions, and the list is demanding on purpose.

The time I met it, the task was a narrow classification step running over every item in a content pipeline, tens of thousands of calls a month, a large model doing something a much smaller one could clearly handle. The prompt was long and identical every time. The volume was flat. The correct answers already existed because somebody had been labelling by hand for months for an unrelated reason. Training that was obviously right and it paid for itself in a few weeks.

The time I did not, I had sixty inconsistent examples and a target based on taste with no number attached to it. There was nothing to score against, so there was no way to know I had failed until I sat down and read the output.

The deciding factor was never the technique. It was whether the task was narrow, high in volume, and already labelled by somebody for their own reasons. Miss any one of those three and the answer is to go back to the prompt. Current model pricing and the comparison tables I keep updated are here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →