What happens when a model is deprecated
I gave the migration an hour.
The notice had been in my inbox about two months by then. A date, a page listing what was going, a suggested replacement picked by somebody who has never read my prompts. The code change was one line in one file and took less time than reading the email had.
Then I ran my own inputs through the replacement and spent the next two days going through disagreements by hand.
That gap between the one line and the two days is the whole subject. Everyone budgets for the config change, and the config change is free.
The one dependency you cannot hold still
Everything else in the stack can be frozen. A library version goes in a lock file and stays where you put it. If the maintainer walks away I can fork it. If the registry pulls it I can copy the source into my own repo and carry it around for the next decade, which is ugly and entirely within my control.
A model has an end date somebody else chooses. You cannot fork it, you cannot vendor a copy, and you usually cannot pay to keep the old one warm past the date, because the reason it is going is that the hardware underneath it is wanted for the successor.
Weights on your own disk do not have this problem, which is a real point in favour of running your own. That argument turns on hardware, licensing and how much idle capacity you can stomach, so I have written it up separately and I am leaving it there.
The thing that makes this dependency awkward is that it breaks quietly. Every other breaking change in software announces itself. A signature moves, a type stops matching, the build goes red, someone has to look at it before anything ships. A model swap compiles.
Most of a mature prompt is scar tissue
Nobody writes the prompt that is in production now. It accumulated.
The transcript step here writes titles. That prompt contains a line telling the model not to use a colon in the title, which exists because the previous model put a colon in roughly every second one. It obeyed. Problem solved, line forgotten, eighteen months pass.
The replacement never had that habit, and now reads a fairly emphatic instruction about colons as a signal that punctuation is a big deal in this task. It started avoiding commas too.
Same story in mail triage. Three sentences in that prompt exist because the old model kept filing supplier invoices as customer mail. On the new one those sentences are somewhere between dead weight and an active push in the wrong direction. I am correcting a mistake nobody is making any more.
My tagging prompt had the format restated at the very bottom, because the old model would drift out of the schema about a third of the way down a long input. The new one holds format from the top, reads the restatement as part of the task, and cheerfully repeats the schema inside its answer.
The fix became the bug. Three times, in three different prompts, in the same afternoon.
The average improves and your specific case comes back
Worth internalising, because it explains why a migration feels worse than the release notes promise.
Your tuning was aimed at failures the vendor has since gone and fixed. So the replacement is better on the broad distribution and worse on the narrow set of inputs you personally worked around. Your headline quality goes up. The one case a customer complained about last year quietly returns.
You find that out in production unless you go looking for it first.
The version that arrives with no notice
A retirement with a date on it is the polite version of this. You get warning, you get months, you get a page to read.
The other version is the model behind the name changing while the name stays the same. Same string in your config, different behaviour, nothing deployed on your side, nothing to read.
I lost four days to that one. My tagger started returning two word labels where it had returned one, the renderer looks those labels up in a dictionary and matched none of them, and about forty paragraphs of a ten minute video got generic fallback footage. Latency did not move. Error rate did not move. Every dashboard stayed green, because dashboards check whether the call worked, and the call worked perfectly. It returned a clean, well formed, completely different answer.
No alert could have fired. I had never written one that looks at what came back instead of whether anything came back.
Pin to a date, then build the alarm
If your provider publishes dated version strings behind the floating alias, point your code at the dated one. Best ratio of effort to pain avoided anywhere in this topic, and it is a one line change.
The dated version has its own end date, often a short one. What pinning buys is a schedule you control rather than one you discover after the fact. Worth a great deal. Not permanence.
For the aliases you cannot pin, five frozen inputs in one file, run every morning, diffed against last week’s output. Mine costs a fraction of a cent a day and messages me when a line changes. It says nothing about whether the new behaviour is better. It says the behaviour moved, which is the only thing I want to know at seven in the morning.
The eval set has to predate the email
When the notice lands there is exactly one question: is the replacement acceptable on my work. You answer it with your own inputs and their known correct answers, and building that set properly is a separate job I have written up elsewhere.
The part that belongs here is timing. A set assembled under a deadline gets built from whatever you can remember, and what you remember are the cases you already handle. It agrees with you. It passes everything, then it passes the replacement, and you ship on the strength of a test that was never capable of failing.
Mine is two hundred and forty rows of CSV, assembled in one afternoon out of logs I already had. Mail bodies with the label they should get. Transcript chunks with the tags they should get.
Two prompts out of eleven
There are eleven prompts in production across three pipelines here and they do not carry equal weight.
If the tagger drifts, a video gets worse footage behind a paragraph and I find out in four days. If mail triage drifts, a customer message lands in a folder I open once a fortnight, which is money and a person waiting.
So two of the eleven get checked first and get checked by hand. The other nine get the eval set and a shrug.
Rank yours by what happens downstream when the answer comes back quietly wrong, which has nothing to do with how clever the prompt is or how long you spent writing it. Then write the ranking down somewhere, because under a deadline the instinct is to check whichever prompt you can find fastest, and that is the one you edited most recently, which is almost never the one that matters.
One file instead of six
The model belongs behind an interface you wrote. One module, every call through it, the model name appearing in exactly one place in the repo.
I did not build it that way. Three pipelines grew over months and the name ended up in six files, two of them with the string typed inline because I was in a hurry that day. When it changed I went looking with grep, missed one, and it failed at three in the morning on a schedule I was asleep for.
An hour of work to fix, and I put it off for about a year.
The interface earns its keep past tidiness. It is where the fallback goes when a call fails. It is where you log which model actually answered, which is what lets you prove drift afterwards rather than argue about whether the output feels different lately. And it is the cheap way to run the old and the new against the same live traffic for a day, which is the test that settles the argument.
If you cannot switch in an afternoon, the vendor owns your roadmap
That is the position and I will defend it.
It never feels like a vendor decision from inside. It feels like a quarter where the feature you promised slips, because two people spent three weeks re qualifying prompts against a replacement on a date chosen by someone who has never met them. That is a roadmap written by a provider, and you agreed to it much earlier by letting the swap get expensive.
Switching cost is the number I watch now, ahead of price per token and well ahead of benchmark position, because it decides whether the other two are decisions I get to make at all.
What I have not been through
One retirement and one silent update is a thin sample and I would rather say so than write as though it were ten.
I have never had a fine tuned model retired underneath me. A fine tune is attached to a base, and when the base goes the tune goes with it, so you are back at whatever training data you kept rather than at a prompt you can sit down and edit. That is a project, not a ticket, and I have not paid that bill.
I also have no idea how this plays out under an enterprise agreement with a real retirement clause in it, because I do not have one. Everything above is written from the position of someone paying with a card on an ordinary API account.
What I pin, what sits in the drift alarm, and the interface module I should have written first are all here.