Open weights or a paid API: what I actually run where
An upscaling tool on this machine spent about three weeks running on the wrong GPU.
It worked the whole time. Correct output, jobs finished, nothing errored. It was slow in a way I had lazily filed under “old hardware, what do you expect”. The tool picks device 0 by default. Device 0 here is the integrated graphics inside the CPU, and the card I actually care about, a secondhand 1080 Ti with 11GB on it, is device 1. The fix was one flag.
Nothing warned me. There is no status page for your own desk.
That half of the argument never makes it into the comparison posts, so this is the version with the failures left in.
The split as it stands today
A 7B open model runs locally and takes the high volume, low stakes work. It reads a transcript and writes the title and the tags that go under it. Dozens of times a week, every week. Nothing it produces reaches a human before I have edited it.
Three things make that job right for a local model. The output is short, the quality bar is competent rather than brilliant, and the volume is high enough that per call pricing would start to register.
The frontier model gets the other kind of work. Long reasoning over messy input, anything that has to hold a thread across many steps, anything where the failure mode is confident plausible nonsense I then have to catch by hand.
I pay that bill without much argument. The real comparison is the token price against how many hours it would take me to repair a worse answer, and my hours are the expensive input here.
Open is three separate claims
The word carries three meanings, and most arguments about it happen because two people picked different ones.
Open weights means you can download the file and run it. That is the property I use and the one nearly everybody means.
The licence is a different question, and licences vary far more than people assume. Some of the most downloaded models carry a monthly user ceiling above which you need a separate agreement, or a clause forbidding you from training a competitor on their output. Read it before you build a business on top. For a weekend project it will never come up.
Open training data, where the dataset itself is published, barely exists and almost nobody needs it.
When somebody says open source model in 2026 they nearly always mean the first one, under a licence they have not opened.
An eight year old card changes the question
Almost every local model guide is written on current hardware and quietly assumes it. Two examples from this desk.
The speech to text component refuses to run in half precision here. The card predates the tensor hardware that makes fp16 fast, so the library errors out and I run it in int8 instead. The silicon decides that one, no setting will save you. I found out the usual way, by following a tutorial and watching it fail. Model support and hardware support are separate questions and only one of them is written on the model card.
The second is thumbnails. I generate candidates locally with an image model, it works, and it is slow enough that the workflow around it had to change. On a current card you sit there and iterate: four images, look, tweak the prompt, four more. Here I queue a batch and go do something else for twenty minutes. Same capability, different working rhythm, and I have never seen a local image generation writeup mention that the rhythm changes at all.
I traded money for latency. Latency has its own cost, it just never turns up on an invoice.
The cost number that actually matters
Most comparisons put a per token price next to an hourly GPU rental price and declare a winner. Wrong unit.
What you want is cost per request at your real utilisation, which means the arithmetic means nothing until you know your volume. Rented hardware bills you for the hours whether you use them or not. Say two dollars an hour: that is about $1,440 a month serving ten requests or ten million. At high steady volume the number gets embarrassing for per token pricing. At my volume it would be absurd, and I would be renting silicon to watch it idle.
My own position is unusual in a way worth stating plainly. The card was already in the machine, bought years ago for something else entirely. So the hardware is sunk cost and a local call costs me electricity. That makes self hosting look far better here than it would for anyone pricing a new build, and I would rather say so than pretend my numbers generalise to yours.
The line item nobody puts in the spreadsheet
Your hours.
The three weeks of the wrong device. The afternoon lost to the precision problem. Every bit of maintenance an API bill would have quietly absorbed on my behalf.
Those hours are real money and for most people they are the largest number in the exercise. Below genuine sustained volume, paying somebody else comes out cheaper even when the token price looks offensive.
Where the quality gap actually sits
The capability difference is real and it is spread very unevenly.
On the short structured work my local model does, I could not pick the winner in a blind test. Titles and tag lists are flat ground.
It opens up sharply once a task needs sustained reasoning: many steps, held context, spotting its own mistake at step six and recovering from it. That is where the frontier models pull away and where I have never once regretted paying.
So the useful question has nothing to do with which model is smarter in general. Ask whether your hardest real task sits on the flat part of that curve or the steep part. Published benchmarks will not answer it, because their inputs are not your inputs. Take ten of your own hard cases and run them through both. It costs an afternoon and it is the only evidence that applies to you.
Context length is the quiet dividing line
This decides more architectures than people admit.
Hosted models take enormous inputs now, and that changes what you can build. You throw a large pile of text at the thing and ask a question about it, and the entire retrieval layer you would otherwise have designed, built and kept alive never gets written.
Open models have improved a lot here. The memory cost of a long input on your own card is still yours to pay, and it climbs fast. On 11GB it climbs fast enough to end the conversation. A design that is trivial against a hosted endpoint can be the single most expensive thing on your own machine.
If your product depends on stuffing a lot of text into every call, price that exact pattern before you commit to self hosting. It is where the two cost curves separate hardest.
The deprecation tax only one side pays
Models get retired.
You tune prompts against one model’s particular quirks, you ship, you sign off on the behaviour, and some months later that version has a shutdown date on it. Now you are re running evals and re qualifying behaviour you had finished thinking about, on somebody else’s schedule.
Weights on my own disk do not do this. What I run today will behave identically in three years, because nothing about it can change unless I change it.
How much that matters depends on how long your thing has to keep behaving the same way. For a demo, irrelevant. Attach a support commitment to it and this becomes one of the stronger arguments for keeping part of the stack local.
From a standing start
Four situations, four different answers:
- a hard data residency rule or a regulator in the room, in which case run it yourself and no pricing argument survives that conversation
- low or spiky volume, so use a hosted model, because idle rented hardware will eat you alive
- high steady volume on a narrow task, so run a small model yourself and take the savings
- genuinely hard reasoning, so pay for the frontier model and stop optimising
Most systems that have been alive a while end up doing all four, routed by task. Mine does. That is what happens when you let each request go to the cheapest thing that can actually do it, and it reads as fence sitting only if you have never had to pay for either side.
The layer I should have written first
Put a thin interface between your application and whichever model answers, and keep an eval suite you can point at anything.
I did not do this at the start. Swapping a component out later cost me more than it should have, because prompts had been tuned to one model’s habits and nothing was measuring whether the replacement was better or worse. I was comparing a vibe against another vibe on a change I could not cheaply undo.
That one habit is worth more than getting the initial choice right, because it makes the choice reversible. Given how fast the pricing and the model list move, reversible is the property that keeps paying.
The current comparison tables and pricing live here.