← all articles

The real math on self hosting a model versus paying an API

The spreadsheet everyone gets wrong

Every few months someone on a team asks why we’re paying an API provider by the token when we could “just run it ourselves.” The pitch usually comes with a napkin calculation: buy a GPU, run an open weight model, stop paying per request. On paper the GPU pays for itself in a few months. In practice most teams that try this end up back on an API within a quarter, quietly, without a retro.

The reason isn’t that self hosting is a bad idea. It’s that the napkin math only counts the sticker price of the card and skips everything that makes inference actually work in production. I’ve run both paths on real workloads and paid real invoices for both, so here’s the version of the math that includes the parts that get left out.

What the API bill actually buys you

When you pay per token, you’re not paying for compute. You’re paying for compute plus someone else absorbing utilization risk. A GPU that’s busy 100% of the time is cheap per request. A GPU that’s busy 8% of the time, which is closer to what most internal tools look like, is brutally expensive per request, because you paid for the other 92% too.

API pricing is built around the provider running huge fleets at high utilization and slicing off marginal capacity to you. That’s the entire economic trick. You’re renting a fraction of a shared, always-warm cluster instead of owning an idle one. For spiky, unpredictable, or low volume traffic, this is very hard to beat with your own hardware, because your hardware doesn’t get to average its idle time across a thousand other customers.

What self hosting actually costs

Break the real cost of self hosting into four buckets, because the GPU purchase price is only one of them.

Hardware or rental. A single consumer card with enough VRAM to run a mid sized open weight model comfortably runs you somewhere in the low thousands of dollars if you buy outright. Cloud GPU rental by the hour avoids the upfront cost but reintroduces the utilization problem: you’re billed whether or not a request comes in, unless you’re diligent about spinning instances down, and spinning down adds cold start latency that breaks anything interactive.

Power. A GPU pulling several hundred watts under load, run continuously, adds up to a real line item on an electricity bill every month, and that’s before you count cooling if you’re running enough cards to notice the heat in the room. This is the cost everyone forgets because it doesn’t show up as a single invoice, it shows up as a slightly higher number on a bill you were already paying.

Depreciation and failure. GPUs degrade, drivers break, and the model that was state of the art when you bought the hardware is not the model you’ll want to be running in a year. Self hosting locks you into a hardware generation. The API model swap is an endpoint change. The self hosted model swap might mean your VRAM budget no longer fits the model you want.

Your time. This is the one that actually kills most self hosting projects. Someone has to manage the inference server, handle batching so throughput doesn’t collapse under concurrent requests, watch for memory leaks, patch the serving stack, and be on call when the box falls over at 2am during a traffic spike. That person’s time has a cost even if it’s your own weekend. On a small team, the engineering hours spent keeping a model server healthy are frequently worth more than the API bill they were trying to avoid.

A worked example, with the assumptions stated

Say you’re running a workload that would cost you around 400 dollars a month on an API, at whatever the current per token rate is for the model tier you need. That’s your number to beat.

Now price the self hosted alternative honestly. A GPU capable of running the equivalent open weight model at usable speed, bought outright, costs a few thousand dollars up front. Spread over two years of expected useful life before you’d want to upgrade anyway, that’s a real monthly cost even before power. Add continuous power draw at typical residential or colo electricity rates and you’ve added another real monthly cost. Add even a few hours a month of your own time to keep the thing running, valued at whatever your time is actually worth, and the “free” self hosted option has a monthly cost that’s frequently in the same range as the API bill you were trying to avoid, sometimes higher, once utilization is anything less than constant.

The math flips hard in the other direction once volume climbs. If that same workload were actually running at 5,000 dollars a month on the API, the fixed costs of self hosting stay roughly the same while the API cost keeps scaling linearly with usage. High and steady volume is where owning the hardware wins, because you stop paying a per unit margin on top of compute you’re using constantly anyway. The crossover point depends entirely on your actual usage pattern, not on a general rule, which is why the spreadsheet has to use your numbers, not a vendor’s marketing numbers.

The variable nobody puts in the spreadsheet: quality gap

The other cost that doesn’t show up as a dollar figure is capability. The frontier hosted models are, as a category, ahead of what you can comfortably run on a single card or two at home. If your task needs strong reasoning, long context handling, or reliable tool use, an open weight model sized to fit your hardware budget may simply not be good enough, and you’ll spend more engineering time working around its weaknesses than you saved on inference cost. That’s a real cost, it’s just paid in debugging time and prompt engineering instead of an invoice.

Where this gap matters less is narrow, well scoped tasks: classification, extraction, summarization of a fixed format, or anything where you can fine tune a smaller open model specifically for that one job. A model tuned tightly for a narrow task can match or beat a general purpose frontier model on that task, because it doesn’t need the general capability, it needs the specific one. This is the strongest case for self hosting: not “replace the API everywhere,” but “replace the API for this one high volume, narrow, well understood task.”

Where each option actually wins

Self hosting wins when you have sustained, predictable, high volume traffic on a task narrow enough that a smaller model handles it well, and when you already have someone who can operate infrastructure without it becoming their whole job. It also wins outright when data can’t leave your environment for compliance or contractual reasons, in which case the cost comparison stops mattering because the API isn’t an option at all.

The API wins for bursty or low volume traffic, for anything that needs frontier level capability, for early stage products where the traffic pattern isn’t known yet, and for teams without spare engineering capacity to babysit a GPU box. It also wins by default any time the true hourly cost of the person who’d be maintaining the self hosted setup is high, because that cost rarely shows up until after the migration.

How to actually decide

Don’t start from “self hosting is cheaper” or “APIs are cheaper.” Start from your actual traffic: requests per day, tokens per request, and how spiky the pattern is. Multiply that against your current API rate to get a real baseline number. Then price hardware or rental at realistic utilization, add power, and add a genuine estimate of the hours per month it’ll take to keep it running, valued honestly. Compare the two totals over a year, not a month, because hardware costs are front loaded and API costs are not.

If the gap is close, stay on the API. Close gaps get eaten by the maintenance time you didn’t count. Only move when the volume is high enough and steady enough that the fixed costs of self hosting are clearly, not marginally, cheaper than what you’re paying per token today.

If you want more breakdowns like this that skip the vendor pitch and show the actual tradeoffs, you can find them on the AI Tool Gazette home page.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →