← all articles

What actually changes when you switch LLM providers

Every few months someone on a team decides it’s time to switch LLM providers. Usually it’s triggered by a price change, a rate limit that’s choking production, or a new model that tests better on the team’s own eval set. The pitch is always “it’s just an API call, we’ll swap the base URL and the key.” It never is. Here’s what actually moves when you change providers, based on what breaks in practice, not what the marketing docs promise.

The API contract is not the hard part

Most providers converged on a similar shape: you send a list of messages with roles, you get back a completion, you can stream it as server-sent events. If your integration is a thin wrapper around one HTTP call, swapping the endpoint and auth header is genuinely a few hours of work. That’s the part everyone budgets for, and it’s the part that was never going to bite you.

The real cost shows up one layer down, in everything you built around that call to make it reliable and cheap.

Prompt formatting doesn’t transfer cleanly

System prompts don’t behave the same way across providers. Some treat the system message as a hard instruction that gets weighted heavily throughout the conversation. Others treat it more like a prefix that competes with everything else in context and can get diluted in long conversations. If you’ve spent months tuning a system prompt against one model’s quirks (the exact phrasing that stops it from padding answers with disclaimers, the specific instruction that gets it to output clean JSON without a markdown fence around it) none of that tuning carries over. You’re not migrating a prompt, you’re re-writing one and re-testing it against the same edge cases that broke it the first time.

This is worse than it sounds because the failures are silent. The new model doesn’t throw an error when it ignores half your system prompt. It just produces slightly worse output, and slightly worse output is exactly the kind of regression that doesn’t show up until a user complains.

Tool calling schemas are not interchangeable

If your app uses function calling or tool use, this is where migrations actually stall. Every provider has its own idea of how a tool schema should be declared, how the model signals it wants to call a tool, and how you’re supposed to feed the result back in. The JSON schema dialect you use to describe a function’s parameters, the required fields, how nested objects and enums get validated, none of it is standardized between providers even though it looks superficially similar.

Worse, models differ in how eagerly they call tools versus answering directly, how they handle a tool call that returns an error, and whether they can chain multiple tool calls in one turn or need a round trip per call. If you built retry and validation logic around one provider’s failure modes (say, a model that occasionally emits malformed JSON and needs a repair pass), that repair logic is tuned to a specific way of being malformed. A different model breaks differently, and your repair pass might not catch it, or might trigger on output that was actually fine.

Context windows and truncation behavior

Context window size gets quoted like a single number, but what matters more for a migration is how a provider handles content near the edge of that window. Some SDKs will hard-reject a request that overflows the limit. Others will silently truncate from one end. If your RAG pipeline or chat history management assumes rejection-and-retry, and the new provider truncates instead, you can end up shipping requests that quietly drop the oldest turns of a conversation or the last retrieved chunk, with no error to catch it.

There’s also a practical difference between a provider’s advertised context window and the length at which output quality actually holds up. A model can accept a huge input and still get worse at following instructions buried in the middle of it. That’s not something you can take on faith from a spec sheet, it’s something you have to notice by running your own real conversations through it, not a synthetic long-context benchmark.

Rate limits and retry logic

Rate limits are usually structured around tokens per minute and requests per minute, but the tiering, burst allowance, and how limits scale with usage history all vary. A retry-with-backoff strategy tuned against one provider’s 429 behavior (how long the backoff needs to be, whether the limit is per-key or per-org, whether concurrent requests share a bucket) doesn’t automatically fit another provider’s limits. If you run a queue that assumes a certain sustained throughput, moving providers can mean redoing your concurrency settings and backoff curve from scratch, not just changing which errors you’re catching.

Pricing structure, not just the price

The headline number people compare is dollars per million tokens, but the structure underneath it matters more for a real bill. Input and output tokens are priced differently everywhere, and the ratio between them isn’t consistent across providers, so a workload that’s output heavy can rank very differently than one that’s input heavy depending on which provider you’re comparing. Some providers also offer prompt caching that discounts repeated context, which matters a lot if you’re sending the same long system prompt or the same retrieved documents on every call. If your current setup leans on caching to keep costs down, and the new provider’s caching works differently or isn’t available for your use case, the real cost delta can look nothing like the sticker price comparison.

I’m not going to quote specific dollar figures here, because they change often enough that anything I write today is stale by the time you read it. The point is: get the actual pricing page for your specific usage pattern (cached vs uncached tokens, input vs output ratio) before you decide a switch saves money. A per-token comparison alone will mislead you.

Output style and refusal behavior shift

Every model has a personality, for lack of a better word, and it’s shaped by its own training and safety tuning, not by the API contract. One model might refuse a request that’s borderline (medical, legal, security related) that another model answers directly with appropriate caveats. One might default to long, hedged answers while another defaults to terse ones. If your product surfaces model output directly to users, this shift is user facing, not internal. You will get support tickets about tone before you get tickets about anything technical.

If you have any kind of safety or content moderation layer sitting in front of the model, it was tuned against one model’s failure patterns. It needs to be re-evaluated, not assumed to carry over.

If you’re running RAG, don’t forget the embeddings

This is the one people forget most often. If you’re switching providers and that provider also handles your embeddings, you can’t just swap the generation model and leave your existing vector store as is. Embedding vectors from one model aren’t compatible with another model’s embedding space. Cosine similarity scores computed against a different embedding model are meaningless. If embeddings are part of the switch, you’re re-embedding your entire corpus and re-indexing, which for a large document store is a real compute and time cost, not a config change.

Even if you keep your embedding provider fixed and only switch the generation model, it’s worth double checking that assumption is actually true in your codebase and not baked into the same SDK call as generation.

What actually needs to happen before you flip the switch

The honest checklist looks like this: rebuild and re-test your system prompt against the new model’s actual behavior, not the old one’s. Rewrite tool schemas and re-test the failure paths, including malformed output. Re-tune retry and backoff logic against the new rate limit behavior. Get real numbers on your own workload’s input/output ratio and caching eligibility before trusting a headline price comparison. Re-run your safety and moderation checks against the new model’s refusal patterns. And if embeddings are involved, budget the time and cost to re-index.

None of this means don’t switch providers. Prices move, rate limits get painful, and a model that’s a better fit for your workload is worth chasing. It just means the switch is an integration project with its own testing cycle, not a config change, and the teams that get burned are the ones who budgeted for the API call and not for everything wired around it.

If you want more of this kind of grounded, no-hype breakdown of how these tools actually work under the hood, you can find the rest of our coverage on the AI Tool Gazette home page.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →