Writing an Adapter So You Can Switch Providers
I built an adapter layer for payment providers a while back, and the surprise was how little of the work was the payment part. Almost all of it went into reconciling how each provider described the same events. Doing the same exercise later for model providers produced an identical shape of problem, so the mistakes transfer.
Why bother, honestly
People say vendor lock-in and picture being unable to switch. That is rarely what happens. Switching stays possible and costs three weeks, so you never do it, and then a pricing change or a capability gap or an outage arrives that you cannot route around.
The lock-in is usually not the model either. Models are more interchangeable than the marketing on either side suggests. What locks you in is everything built around one provider’s particular shape: how you construct a request, read a response, handle tools, interpret errors. That shape spreads across a hundred small places in your codebase, and pulling it out later is the three weeks.
An adapter buys you exactly one thing: the option to move stays cheap. I have kept second providers configured and switched off for months without using them and still think it was worth it, because the once I needed to move, I moved in an afternoon.
What actually differs
More than people expect.
Message format, meaning how a conversation is represented, how the system instruction is passed, whether it is a separate field or a message with a special role, how multi-part content is structured.
Tool calling differs most, and it is where naive adapters break. The schema describing a tool, how the model signals it wants one, how you pass a result back, whether multiple calls arrive at once, what happens when it calls a tool that does not exist. All different across providers. An abstraction designed around text in and text out falls apart here.
Streaming differs: event types, how partial content arrives, how tool calls arrive mid-stream, how the end is signalled.
Token counting differs, which matters more than it sounds, because cost tracking and context management both depend on it and the same text does not tokenise to the same number.
Stop reasons differ. Why generation ended, whether it hit a limit, whether it was cut off, and the vocabulary for saying so.
Errors differ most annoyingly of all.
Start naive, then fix errors properly
The naive version is a function mapping your internal request shape to a provider, plus another mapping the response back. An hour of work, and genuinely the right starting point. It handles plain text generation fine, then breaks on tools and streaming, and breaks worst of all on errors.
Errors are hard because a status code is not a taxonomy. One provider returns 429 for both “you are going too fast” and “you are out of credit”, completely different situations, one of which you retry and one of which you must not. Another returns 503 for an overload you should retry while a third uses 500 for the same condition. Some put the retry delay in a header, some in the body, some nowhere.
What your adapter should produce is not the provider’s error but a small set of your own categories, with every provider mapped onto them. At minimum: retry after a delay, retry immediately with a smaller request, do not retry because the input is wrong, do not retry because the account is wrong, and something unexpected happened.
That last one matters. Leave an escape hatch for errors you have not seen and log them loudly with the raw provider response attached. The whole point is that you will meet shapes you did not anticipate, and you want to learn that from a log rather than from silently retrying something you should not.
Register the second provider disabled
This is a process point rather than a code point, and it is the one I would argue hardest for.
On the payments side the adapter went live with the incumbent still doing everything, and the alternatives configured, wired, and switched off. My rule was that flipping the default was not allowed until reconciliation, the thing checking what actually happened against what we recorded, had caught up and agreed for the new path.
That discipline separates an adapter that reduces risk from one that adds it. On the day you switch, you are changing provider and exercising a new abstraction for the first time at the same moment, and when something breaks you cannot tell which caused it. Registering early and dark splits those two events apart.
The stronger version is shadow mode. Send the same request to both, use the incumbent’s answer, record the other. You pay twice for a while, which is why you sample rather than mirror everything, and you get real comparative data on quality, latency, and failure rate before committing anything.
Never match on a prefix of an identifier
A specific trap from the payments work that generalises exactly.
Providers format references differently, and there is always a temptation to find your record by matching the beginning of a reference string. It looks unique. It is not. Formats overlap, providers reuse patterns, and one day two things match and you attribute an event to the wrong record.
Store the provider name and the provider’s identifier as separate fields and match both, exactly. One extra column, and it removes a class of bug that is genuinely awful to debug afterwards.
Do not build a lowest common denominator
The tempting move when providers differ is to expose only what they all support. Clean interface, and it quietly throws away capabilities you are paying for.
When one provider does something useful the others do not, I would rather the adapter expose it and fail loudly on a provider that cannot, than pretend the capability does not exist. A clear error at the boundary beats silently degraded behaviour nobody notices.
Cost and latency do not port either
The headline price per token is the easy part. The same text tokenises differently, so a request of a given size on one provider is a different size on another, and discounts you relied on may not exist on the other side. Caching discounts especially are structured differently everywhere and sometimes require shaping requests a particular way to qualify. A switch that looks cheaper on the published rate can cost more in practice, and you find out by running real traffic.
Latency is the same, and the number that matters depends on the job. Streaming to a person makes time-to-first-token the thing they feel, with total time barely registering. A batch job nobody watches makes throughput everything and first-token time irrelevant. Providers differ on those independently, so one can win on one measure and lose the other, and choosing on a single average will mislead you.
Record which provider served each request
Easy, and almost nobody does it. Store it next to the output.
It sounds trivial and it is the difference between explaining a quality complaint during a rollout and guessing. Whenever two providers are live at once, un-attributed output is nearly useless for diagnosis.
The end state may not be switching
Once the seam exists, routing becomes possible. Send cheap high-volume work to one provider and hard work to another, route by task type, or fail over automatically when one starts erroring.
That is a better use of the abstraction than a one-time migration, and it is only available because you built the seam. Hold off until the basic path is solid, but that is where this leads.
Prompts do not port
The biggest thing people miss. The same prompt sent to a different model produces different output, sometimes slightly, sometimes substantially, and the differences concentrate exactly where you care: formatting, how strictly instructions are followed, behaviour at the edges.
So switching providers is not a configuration change. It requires re-running whatever evaluation you have, and without one you are switching blind and your users will tell you.
The adapter is necessary and nowhere near sufficient. It makes the plumbing swappable and does nothing about the behaviour, and behaviour is what your product depends on.
Test it once, run it against everything
An adapter is unusually testable and people skip it.
Write one set of tests and run them against every provider you support. Same input, asserting on your internal shape rather than the provider’s, so the tests are identical across all of them. That gives you a conformance suite: a new provider is supported when it passes, and a provider changing something under you fails a test instead of surprising you in production.
Include the error cases. Deliberately send a request that is too large, or an invalid tool schema, and assert you get back the right category. Those are exactly the paths that never get exercised until the day they matter.
The cost of the layer
Be honest about it. An adapter is another layer between you and the thing actually running, so a failure now has two places to look. It also tends to accumulate provider-specific special cases inside itself, at which point it is a pile of conditionals wearing an abstraction.
If you only ever intend to use one provider and would genuinely accept the rewrite if that changed, skip it. It earns its keep once you actually have two.
My preference: build the seam, not the framework. One module owning every call to a provider, an internal request and response shape of your own, and your own error categories. Resist configuration systems and plugin architectures around it, because the value is entirely in having one place to change, and the seam alone gives you that.
Limits
I have done this properly for payment providers and partially for model providers, and payments is the case I have watched in production long enough to trust. Providers also keep converging on each other’s conventions, so some specific differences above will be less true in a year.
The structural advice does not depend on any of that. One seam, your own error taxonomy, register disabled and prove it before flipping.
For more breakdowns like this on how AI tools actually work under the hood, head back to the AI Tool Gazette homepage.