Playbook · 6 minute read
How to Migrate From One LLM Provider to Another
Changing model provider means handling interface differences, behavioural differences, and contractual and data-handling differences at once. The sequence that works is abstracting the interface first, comparing behaviour on your own cases, resolving the commercial and data questions, and cutting over gradually.
Changing model provider involves more than swapping an endpoint. Interfaces differ, behaviour differs more, and the contractual and data-handling questions run on their own timeline. This playbook covers the sequence, drawing on FISTA Solutions' AI enablement work.
When is this worth doing?
When pricing, availability, capability, data handling, or contractual terms make the current provider the wrong choice — or when a resilience requirement means you need a tested alternative regardless.
It is not worth doing for a modest capability difference on a benchmark. Migration costs engineering time and regression risk, and the gain has to be visible on your own cases rather than on a published comparison.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Abstract the interface | A thin layer over provider calls |
| 2. Resolve contract and data | Terms, locations, retention, assessment |
| 3. Compare on your cases | Not on published benchmarks |
| 4. Rework prompts | Formatting and refusals differ |
| 5. Shadow run | Real traffic, no dependency |
| 6. Cut over gradually | Old provider available throughout |
Step 1 — Abstract the interface first
Build a thin layer between your application and the provider: a function that takes your parameters and returns your shape, with the provider-specific handling inside it.
This costs little and makes everything afterwards straightforward. Systems calling provider SDKs directly from a dozen places have made switching a refactoring project rather than a configuration change.
Keep the abstraction thin. Layers that try to unify every provider capability become their own maintenance burden and lag behind what providers offer. See what is a fallback chain.
Step 2 — Resolve the contract and data questions early
Start the commercial and security work in parallel with the technical evaluation, because it usually takes longer.
What data will be processed where, what the provider retains, whether inputs are used for training, sub-processing arrangements, and what your own security assessment requires. In regulated contexts, add the register and contractual provisions your regime expects.
Teams that complete the technical migration and then wait six weeks for a security review have sequenced it wrong. See AI and DORA regulation.
Step 3 — Compare on your own cases
Run your evaluation set against the candidate and compare against your current baseline.
Published benchmarks tell you about general capability on tasks that are not yours. What matters is whether the candidate handles your difficult cases, your formats, and your domain vocabulary at least as well.
Pay attention to refusal behaviour. Models differ considerably in what they decline, and a candidate that refuses a category of legitimate requests your users make is unusable regardless of its scores.
Step 4 — Rework the prompts
Expect real work here. Instruction-following, default verbosity, output formatting, and structured output reliability all differ.
Structured output is where most breakage happens. A prompt producing clean JSON from one model may produce JSON wrapped in commentary from another, and every downstream parser breaks silently.
Re-run the evaluation after reworking, because prompt changes tuned for the new provider may alter behaviour on cases that were previously fine.
Step 5 — Shadow run against real traffic
Send production traffic to the new provider alongside the current one without using its outputs, and compare at volume.
This catches what your evaluation set does not: the unusual inputs, the long tail, and the interactions between real user behaviour and the model's differences.
Watch latency and cost alongside quality. A provider that is cheaper per token can be more expensive per task, and one that is better can be slower than an interactive product tolerates. See how to run a shadow deployment for ai.
Step 6 — Cut over gradually
Move traffic in stages, watching quality, cost, and latency at each step, with the old provider still available.
Keep both live until the new one has seen the tail of your distribution — usually weeks. Rare inputs arrive rarely, and they are where the regressions you did not anticipate live.
Do not cancel the old contract the day cutover completes. The overlap costs money and buys the ability to reverse a decision that turns out badly, which is worth considerably more.
Should you run more than one provider permanently?
Sometimes. Multi-provider operation gives resilience and negotiating position, at the cost of maintaining prompts and evaluation for both.
The pragmatic middle is a primary provider with a tested secondary path that is exercised periodically rather than run continuously. That preserves the option at a fraction of the maintenance cost, and it satisfies most resilience expectations.
What about fine-tuned models?
They do not transfer. A fine-tuned model is specific to its base, and migration means retraining on the new provider with the same data and re-validating the result.
That converts a migration into a project, and it is worth knowing before committing to fine-tuning. Where switching flexibility matters, prompting and retrieval preserve it in a way fine-tuning does not. See what is lora fine-tuning.
Who needs to be involved?
An engineer who can change the integration, someone who can judge output quality, and whoever owns commercial and security relationships.
The commercial and security participation matters from the start rather than at the end, because those paths are the long pole.
How long does it take?
Four to eight weeks for a well-abstracted system, dominated by shadow running and by contract and security review. Longer where the integration is scattered or a fine-tuned model is involved.
What are the common failure modes?
Calling provider SDKs from everywhere. Starting commercial work after the technical work. Comparing on benchmarks. Assuming prompts transfer. Cutting over at once. And cancelling the old contract immediately.
How do you know it worked?
Quality at least matching the baseline on your own cases, no downstream parsers broken, cost and latency within budget, and a tested path back that you did not need.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Count how many places in your codebase call a provider directly. If the answer is more than one, building the abstraction is the first task regardless of whether you migrate.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: a thin provider abstraction built before it is needed, commercial and security paths started in parallel with the technical work, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to run a model migration.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is the hardest part of switching provider?
Behavioural differences rather than API differences. The interfaces are broadly similar and easy to abstract; the way models follow instructions, format output, and refuse requests differs in ways that affect everything built on top.
02Should you build an abstraction layer?
Yes, and before you need it. A thin layer over provider calls makes comparison and cutover straightforward, and it is the practical form of the exit strategy that resilience and outsourcing requirements increasingly expect.
03What non-technical issues arise?
Contract terms, data processing locations, retention policies, sub-processing, and any security assessment the new provider needs to pass. Those run on their own timelines and frequently take longer than the engineering.
04Do prompts transfer?
Rarely unchanged. Verbosity, formatting, instruction-following, and refusal behaviour all differ, and structured output in particular tends to break. Budget prompt rework as a distinct piece of the migration.
05How do you keep the option open long term?
By maintaining the abstraction, evaluating a second provider periodically, and keeping the switch tested. An exit path that has never been exercised is a plan rather than a capability.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.