Cost · 5 minute read
AI Migration Cost: Moving Models, Providers and Platforms
AI migration cost is driven by behaviour differences rather than by interface changes. Prompts tuned for one model perform differently on another, evaluations must be rerun, embeddings must be regenerated if the model changes, and downstream logic tuned to the old behaviour needs revisiting.
Migrating between models or providers looks like an interface change and is a behaviour change. The API call is trivial to update; everything that was tuned to how the previous model behaved is not. This guide covers what that actually involves, drawing on FISTA Solutions' AI enablement work. It complements what is champion-challenger testing and what is a prompt template.
Why are behaviour differences the main cost?
Because the system was calibrated against the old behaviour. Prompts were refined until they produced the right output from that model. Confidence thresholds were set from that model's calibration. Output parsing was written for its formatting habits. Guardrails were tuned against its failure modes.
A different model does all of those differently, and every one of those calibrations needs re-measuring. The interface change takes an afternoon; the recalibration takes considerably longer.
| Element | Ports cleanly | Requires rework |
|---|---|---|
| API call structure | Mostly | Minor |
| Prompt wording | No | Yes, per model |
| Confidence thresholds | No | Re-measure |
| Output parsing | Sometimes | Verify |
| Embeddings and index | No | Full re-embedding |
| Evaluation results | No | Rerun in full |
Do prompts port between models?
Not cleanly. A prompt refined over weeks against one model frequently performs noticeably worse on another — following some instructions less reliably, producing different formats, refusing different requests.
The arrangement that works maintains provider-specific prompt versions, each evaluated against the same suite. That is more work than a single prompt and it is what makes a multi-provider or migrating estate actually function.
What makes re-embedding expensive?
Scale and coupling. Vectors from different embedding models occupy different spaces and cannot be compared, so changing the embedding model means regenerating every vector in the corpus and rebuilding the index.
For a small corpus that is an afternoon. For millions of documents it is a project with meaningful compute cost, a cutover plan, and a period during which both indexes exist. Planning that path before the corpus grows large is what keeps the option open. See what is an embedding model.
What downstream logic is affected?
More than expected. Confidence thresholds for routing and abstention. Output parsing and schema validation. Retry and fallback logic. Guardrail patterns tuned against specific failure modes. Cost and latency assumptions built into capacity planning.
Each was set empirically against the previous model, and each needs re-measuring rather than assuming it transfers.
Why must evaluation be rerun in full?
Because a model change affects everything, not a subset. A partial evaluation that samples a few categories will miss the category where behaviour diverged most, and that category is exactly what will surface in production.
Full re-evaluation is also the only way to make the migration decision on evidence rather than on the provider's benchmark claims.
How is future migration cost reduced?
By abstracting the model call behind an interface so application code does not depend on a provider's specifics. By keeping prompts as versioned artefacts with evaluation results attached, so a new provider's version can be developed and compared. And by maintaining an evaluation suite that can be rerun on demand.
Those three turn a future migration from an investigation into a measurement exercise. They cost little to establish early and are expensive to retrofit.
What about platform migrations?
Broader and frequently harder, because a platform migration touches orchestration, observability, evaluation tooling, and deployment alongside the model. Those have their own data — traces, evaluation history, prompt versions — that either migrates or is lost, and losing evaluation history means losing the ability to compare against what came before.
What should you do first?
Check whether your application code calls a provider SDK directly. If it does, the first step in reducing migration cost is an abstraction layer, and it is worth doing regardless of whether a migration is planned.
What triggers a migration?
Usually one of four things: a provider deprecating a model version, a material price change, a capability the current provider does not offer, or a residency or contractual requirement that the current arrangement cannot satisfy.
The first is the most common and the least anticipated. Providers retire model versions on their own schedule, and an organisation pinned to a version will eventually be moved whether or not it planned to be. Treating migration as inevitable rather than optional changes how much abstraction is worth building.
How should a migration be sequenced?
By risk. Evaluate the candidate offline against the full suite, then run it in shadow against production traffic to see behaviour on real inputs, then move a small share of live traffic, then expand. Each stage can stop the migration cheaply.
Migrations that switch everything on a date, because a deprecation deadline arrived, are the expensive kind — which is another argument for treating the deprecation calendar as a planning input rather than a surprise.
How FISTA Solutions helps
FISTA Solutions abstracts model calls behind stable interfaces, maintains provider-specific prompt versions with attached evaluation results, plans re-embedding paths before corpora grow, reruns evaluation in full on any model change, and treats downstream calibration as part of the migration rather than an afterthought, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To keep your options open without paying for them twice, message FISTA on WhatsApp, or read what is champion-challenger testing.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why are behaviour differences the main cost?
Because prompts, thresholds, and downstream logic were tuned to how the previous model behaved. A different model follows instructions differently, refuses different things, and produces different formats, and everything calibrated against the old behaviour needs revisiting.
02Do prompts port between models?
Not cleanly. A prompt refined against one model frequently performs noticeably worse on another, and the differences are not predictable from documentation. Provider-specific prompt versions, each evaluated, is the arrangement that works.
03What makes re-embedding expensive?
Scale and coupling. Vectors from different embedding models are not comparable, so changing the model means regenerating every vector and rebuilding the index. For millions of documents that is a project with compute cost and a cutover plan.
04What downstream logic is affected?
Confidence thresholds, output parsing, retry logic, guardrail patterns, and anything calibrated against the previous model's behaviour. Each was tuned empirically and each needs re-measuring rather than assuming it transfers.
05How is future migration cost reduced?
By abstracting the model call behind an interface, keeping prompts as versioned artefacts with evaluation results attached, and maintaining an evaluation suite that can be rerun. Those three make a future migration a measurement exercise rather than an investigation.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.