Playbook · 7 minute read
How to Run a Model Migration Without Breaking Behaviour
A model migration works when you establish what the current model does before changing anything, run the candidate in shadow against real traffic, compare behaviour on cases that matter, and cut over with a tested rollback path. Migrations that skip the baseline discover regressions from users.
Model migrations go wrong for one reason more than any other: nobody established what the old model did before replacing it. Everything else follows from that. This playbook covers the sequence that avoids it, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
When a provider deprecates a model you depend on, when a newer model offers materially better quality or cost, or when a commercial or regulatory reason forces a change.
It is not worth doing for a marginal benchmark improvement. Migration carries real cost in engineering time and regression risk, and a model that scores slightly higher on a public benchmark may perform worse on your specific inputs.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Build the baseline | Capture what the current model does |
| 2. Assemble the comparison set | Cases that actually matter |
| 3. Run offline comparison | First pass on captured inputs |
| 4. Shadow run in production | Real traffic, no dependency |
| 5. Adjust prompts and settings | Same prompt rarely transfers cleanly |
| 6. Cut over with rollback | Gradually, reversibly, watched |
Step 1 â Build the behavioural baseline
Capture what the current model does on real inputs before changing anything: outputs, latency, cost per request, refusal patterns, and failure modes.
Without this, the migration has no reference point. Teams that skip it end up arguing about whether a behaviour is a regression or something the old model did too, and nobody can settle it.
If you do not already log inputs and outputs, start now and wait. A fortnight of captured traffic is worth more than any amount of reasoning about expected behaviour.
Step 2 â Assemble a comparison set that matters
Build a set of cases covering your known failure modes, your highest-volume patterns, your edge cases, and your highest-consequence decisions.
Average quality across a generic benchmark is close to irrelevant. What matters is whether the new model handles your difficult cases at least as well, and whether it introduces new failure modes on cases the old one handled fine.
Include cases the current model gets wrong. A migration that fixes existing problems is worth more than one that merely avoids new ones, and you will only know if you tested for it. See what is an evaluation rubric.
Step 3 â Run the offline comparison first
Run both models over the captured inputs and compare. This is cheap, fast, and catches the obvious problems before anything touches production.
Expect differences in formatting, length, and structure even where the substance is equivalent. Those matter if anything downstream parses the output, which is a common and easily missed source of breakage.
Score the comparison against your rubric rather than reading a sample and forming an impression. Impressions favour whichever model writes more fluently, which is not the same as being right.
Step 4 â Shadow run against real traffic
Send live traffic to the candidate alongside the current model, without using its outputs. Compare at production volume, on inputs nobody selected.
This is where the real differences appear. Offline comparison uses inputs you captured; shadow running uses everything, including the unusual cases that never made it into a test set.
Watch cost and latency as well as quality. A model that is better and slower may be unusable in an interactive product, and a model that is cheaper per token can be more expensive per task if it needs more tokens to get there. See how to run a shadow deployment for ai.
Step 5 â Adjust prompts and settings
Expect to change prompts. They were tuned against the old model's behaviour, and the same text on a different model frequently produces different verbosity, formatting, refusal behaviour, and instruction-following.
Budget this explicitly rather than treating the migration as a configuration change. A migration where prompts transfer unchanged is unusual enough that discovering it is a pleasant surprise rather than the plan.
Re-run the comparison after adjustment. A prompt tuned for the new model may change behaviour on cases that were previously fine, and that has to be checked rather than assumed.
Step 6 â Cut over gradually with a rollback path
Move traffic in stages, watching quality and cost at each step, with a tested path back to the old model.
Tested matters. A rollback plan that has never been exercised is a hope, and the moment you need it is the worst moment to find out the configuration does not work.
Keep the old path available until the new model has seen the tail of your input distribution, which usually takes weeks. Rare inputs are rare, and they are where the regressions you did not anticipate live.
What if the old model is being deprecated?
Then the timeline is not yours, and the work should start as soon as the deprecation is announced rather than near the deadline.
The common failure is discovering late that prompts need substantial rework, leaving no time for shadow running. Starting early converts a forced migration into a normal one, and it preserves the option of choosing a different provider rather than taking whatever the incumbent offers.
How do you handle downstream consumers?
Tell them, and give them the comparison. Downstream systems and teams that built around the old model's output characteristics may need to adjust.
The specific risk is anything parsing structured output. A change in how the model formats a list or a JSON block breaks parsers silently, and the failure appears somewhere unrelated. Check every consumer of the output during shadow running rather than after cutover.
Who needs to be involved?
Whoever owns the system, someone who can run the evaluation, and the owners of any downstream consumers of the output.
A migration run entirely by one engineer without downstream involvement is how parsers break in production.
How long does it take?
Two to six weeks for a well-instrumented system, depending on how long shadow running needs to see the input distribution. Longer if the baseline has to be captured first, because that wait is unavoidable.
What are the common failure modes?
Migrating without a baseline. Comparing on benchmarks rather than your own cases. Assuming prompts transfer. Cutting over all at once. Ignoring cost and latency changes. And forgetting downstream parsers.
How do you know it worked?
Quality on your comparison set at least matching the baseline, no new failure modes at production volume, cost and latency within budget, and downstream consumers unaffected.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Start capturing inputs and outputs from your current model today, even if the migration is months away. That log is the baseline, and it cannot be created retrospectively.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: a behavioural baseline captured before anything changes, shadow running at production volume with cost and latency watched alongside quality, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to migrate from one LLM provider to another.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why do model migrations go wrong?
Because nobody established what the old model actually did. Without a baseline of behaviour on real inputs, there is no way to tell whether the new model is better, worse, or differently wrong, so regressions surface through user complaints.
02What is shadow running?
Sending real production traffic to the candidate model alongside the current one, without using its outputs, so you can compare behaviour on genuine inputs at genuine volume before anything depends on it.
03What should you compare?
Behaviour on the cases that matter: your known failure modes, your highest-volume patterns, and your highest-consequence decisions. Average quality across a benchmark tells you little about whether your specific system will regress.
04Do prompts need changing?
Usually. Prompts are tuned against a model's behaviour, and the same prompt on a different model frequently produces different formatting, verbosity, or refusal patterns. Budget prompt adjustment as part of the migration rather than hoping.
05How long should you keep a rollback path?
Until the new model has run at full traffic for long enough to see the tail of your input distribution â usually weeks rather than days. Unusual inputs arrive rarely and are where regressions hide.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.