Playbook · 6 minute read
How to Run a Shadow Deployment for an AI System
Shadow deployment runs a candidate against real production traffic without using its outputs, which surfaces behaviour no evaluation set can. It works when side effects are genuinely suppressed, outputs are compared systematically rather than sampled by impression, and the extra cost is bounded.
Shadow deployment is the cheapest way to learn how a change behaves on real traffic before anything depends on it. It works when side effects are genuinely suppressed and the comparison is systematic. This playbook covers both, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
Before promoting a model change, a significant prompt change, a retrieval change, or a new provider â anywhere the evaluation set gives you confidence but not certainty.
It is not worth doing for small changes with strong evaluation coverage. Shadow running costs inference spend and engineering time, and spending both on a change your suite already validates is unnecessary.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Set promotion criteria | Before starting, in writing |
| 2. Suppress side effects | Every tool that writes |
| 3. Mirror a sample | Bound the cost |
| 4. Capture both outputs | With enough context to compare |
| 5. Compare systematically | Against the rubric, not by reading |
| 6. Decide against the criteria | Promote, extend, or abandon |
Step 1 â Set the promotion criteria first
Write down what would make you promote the candidate: quality thresholds on your case categories, acceptable cost, acceptable latency, and what constitutes a blocking new failure.
Deciding afterwards produces post-hoc justification. Teams that run a shadow deployment without criteria tend to promote, because the effort of running it creates momentum.
Include what would make you abandon rather than extend. A candidate that is slightly worse and considerably cheaper is a judgement call, and making it in advance is easier than making it under pressure.
Step 2 â Suppress every side effect
For an agent, this is the critical step. Every tool that writes, sends, charges, or changes state must be replaced with a recording stub that captures what would have happened.
A shadow agent whose tools actually execute is a second production system operating unmonitored, and the failure mode is duplicate actions on real records.
Verify the suppression explicitly before mirroring any traffic. Check each tool individually rather than trusting a configuration flag, because the one that was missed is always the one that writes.
Step 3 â Mirror a bounded sample
Mirroring all traffic doubles inference cost for the duration. A sample of ten to twenty per cent usually provides enough signal.
Sample in a way that preserves the distribution rather than taking the first N requests, which biases towards whatever time of day the sampling started.
Set a budget and an end date before starting. Shadow deployments left running indefinitely become a permanent cost nobody can account for, and they stop being watched within a fortnight.
Step 4 â Capture both outputs with context
Record the input, the current version's output, the candidate's output, and enough context â retrieved documents, tool calls, versions â to understand disagreements.
Without the context, a comparison shows that two versions differed and gives no way to establish why, which converts the analysis into guesswork at exactly the moment it should be precise.
Store them in a form you can query and sample from, not in raw logs somebody has to grep.
Step 5 â Compare systematically
Score both against the rubric, automatically where possible, and sample the disagreements for human review.
Disagreements are where the information is. Cases where both versions produced equivalent output tell you little; cases where they diverged tell you what the change actually did.
Avoid reading a handful of side-by-side outputs and forming an impression. That method reliably favours whichever version writes more fluently, which is not the same as being right. See what is an evaluation rubric.
Step 6 â Decide against the criteria
Promote, extend, or abandon, measured against what you wrote down at the start.
Extending is legitimate when the sample was too small to be conclusive or when a specific category needs more evidence. Extending because the result was disappointing and you hope it improves is not, and the distinction is worth being honest about.
Record the decision and the evidence. The next person considering a similar change benefits from knowing what happened, and so do you in six months.
What about latency-sensitive systems?
Mirror asynchronously so the shadow call is not in the user's path.
A synchronous shadow call adds its latency to every mirrored request, which degrades the experience you are trying to protect and contaminates the latency comparison. Fire the shadow request separately and correlate afterwards.
How does this differ from a champion-challenger test?
Shadow deployment observes without using outputs; champion-challenger serves outputs to a portion of real users and measures the outcome.
Shadow running is safer and answers a narrower question: does the candidate behave acceptably. Champion-challenger answers whether it produces better results for users, which requires exposing them to it.
The sequence is usually shadow first, then champion-challenger for the changes that pass. See how to run an ai champion challenger test.
Who needs to be involved?
An engineer to wire the mirroring and suppression, and someone who can judge output quality in the domain.
The suppression verification is worth a second pair of eyes. It is the step where a mistake has real consequences.
How long does it take?
One to three weeks of mirroring for most systems, plus setup. Long enough to see the input distribution including the weekly cycle, short enough that the extra cost stays bounded.
What are the common failure modes?
Side effects not fully suppressed. Mirroring everything and doubling spend. No promotion criteria. Comparing by reading samples. Synchronous mirroring on latency-sensitive paths. And leaving it running indefinitely.
How do you know it worked?
A promotion decision supported by systematic comparison, no side effects from the shadow version, cost within the budget set, and a written record of what the change did.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
List every tool the candidate can call and mark which have side effects. That list is the suppression work, and it should be done before any traffic is mirrored.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: side effects suppressed and verified tool by tool before any traffic is mirrored, promotion criteria written down before the comparison starts, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to run an AI champion challenger test.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is a shadow deployment?
Running a candidate version against real production traffic in parallel with the current one, without using its outputs for anything. It gives you behaviour on genuine inputs at genuine volume before anything depends on it.
02What must be suppressed?
Every side effect: tool calls that write, emails, payments, record updates, and anything else with an external consequence. A shadow agent whose tools actually execute is not a shadow deployment; it is an unmonitored second system.
03How should outputs be compared?
Systematically, against your rubric, with disagreements between versions sampled for human review. Reading a handful of side-by-side outputs produces an impression biased towards whichever version writes more fluently.
04Does shadow running cost double?
Roughly, for the mirrored portion. Bound it by mirroring a sample rather than all traffic, and by limiting the period. A ten per cent sample over two weeks usually gives enough signal at a tenth of the cost.
05When do you promote?
When the criteria you set before starting are met: quality at or above baseline on the cases that matter, no new failure categories, and cost and latency within budget. Deciding afterwards produces post-hoc justification.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.