Playbook ¡ 6 minute read
How to Run an AI Champion Challenger Test Properly
Champion-challenger testing exposes a portion of real traffic to a candidate and measures the outcome. It works when the metric reflects the business result rather than model quality, the split is random and stable, the run is long enough, and the decision criteria were set beforehand.
Champion-challenger testing answers a question shadow deployment cannot: whether a change produces better outcomes for real users. It also exposes those users to the candidate, which raises the bar for doing it carefully. This playbook covers how, drawing on FISTA Solutions' AI enablement work.
When is this worth doing?
When a change has passed evaluation and shadow comparison, the volume is high enough to measure, and the outcome genuinely depends on user behaviour rather than on output quality alone.
It is not worth doing at low volume, where nothing reaches significance, or for changes where the better version is not in question. Running experiments on obvious improvements is a way of feeling rigorous while delaying value.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Choose the outcome metric | Business result, not model score |
| 2. Set guardrails | What must not get worse |
| 3. Decide criteria and duration | Before starting |
| 4. Split consistently | Per user, randomly, stably |
| 5. Run the full period | Including the weekly cycle |
| 6. Decide against the criteria | Without post-hoc reasoning |
Step 1 â Choose an outcome metric
Pick the business result the system exists to affect: resolution rate, time to complete a task, escalation rate, downstream error rate, or conversion.
Model quality scores are inputs. A change that improves a rubric score while leaving resolution rate unchanged has improved something that does not matter to anyone outside the team.
The metric also has to move within a measurable period. Outcomes that take months to materialise cannot be measured in a two-week experiment, and pretending otherwise produces confident conclusions from noise.
Step 2 â Set guardrail metrics
Decide what must not get worse: escalation rate, complaint volume, cost per task, latency, and any safety measure that applies.
Guardrails catch the changes that improve the primary metric by causing damage elsewhere. A support agent that resolves more cases by refusing to escalate scores well and serves users badly, and only a guardrail on escalation appropriateness catches it.
Set thresholds that stop the experiment automatically rather than requiring someone to notice.
Step 3 â Decide criteria and duration before starting
Write down what result would cause you to adopt the candidate, what would cause you to reject it, and how long the test will run.
This is what prevents the most common failure: stopping when the numbers look favourable. Early results fluctuate, and a test stopped at a promising moment reliably overstates the effect.
Duration should cover at least one full weekly cycle, and long enough to accumulate the events needed for the difference to be distinguishable from noise. Calculate that in advance rather than watching until it looks convincing.
Step 4 â Split consistently per user
Randomise at the user or session level and keep the assignment stable, so the same person gets the same version throughout.
Per-request splitting produces users experiencing both versions, which confuses their behaviour and contaminates the measurement. It also produces a worse experience, because consistency matters to users more than either version's quality.
Check the split is actually balanced on the characteristics that matter. A random assignment that happens to put most enterprise customers in one arm will produce a difference that is not about the change.
Step 5 â Run the full period without peeking at the decision
Monitor guardrails continuously and resist evaluating the primary metric until the period completes.
Guardrails are different: a safety or cost guardrail breaching should stop the test immediately, and that is not peeking, it is the control working.
The primary metric is where discipline matters. Teams that check daily find a moment when the numbers favour their preferred outcome, and stopping then is indistinguishable from choosing the result.
Step 6 â Decide against the criteria you set
Compare the result against what you wrote down, and record the decision with the evidence.
If the result is inconclusive, that is a result. The honest options are running longer, accepting that the change does not matter enough to measure, or shipping it for other reasons and saying so.
What undermines the practice is reinterpreting the criteria after seeing the data. Once a team does that, nobody trusts the next experiment, and the whole apparatus becomes decoration.
What about segment differences?
A change frequently helps one group and harms another, and an aggregate result hides both.
Pre-register the segments you will examine â customer type, language, request category â rather than searching for a segment where the result looks good afterwards. The second approach finds a favourable segment in almost any dataset.
Where a change helps one segment and harms another, the answer is usually routing rather than a single decision.
How does this interact with evaluation?
Evaluation measures output quality against a rubric; experiments measure outcomes with real users. Both are necessary and neither substitutes.
A change that improves evaluation scores and worsens outcomes is telling you the rubric is measuring the wrong thing, which is valuable information about the rubric. Treat that divergence as a finding rather than an anomaly. See how to run an ai evaluation program.
Who needs to be involved?
Someone who can define the business outcome, an engineer to implement the split, and someone with the standing to accept a negative result.
That last role matters. Experiments run by the person who built the change, with no one else empowered to call it, produce predictable conclusions.
How long does it take?
Two to four weeks typically: long enough for the weekly cycle and sufficient events, short enough that the organisation does not lose interest.
What are the common failure modes?
Measuring model quality instead of outcomes. No guardrails. Stopping early. Per-request splitting. Searching for favourable segments afterwards. And reinterpreting criteria after seeing the data.
How do you know it worked?
A decision supported by pre-registered criteria, guardrails held, segments examined as planned, and the result recorded whether or not it was the hoped-for one.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Write down the single business outcome the change is meant to improve, and check whether you currently measure it. Frequently the answer is no, and that is the first piece of work.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: outcome metrics and guardrails agreed before the split is built, decision criteria written down so results are read rather than interpreted, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read what is champion-challenger testing.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What should the metric be?
The business outcome the system exists to affect: resolution rate, time to complete, escalation rate, conversion, or error rate downstream. Model quality scores are inputs to that and not substitutes for it.
02How should traffic be split?
Randomly and consistently per user or per session, so the same person gets a stable experience. Splitting per request produces users who see both versions, which confuses their behaviour and your measurement.
03How long should it run?
Long enough to cover the full weekly cycle and accumulate enough events for the difference to be distinguishable from noise. Stopping early when the numbers look good is the most common way these tests mislead.
04What are guardrail metrics?
Measures that should not get worse even if the primary improves: escalation rate, complaint rate, cost per task, latency, and safety incidents. A change that improves resolution while doubling complaints has not succeeded.
05When should you not run an experiment?
When the change is obviously correct, when the volume is too low to measure anything, or when exposing users to a worse version has real consequences. Shadow deployment covers those cases better.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.